🔎 Reading Quality Signals
Install count, source, description, scope, and maintenance—the signals that distinguish a mature skill from an abandoned repo. Plus another practical checklist.
The install count that skills.sh shows next to each skill. It's the most visible popularity signal — find-skills has 1,802,925, frontend-design 488,299.
High install counts correlate with quality, but they aren't proof. There's early-adopter bias and a brand effect. Knowing how to read this helps you avoid installing popular junk.
power law · early-mover bias · brand effect · 0.3% have 100k+
The owner/repo repository that hosts the skill. vercel-labs, anthropics, and microsoft carry institutional reputations; a random 3-star repo does not.
The source is the second-best signal after the description. Official sources keep the skill alive, follow the spec, and respond to issues.
vendor-backed · anthropics/skills · vercel-labs · microsoft/azure-skills · random repo
The frontmatter description field. A good one says WHAT the skill does AND WHEN to use it. It’s the trigger the agent reads to decide whether to activate it.
A clear description shows that the author understood the problem. A vague description means a skill that triggers incorrectly or never triggers. It’s the best predictor of quality.
activation trigger · WHAT + WHEN · precise trigger · sub-trigger
How much the skill tries to cover. An atomic skill solves one problem well; an “everything-about-X” skill tries to solve ten and doesn’t solve any of them properly.
Focused skills are absorbed into context more effectively, trigger precisely, and compose with others. A compact size (body <500 lines) is a sign of discipline.
atomicity · one responsibility · composition · <500 lines
Commit frequency, issues answered, and releases. Since skills are installed via symlink, npx skills update pulls improvements from the source automatically.
A living skill keeps up with changes to the framework and the spec itself. An abandoned one ages and starts giving the agent incorrect instructions.
latest commit · open issues · npx skills update · symlink vs. copy
A 7-question checklist that brings all the signals—installs, source, description, scope, maintenance, license, and fit—together in a 2-minute decision.
Turns intuition into a repeatable process. You stop installing based on hype and start installing based on evidence.
2-minute evaluation · evidence > hype · fit with your stack · read the SKILL.md
🏆 Showcase: the best in each group
Six catalog groups, their reference skills, and why they’re good—using real installs and sources from skills.sh.
UI, React, CSS, and design skills. References: frontend-design (488k, anthropics), vercel-react-best-practices (443k), and web-design-guidelines (358k).
It’s the most mature group for frontend developers. Official sources + very high install counts = low installation risk.
frontend-design · vercel-react-best-practices · web-design-guidelines
The largest group by volume. References: agent-browser (332k, vercel-labs), skill-creator (246k, anthropics), and vercel-composition-patterns (195k).
It’s the heart of the ecosystem. skill-creator is the canonical tool for creating skills—it’s meta and essential.
agent-browser · skill-creator · vercel-composition-patterns
SQL, Postgres, and data modeling skills. References: supabase-postgres-best-practices (203k, supabase), supabase (99k), and neon-postgres (38k).
Here, the source is vendor-backed (supabase, neondatabase): the company that builds the database writes the best practices. Maximum confidence.
supabase-postgres-best-practices · supabase · neon-postgres
The microsoft/azure-skills suite: foundry, ai, deploy, diagnostics, and prepare—all between ~358k and 360k installs.
Few skills, but massive installs (8.3M). Shows how a cohesive vendor suite dominates an infrastructure niche.
microsoft-foundry · azure-ai · azure-deploy · azure-diagnostics · azure-prepare
Testing and quality skills. References: test-driven-development (107k, obra/superpowers) and playwright-best-practices (45k).
Process skills (such as TDD) show that skills aren’t only about stacks—they also encode engineering discipline.
test-driven-development · playwright-best-practices · process skill
Business skills: marketing-psychology (84k, coreyhaines31), lark-slides (139k, larksuite), and stripe-best-practices (stripe/ai).
It proves the ecosystem goes far beyond developers. There are entire groups (Marketing 2.3M, Office, Finance) for non-technical work.
marketing-psychology · lark-slides · stripe-best-practices
⭐ The Best by Quality Signal
For each quality signal, the real skill that’s its perfect example—description, atomic scope, strong source, and clear discoverability, with the catalog’s actual descriptions.
frontend-design (488k, anthropics) says what to do and when, and gives examples: "Use this skill when the user asks to build web components, pages...". supabase (99k) is "pushy" with "Triggers:...".
The description is the routing trigger. Seeing the perfect model gives you a pattern to copy when writing yours.
what + when + examples · “pushy” · Triggers: · combating under-triggering
vercel-react-best-practices (443k) and test-driven-development (107k, from obra/superpowers): one responsibility each, done exceptionally well.
Atomic scope triggers precisely and works well with other skills. It's the signal that separates a mature skill from an "everything-about-X" skill.
one responsibility · sentence test without "and" · composition · process skill
The microsoft/azure-skills suite (~358k each) and vercel-labs: official vendor, cohesive suite, actively maintained via symlink.
Vendor-backed is the highest level of trust—the people who build the product know the edge cases no one documents.
vendor-backed · cohesive suite · npx skills update · actively maintained
find-skills (1,802,925, vercel-labs), the most installed in the catalog: its name and purpose are instantly obvious.
If the name needs explaining, discovery has failed. Self-explanatory name+description become the entry point.
self-explanatory name · entry point · top 100 = 43.7% · 0.3% have 100k+
skill-creator (246k, anthropics) combines a clear description + focused scope + official source + obvious discoverability.
Getting several signals right doesn’t add—it multiplies. That’s why these skills lead the catalog.
compound effect · meta-skill · mutual reinforcement of signals
The reference table: each quality signal with its model skill, source, and real install count.
It becomes a benchmark: when evaluating or creating a skill, compare it with the leader for the signal that matters most to you.
answer key · signal leader · copy the pattern, don't make one up
🛠️ How to Create a Skill That Looks (and Is) High-Quality
The hands-on part: a description that triggers correctly, atomic scope, a polish checklist, before-and-after examples of a weak skill becoming a strong one, and a ready-to-use description template.
Write the description in three parts: the result it delivers, when to use it (concrete triggers), and a pushy stance against under-triggering.
It’s the routing trigger that’s always in context. Get this wrong = a skill that triggers incorrectly or never triggers.
result, not activity · “Use when...” · “Triggers:” · the 100-word rule
Keep the skill focused on a single responsibility, like vercel-react-best-practices and the azure-skills suite. Break up anything large.
An atomic skill is absorbed into context, triggers precisely, is testable, and composes. “Everything about X” fails at everything.
single-sentence test · <500 lines · no "and"/"also" · composition
A 7-item checklist—description, scope, <500 lines, explain why, test prompts, scripts, obvious name—before publishing.
Separates a skill that seems high-quality from one that is. Inspired by the canonical skill-creator workflow.
checklist of 7 · polish = remove · imperative · bundled resources
Same intent ("help with SQL"), two executions side by side: the "sql-helper and more" that nobody installs versus the postgres-best-practices pattern.
Show that the quality leap comes from routing and focus, not technical content.
code box ✗/✓ · "and more" = infinite scope · routing and focus
A frontmatter template that combines frontend-design and supabase patterns, with a completed example for a migrations skill.
Get rid of the “blank page.” You fill in the brackets and calibrate the triggers against near-misses.
template · outcome verb · examples + Triggers · calibrate near-misses
Capture intent → draft → 2-3 test prompts vs. baseline → generalize and trim → optimize the description → package.
Creating a good skill is iterative, not a guess. The loop ensures it actually helps before you publish it.
with-skill vs baseline · should-trigger/should-not · iterate · package
🚀 Advanced Tips: Signals Only Experts See
The finer signals: triggering precision, context efficiency, reading transcripts, eval pass rate and discriminating assertions, variance, and why a popular skill can still be bad.
The skill fires in the right cases (should-trigger) and stays quiet in similar-but-wrong ones (near-misses). A false positive is worse than not triggering.
It’s the invisible signal: a skill that triggers incorrectly injects irrelevant context and degrades the response.
should-trigger · should-NOT-trigger · false positive/negative · precision vs. coverage
The 3 levels of progressive disclosure: the description is always in context, the body loads when triggered, and bundled resources load on demand. A skill that uses too many tokens takes up space.
30 bloated descriptions cost thousands of tokens before any task begins. A dense, concise description is a budget.
3 levels · cost of always-active skills · push details to level 3
Reread real sessions looking for repetition, non-triggering (a skill that should have activated), and incorrect triggering (near-misses).
Transcript is the best free eval. Each pattern you find is a concrete adjustment to a skill or its description.
repetition → script · no trigger → weak trigger · wrong trigger → near-miss
A discriminating assertion fails without the skill and passes with it. The number that matters is the delta between with-skill and baseline.
A 100% pass rate that doesn’t distinguish the baseline from the with-skill results is false reassurance—it doesn’t measure the skill’s contribution.
discriminating assertion · delta vs. baseline · always run the baseline
The model is stochastic: the same skill passes one run and fails another. Run the eval several times and look at the mean ± standard deviation.
90% flaky is worse than 80% stable. Stability is quality; a single run is misleading.
stochastic · mean ± σ · imperative reduces variance · explain why
Power law: the top 100 account for 43.7% of installs; only 0.3% (131 skills) exceed 100k. There’s early-mover and brand bias.
A popular skill may be abandoned, built for a different stack, or bloated. Install is where you start investigating, not where you finish.
early-mover bias · stack fit · maintenance · evidence > hype
Learning path overview
Installs mislead; the description doesn't. Learn to separate hype from evidence.
Six groups, the reference skills, and why—with real numbers.
Each signal has a champion. find-skills, frontend-design, azure: emulate the models.
A description that triggers, atomic scope, and the ready-to-use template. From weak to strong.
Precise triggering, context cost, honest evals. Popularity isn’t proof.