🔒 Internal Skills (Private)
Not every skill belongs in the public catalog. Proprietary knowledge, internal conventions, and your team’s workflows live as internal skills — marked with metadata.internal and installed only with the flag INSTALL_INTERNAL_SKILLS. They're outside skills.sh and the public install count.
mark a skill as internal:
--- name: deploy-runbook-acme description: Runbook interno de deploy da Acme. Acione em deploy, rollback ou incidente em produção. metadata: internal: true --- # só instala com a flag explícita: INSTALL_INTERNAL_SKILLS=1 npx skills add acme/skills-privadas
Good candidates for internal use
- ›Team deploy/incident runbooks
- ›Proprietary coding conventions
- ›Integrations with internal systems
Why separate it from the public
- ›Prevents sensitive context from leaking into the catalog
- ›Doesn’t clutter skills.sh with noise that’s just for you
- ›Install only with explicit opt-in, never by accident
🛡️ No Surprises and Security
The golden rule: the skill never does what the description doesn’t promise. Because skills are installed via symlink and updated with one command, a malicious or careless repo can contaminate many people at once. The trust of the entire ecosystem depends on this.
✗ Surprise / risk
- ✗Script that exfiltrates data or sends hidden telemetry
- ✗
rm -rforgit push --forcehidden in one step - ✗Accesses secrets/credentials without the user asking
- ✗Downloads and runs code from an external source at runtime
✓ Reliable
- ✓Does exactly what the description says
- ✓Destructive actions require explicit confirmation
- ✓Least privilege: touches only what’s necessary
- ✓Auditable scripts, with no obscure dependencies
💡 Audit it as someone installing it would
Before publishing, read your own scripts/ through the eyes of a skeptical user. No malware, no exfiltration, no surprise side effects. One breach of trust burns the reputation of the entire repo—and spills over to skills.sh.
🏷️ Versioning strategy
At scale, "push to main" without a strategy creates instability. Because the symlink propagates every change immediately, you need versioning discipline that separates what's safe from what's risky.
What each type of change requires
Text edits, body clarification. Low risk — can go straight through, with a quick eval.
New additive behavior. Still compatible—run full evals first.
A change to the trigger or expected behavior. Breaks for those who depended on it — communicate and document in the CHANGELOG.
tag a release and maintain a CHANGELOG:
## [1.2.0] - 2026-06 ### Changed - description agora cobre o near-miss de "refatorar" (minor) ### Fixed - corpo: removido MUST que causava overfit em projetos sem testes git tag v1.2.0 && git push --tags
The trigger-change rule
Change the description is the most dangerous change there is: it changes when the skill triggers for everyone who installed it. Treat every description edit as a potential major until the evals prove otherwise.
👥 Team rollout
Sharing skills with a team is different from publishing them to the world. You want consistency (everyone using the same set), version control, and an adoption path that doesn't catch anyone by surprise.
A curated team repo
Centralize approved skills in a single public or internal repo. It becomes the single source of truth — no one installs random skills directly in the shared project.
Pilot before the general rollout
Have a small group install and use it for a week, then report back. Adjust the trigger and body based on real feedback before rolling it out to everyone.
Communicate before each update
How npx skills update propagates everything; notify the team before trigger changes. A changelog in a shared channel prevents the question, "why did the skill start triggering on this?".
💡 Onboarding pro tip
Document the team's skill set in the curated repo's README, with a npx skills add line by line. Anyone joining can follow the list and be productive on day one, with the same agent behavior as everyone else.
📈 Measure adoption
A skill at scale is a product—and products are measured. Two metrics matter: install count (how many adopted it) and triggering (whether it triggers in the right cases). One high and the other low tell very different stories.
Install count
Measures discovery and adoption. Low installs with good triggering = a naming/description or niche problem, not a quality problem. Don’t chase the leaderboard.
Triggering
Measures whether the skill appears at the right time. High installs + poor triggering is the worst case: lots of frustrated users. Measure with should-trigger / should-not-trigger / near-miss evals.
♻️ Keep skills alive at scale
A skill is easy to maintain. Twenty skills become a portfolio—and a portfolio without maintenance goes stale. Pro tips for keeping dozens of skills useful without turning them into chaos:
✗ Portfolio that rots
- ✗Duplicate skills with overlapping triggers
- ✗No evals — breakages go unnoticed
- ✗Dead skills nobody uses, cluttering search
- ✗Bodies that have grown to 800 lines over time
✓ Living portfolio
- ✓Each skill is atomic, with non-overlapping triggers
- ✓Eval suite runs in CI on every push
- ✓Periodic review: retire what doesn’t trigger
- ✓Lean body; details go in references/
Final pro tips
- ★Evals in CI: run should-trigger / should-not-trigger automatically on every PR — catch trigger regressions early.
- ★Audit for overlap: Two skills triggered by the same prompt compete and cause confusion. Merge them or differentiate their triggers.
- ★Retire without mercy: a skill with a stalled install count and poor triggering just adds noise. Deprecate it in the CHANGELOG and remove it.
- ★Keep it concise: with every release, ask what can be cut. A body under 500 lines isn’t a goal; it’s hygiene.
💡 The closing
You’ve completed all 5 tracks: ecosystem overview, quality and the best skills, anatomy, the creation loop, and the mental models for decisions, publishing, and scaling. Now it’s time to practice: choose a repeatable workflow, write, publish, measure, and iterate.
✅ Module Summary
Next:
Module 5.6 — 2026 Rules: degrees of freedom, description limits, hooks in the skill, the 5,000-token cutoff, and the 10-rule checklist with the validator.