🩺 Audit for real
So far, you've learned to classify each line on paper. Now comes the operational question: how do you run the audit on your own config and turn the report into safe cuts? The skill reads and diagnoses; you apply the changes in another session, with the config under git.
What to look at: the solid line (purple) is what the skill does on its own—read the config and produce a report. The dashed line (cyan) crosses a boundary: everything to the right is a human decision, made in another session. An audit that starts making changes right away is an audit you can’t trust.
Track map
Detailed content
🩺 Running the audit-ablacao skill
Install and run the skill on your own configuration, and read the 10-section report knowing what each section requires of you.
The skill reads your CLAUDE.md (project-level and global), your skills, your hooks, and the settings.json, classifies each instruction and proposes a minimal version. It never edits, moves, deletes, overwrites config, or commits.
Because the boundary is what makes the audit reliable. You wouldn’t let a tool that starts making changes run globally — and then you’d never audit anything. Knowing it’s read-only means you can run it without worry, even before git is in the config.
Diagnosis ≠ treatment. Applying changes is a separate request, by design. And don’t confuse it with memory-audit: one handles the session’s stored memory, while the other handles the agent’s configuration.
There are two possible destinations: ~/.claude/skills/audit-ablacao/SKILL.md— which applies to all projects, or .claude/skills/audit-ablacao/SKILL.md within the repository, where it applies only there. In both cases, it’s a cp of the SKILL.md and a session restart for Claude Code to load it.
Choosing the destination is already a config architecture decision. A skill inside the project is versioned with the code, goes into the PR, can be reviewed, and doesn’t spill over into other projects. A global skill is available in any session, but takes up a slot in the list of skills every project sees.
The repository root already contains the SKILL.md installable — this is why the guide stays in guia/ and not in the root. If the skill doesn't appear, you almost always need to restart the session.
Three scopes with distinct triggers. Global (~/.claude/): about every ~6 months or with a major model release. A project: when the CLAUDE.md of it exceeds ~150 lines. A set of skills: when 3 or more compete for the same trigger.
Too broad a scope produces a report you can't act on; too narrow a scope hides redundancy, which is almost always between global and project-level instructions. Auditing everything at once is the most common way to end up cutting nothing.
Triggers are measurable signals, not impressions: number of lines, number of conflicting skills, model release. An objective signal is what keeps the audit from becoming a task that never happens.
Executive summary, metrics, problems by file, candidates for removal, redundancies and conflicts, skills, CLAUDE.md proposed minimum, proposed skills, ablation test plan (versions A/B/C), and Top 10 changes by impact ÷ risk.
Each section asks something different of you. Metrics and issues by file are for reading; removal candidates and conflicts require your judgment; the CLAUDE.md the minimum is a proposal, not a verdict; the Top 10 is the only section that becomes an immediate work plan.
Impact ÷ risk orders the Top 10 because the goal isn't to reduce as much as possible, but to maximize quality + autonomy + verifiability ÷ complexity. High-impact, high-risk cuts come later, with testing.
Each skill gets one of seven verdicts: KEEP, SIMPLIFY, MERGE, SPLIT, LOAD-ON-DEMAND, CONVERT-TO-CONTEXT or DELETE-CANDIDATE. It’s a separate axis from the line-by-line verdict in CLAUDE.md.
A skill has a boundary, name, and scope—it’s easy to audit and retire, unlike “that paragraph in the middle of the CLAUDE.md". That's why a skill is the right unit to fix: you can disable one and measure the effect.
MERGE resolves skills that compete for the same trigger; SPLIT resolves a skill that does too many things; CONVERT-TO-CONTEXT is for what was information disguised as a procedure; LOAD-ON-DEMAND removes the context cost from runs that don’t use it.
Run /audit-ablacao within the selected scope, take the first three candidates for removal, open the cited file, check the passage in context, and make your own call: agree, disagree, or send it to TEST.
This is the habit that separates an audit from blind faith. Checking the passage in the original file catches the report's two most common errors: a quote taken out of context and a line that seems redundant but contains a detail found only there.
When in doubt, TEST — never REMOVE. And the audit deliberately preserves what the model can’t infer: project identity, paths and sources of truth, branding, security, compliance, integrations, and interface contracts.
🔧 From report to cuts: the skill is the right unit
Turn the Top 10 into safely applied changes by moving procedures from CLAUDE.md into on-demand skills.
The report is produced in one session; the cuts happen in another. Before applying anything, the config needs to be under version control—a clean commit first, then a commit for the cut, so reverting is a command, not an archaeological dig.
Applying changes in the same session you audited contaminates your judgment: the model already has the full report in context and tends to defend its own conclusions. A new session reads the config as it is, not as the report said it was.
Reversibility is a prerequisite for courage. With git in the config, an aggressive cut costs a git revert; without git, it’s hard to remember what was written — and nobody does.
Start with redundancies and conflicts — the same rule in three places, two rules that contradict each other. Then tackle micromanagement and outdated legacy instructions. Last, address what was marked as TEST— which only goes after it’s been measured.
Redundancy and conflict are the highest-impact, lowest-risk cuts: deleting the copy doesn’t change behavior because the rule still exists in one place. Starting with them gives you a real reduction without risking anything—and clears the way to assess the rest.
Conflict is worse than redundancy: with two contradictory rules, the model picks one, and you don’t know which. Resolving a conflict is your decision about which rule applies, not an automatic cut.
A rule that lives in the CLAUDE.md is read on every run, including the 90% that have nothing to do with it. The same rule inside a skill only costs context when the task calls for it. That's exactly the MOVE e o LOAD-ON-DEMAND of the report.
It's the cheapest way to slim down the config without losing anything—no instructions are discarded; they just move somewhere else. If you're afraid to delete things, moving them is the first cut you can make.
The partitioning criterion: the CLAUDE.md keeps only what's true always — identity, guardrails, sources of truth, security. Everything else (procedure, format, recipe, integration) becomes a skill.
Invoke via /nome-da-skill instead of hoping the description matches your phrasing. When several skills compete for the same topic, explicit invocation breaks the tie.
This eliminates an entire category of bloated lines: the routing rule in the CLAUDE.md (“when the user asks for X, use skill Y”). If you call it directly, the routing rule doesn’t need to exist — and it was read on every run.
Description-based triggers are probabilistic; explicit invocation is deterministic. Replacing one with the other reduces context and increases predictability at the same time — a rare win on both fronts.
When the model stumbles, there are three remedies. Better prompt: the instruction was unclear. Skill: a repeatable procedure is missing. MCP: context it can't access on its own is missing.
Choosing the wrong one bloats the CLAUDE.md. The most common reflex is to dump one more global rule onto a problem caused by a missing tool or procedure—and that rule stays there forever, reread every time it runs.
The diagnosis comes from the observed failure, not a guess: if the model didn’t know something, it’s context/MCP; if it knew but did things out of order, it’s procedure/skill; if it misunderstood the request, it’s the prompt.
The track’s closing exercise: remove three items from the Top 10, apply them in a separate session, and record the before and after using two numbers — lines of the CLAUDE.md and the number of rules—plus one more sentence about the observed behavior.
Without a record, you don’t know whether the config got smaller or you just reorganized it. And the success criterion isn’t reduction: it’s unchanged behavior with fewer lines. If something broke, you can identify the wrong cut because there were only three.
Three at a time is intentionally a small batch—that’s what keeps the cause identifiable. Then comes using it in real work for a few days, which is where Track 4 takes up the subject with the A/B/C plan.