Reference · 2026-09-24

Effort in practice

Two experiments ran the same request at every effort level: one with Claude Opus 5.5 and one with GPT-6 Astra. Both show the same pattern: the top level was never the preferred one.

This page is a reference: it does not change the stack or the guide rules. It is third-party data, with one run per level and one person judging. Text sources: padroes-esforco · Opus 5.5 video · astra-effort.
Source A · Claude Opus 5.5

The same 3D app at six levels

Task: a long /goal that turns ~105 GB of recordings from a virtual event into an explorable 3D conference. The author preferred xhigh (“Extra” in the video). Max cost twice as much and had more bugs.

12.9×
cost of max relative to low
2.3×
increase in checks over the same range
1 / 6
runs that asked any question (0 used a subagent)

How much each metric grows (low = 1×)

Cost and time shoot up; checks barely move.

Show table
LevelTokensCostTimeChecks
low1.001.001.001.00
medium2.193.184.371.05
high2.934.174.011.00
xhigh ★3.796.635.381.55
max6.1812.888.852.32
ultracode—4.785.681.91

Estimated cost per level (US$)

Estimated at API prices; the author ran it on a subscription. Highlighted: the preferred one.

Show table
LevelTimeCostTokensChecksQuestions
low16m43s3.91191k220
medium1h13m12.44419k230
high1h07m16.31~559k221
xhigh ★1h30m25.92~723k340
max2h28m50.381.18M510
ultracode1h35m18.69?420

~ = number spoken in truncated form in the video. Ultracode tokens were said as “66 thousand”, which does not match a 1h35m run; they are treated as unknown.

low → medium

Biggest perceived jump: correct branding, real videos playing, booths and a VIP area.

high → xhigh

Interaction gains: talk to characters, sit in talks, open resources.

xhigh → max

2× the cost and +1h of runtime, with poor walking, glitches and a video that did not play. It regressed.

Source B · GPT-6 Astra

Research + micro-SaaS under seven conditions

Data from the kit translated in astra-effort: the same research-and-build task, 30 to 46 minutes. The author preferred medium, which was the fastest and still delivered the verified flow. Sol high and Astra ultra (with 3 subagents) appear as separate conditions.

Time per condition (minutes)

Low took longer than medium. Low effort does not guarantee less work.

Output tokens (thousands)

Always rises with effort. The reasoning budget gets spent whether or not it pays off.

Show table
ConditionTimeProcessed tokensOutput tokensProducer note
low37m37s14.81M51.5kmore time and tokens than medium
medium ★30m41s10.23M42.2kpreferred: verified flow in the shortest time
high41m21s14.33M57.4kgood with interacting constraints
xhigh45m14s12.29M67.7kfound evidence against its own idea
max46m10s10.55M70.2kthe longest; precision in values and rules
ultra (3 children)42m10s21.73M96.0k+7.12M from children; similar product
Sol high32m49s6.70M54.2kcheap, but one broken decision path

Processed tokens include ~98% cached input and do not represent cost. Do not compare them with Source A tokens in absolute terms: compare only the shape of the curves.

Crossing the sources

Eight patterns

Includes Source C, an internal test recorded in maestro-roteador: from high to max, “the difference was a favicon, for 2–5× the tokens”.

The top level was never preferred

Preferred: xhigh on Opus and medium on Astra; in maestro, high was practically equal to max. In the Opus video, max regressed.

The sweet spot depends on the task

Bounded ~40-min task → medium. Hours-long autonomous build with lots of material → xhigh. What decides is how much there is to read and decide.

Reasoning is always spent; the payoff is not always there

Output tokens rise monotonically. Cost rises 12.9× from low to max. The budget is consumed either way.

Total cost is not linear

Astra low took longer than medium; Opus high finished before medium. Low effort does not guarantee less work.

More checks do not mean fewer bugs

Opus max ran 51 checks and still had more glitches than xhigh, with 34.

Delegation is a separate condition

Astra ultra added +7.12M tokens with 3 subagents without changing the product. Opus ultracode delegated nothing and behaved like xhigh.

Effort does not buy questions

There was 1 question in 6 runs. Ambiguity has to be resolved in the request, not by raising effort.

Effort buys deliberation, not taste

The jump in “feel” came between low and medium. From there on, the gain was interaction and detail, which matches maestro’s model × effort axis.

🧭 Nei’s take: how I use effort

medium · default→high · when more reasoning is needed→xhigh · extreme cases

Running Opus 5.5 on medium is the best option for everyday work. I go up to high when the task needs more reasoning and use xhigh only in extreme cases.

The same goes for GPT-6 Astra. Max stays out: in both sources it cost more and did not deliver more.

⚖️ Conclusion

Effort is an exploration budget, not a quality knob.

Start at the lowest plausible level; when in doubt, medium. Go up only with evidence of an insufficient result, and treat max/ultra as a justified exception. This rule is already in the guide and in maestro-roteador.

What this material adds is an observed ceiling: in long autonomous builds with Opus 5.5, xhigh was the best and max did not pay off.

Data limits

  • One run per level in each source; run-to-run variance was not measured.
  • Visual judgment by one person. Source A has a sponsor; Source B was translated without a new run.
  • Source A costs are estimates at API prices.
  • There is no INEMA test of our own yet. To run one, use avaliacao/bateria.md.