Two experiments ran the same request at every effort level: one with Claude Opus 5.5 and one with GPT-6 Astra. Both show the same pattern: the top level was never the preferred one.
Task: a long /goal that turns ~105 GB of recordings from a virtual event into an explorable 3D conference. The author preferred xhigh (“Extra” in the video). Max cost twice as much and had more bugs.
Cost and time shoot up; checks barely move.
| Level | Tokens | Cost | Time | Checks |
|---|---|---|---|---|
| low | 1.00 | 1.00 | 1.00 | 1.00 |
| medium | 2.19 | 3.18 | 4.37 | 1.05 |
| high | 2.93 | 4.17 | 4.01 | 1.00 |
| xhigh ★ | 3.79 | 6.63 | 5.38 | 1.55 |
| max | 6.18 | 12.88 | 8.85 | 2.32 |
| ultracode | — | 4.78 | 5.68 | 1.91 |
Estimated at API prices; the author ran it on a subscription. Highlighted: the preferred one.
| Level | Time | Cost | Tokens | Checks | Questions |
|---|---|---|---|---|---|
| low | 16m43s | 3.91 | 191k | 22 | 0 |
| medium | 1h13m | 12.44 | 419k | 23 | 0 |
| high | 1h07m | 16.31 | ~559k | 22 | 1 |
| xhigh ★ | 1h30m | 25.92 | ~723k | 34 | 0 |
| max | 2h28m | 50.38 | 1.18M | 51 | 0 |
| ultracode | 1h35m | 18.69 | ? | 42 | 0 |
~ = number spoken in truncated form in the video. Ultracode tokens were said as “66 thousand”, which does not match a 1h35m run; they are treated as unknown.
Biggest perceived jump: correct branding, real videos playing, booths and a VIP area.
Interaction gains: talk to characters, sit in talks, open resources.
2× the cost and +1h of runtime, with poor walking, glitches and a video that did not play. It regressed.
Data from the kit translated in astra-effort: the same research-and-build task, 30 to 46 minutes. The author preferred medium, which was the fastest and still delivered the verified flow. Sol high and Astra ultra (with 3 subagents) appear as separate conditions.
Low took longer than medium. Low effort does not guarantee less work.
Always rises with effort. The reasoning budget gets spent whether or not it pays off.
| Condition | Time | Processed tokens | Output tokens | Producer note |
|---|---|---|---|---|
| low | 37m37s | 14.81M | 51.5k | more time and tokens than medium |
| medium ★ | 30m41s | 10.23M | 42.2k | preferred: verified flow in the shortest time |
| high | 41m21s | 14.33M | 57.4k | good with interacting constraints |
| xhigh | 45m14s | 12.29M | 67.7k | found evidence against its own idea |
| max | 46m10s | 10.55M | 70.2k | the longest; precision in values and rules |
| ultra (3 children) | 42m10s | 21.73M | 96.0k | +7.12M from children; similar product |
| Sol high | 32m49s | 6.70M | 54.2k | cheap, but one broken decision path |
Processed tokens include ~98% cached input and do not represent cost. Do not compare them with Source A tokens in absolute terms: compare only the shape of the curves.
Includes Source C, an internal test recorded in maestro-roteador: from high to max, “the difference was a favicon, for 2–5× the tokens”.
Preferred: xhigh on Opus and medium on Astra; in maestro, high was practically equal to max. In the Opus video, max regressed.
Bounded ~40-min task → medium. Hours-long autonomous build with lots of material → xhigh. What decides is how much there is to read and decide.
Output tokens rise monotonically. Cost rises 12.9× from low to max. The budget is consumed either way.
Astra low took longer than medium; Opus high finished before medium. Low effort does not guarantee less work.
Opus max ran 51 checks and still had more glitches than xhigh, with 34.
Astra ultra added +7.12M tokens with 3 subagents without changing the product. Opus ultracode delegated nothing and behaved like xhigh.
There was 1 question in 6 runs. Ambiguity has to be resolved in the request, not by raising effort.
The jump in “feel” came between low and medium. From there on, the gain was interaction and detail, which matches maestro’s model × effort axis.
Running Opus 5.5 on medium is the best option for everyday work. I go up to high when the task needs more reasoning and use xhigh only in extreme cases.
The same goes for GPT-6 Astra. Max stays out: in both sources it cost more and did not deliver more.
Effort is an exploration budget, not a quality knob.
Start at the lowest plausible level; when in doubt, medium. Go up only with evidence of an insufficient result, and treat max/ultra as a justified exception. This rule is already in the guide and in maestro-roteador.
What this material adds is an observed ceiling: in long autonomous builds with Opus 5.5, xhigh was the best and max did not pay off.
avaliacao/bateria.md.