On this page
Producer review — first attempts
Review completed on September 7, 2026, Toronto time (September 8 in UTC). The first seven attempts were frozen before the producer's tests. No corrections or guidance were sent to participants. The original source files were not edited. Servers were restarted for the review; several simultaneous startups collided on inspection ports and then worked in sequence. Restart time is not included in participants' durations.
This is a translation of the historical review, not a new test run. The English version is at ../original-en/evidence/PRODUCER-CHECKS.md.
Measurement
Time: the first task_complete.duration_ms for each task, rounded to the nearest second. Includes time waiting for tools, research, and building, not just model generation. All seven runs took place simultaneously on the same machine; compare the observed runs, not universal speed.
Tokens: total_token_usage final accumulated total at the first completion, independently reconciled with the sum of last_token_usage of the distinct accumulated updates. The seven reconcile. total = input + output; reasoning_output is a subset of output, e cached_input is a subset of input. Do not count any of them twice. Ultra includes the main task and three child-task totals, each capped at the main task’s completion time. The child counters start independently at around 30,000, without inheriting the main task’s accumulated total. The sum is 21,730,368. These are processed tokens, mostly cached or reused input—not new context, credits, money, or charges.
Primary source: results.json; the source session hashes were preserved. Exact counts and raw/normalized links are in that file. XHIGH contains eight post or conversation URLs, but only seven conversations: the second X status is an author reply in the same conversation. The others match the conversation counts identified in the record. One original Reddit source for Low and Medium was reopened; the Medium source does describe additional rooms, stairs, and ducts outside the contract. The counts measure the extent of the research record, not validated purchase demand.
Boards: active, non-excluded native Excalidraw elements. Embedded images were counted in the files; drawings and shapes do not count as embedded images. All seven board JSONs are valid; their text was read to check the six required zones and connections among evidence, product decisions, journey, screens, architecture, and MVP. Sol treats visual direction as an additional zone and does not embed raster images. The counts are descriptive, not quality scores.
LOW · Astra / Clearshift
In the browser, the producer opened CS-1042, edited the correction instruction, saved the assignment, and loaded a new sample check. Closure remained disabled until the flagged pattern was reviewed. After completing the check and reloading, the dashboard showed Open 2 and Closed 1. Passed for this correction, new check, and save workflow.
Eight Reddit conversations; none on X. The board has 161 elements and three real app screenshots. The owner’s bounded workflow is coherent, with explicit alternatives and a proposed service/data advantage. More time and recorded processed tokens than Medium, despite lower effort.
MED · Astra / Scopewell
In the browser, in a preserved earlier verification, the producer changed labor from 45 to 60 minutes; the monthly proposal went from $164 to $212. They prepared the proposal, recorded a fictional acceptance, and reloaded: one recorded change and $212 in add-ons persisted. Passed for price changes, approval, and persistence.
Six Reddit conversations; none on X. The board has 133 elements and an embedded interface sketch. Detailed counterevidence, an established paid competitor, day-one utility, and a plan to disprove the idea. It has the fewest elements, not the greatest visual depth. Lowest recorded completion time and fewer tokens processed than Low and High. The author recommends it as a starting point for this research and build task because it delivered enough verified functionality sooner.
HIGH · Astra / Fieldwork
In the browser, the producer marked Cedar unavailable. They reduced Emma’s latest finish time to 10:00, resulting in “No feasible cover”. They restored 12:00 and chose Elm 09:00–11:00 and Birch 14:15–16:15. They reviewed the two changed drafts and saved. After reloading, Cedar was empty, Emma was at Elm, and Priya was at Birch; four other visits were unchanged. Passed for constraint validation, reassignment, and persistence.
Seven Reddit sources and one on X. The board has 186 elements and three app screenshots. A scheduling problem deeper than a simple form; real conditional times and an infeasible state provide meaningful depth. Fixed travel buffer, with no real-time route optimization.
XHIGH · Astra / Scopekeep
In the browser, the producer changed the labor rate from $95 to $110 for three hours; the total became $530. They used a fictional response, recorded approval, and prepared billing. After reloading, a $530 line was ready to bill, at revision 1. Passed for recalculation, approval, billing, and persistence.
Six Reddit conversations and one on X, after removing the duplicate reply. Board with 155 elements and one embedded visual. Strong counterexample: a roofing business owner says Grok and the open-source software Slowbooks already produce service change requests; this directly challenges the proposed business. The producer's attempt to access X failed, so the source remains identified as read by the participant, with no independent re-verification in this review. This is an editorial distinction about business depth, not a universal factual ranking.
MAX · Astra / Fieldnote
In the browser, the producer changed labor to $275,25 and materials to $84,75. The preview showed exactly $360. The simulated approval persisted after reloading; Morgan’s approved add-ons became $840 and the service total became $25.640. Passed for precise pricing, approval, and persistence.
Seven Reddit sources, none on X. Board with 155 native elements and three app screenshots. The model created includes state transition validation, exact cents, and removal of the previous approval on rejected revisions. The participant noticed that Escape closed the board, fixed it, and tested again. The download event wasn't verified by the participant; a selectable export option is available. Longest recorded duration, 46:10, exceeding the self-managed 45-minute limit, including packaging.
HIGH · Sol / RelayOps
In the browser, the producer selected “Split with Crew 1” (11 additional minutes of travel) and then “Prepare updates”. The summary still said “Move Fernbank” and “Brief Crew 2”, matching the compression option. They entered a custom message to the customer and simulated the updates; the result showed fixed values of 0 minutes of travel and 3/3 appointments. Reloading preserved the resolved state.
Partial result: progression between steps and persistence worked; the selected plan doesn't determine the summary or later output. Confirmed change and result JSXs with fixed values in app/page.tsx, lines 115 and 118. The historical instruction was not to correct it before filming.
Six Reddit links and two on X. Board with 215 elements, vector screen diagrams, and no embedded images. The short research log, at 369 words, is supported by the app and board content; don't judge depth by file size alone. Lowest total tokens processed, second-lowest time. Weakest connection between decision and output among the verified paths.
ULTRA · Astra / Accord
In the browser, the producer created a new custom pantry finishing change: 2 hours at $85 plus $30 came to $200. They created a pending request; approval remained disabled until a note was added. They entered a fictional approval and recorded it; after reloading, the title, scope, $200, and approval note remained exact. Passed for creating a new record, calculation, approval, and persistence.
Seven Reddit conversations and two on X. Board with 240 elements, two embedded images, and six main zones. Three Ultra subagents researched Reddit, X, and alternatives; research and alternatives specialists later audited the implementation and extreme numeric cases. Competitor analysis included invoice-only change requests in Joist and reapproval behavior in Jobber, going beyond marketing claims. Best fit here for delegating separate research and verification workstreams; more work processed than Medium, without a cheaper model team. The aggregate includes 7.12 million child tokens.
How to interpret the recommendation
Choose a starting configuration for a bounded research and prototyping task. Medium: shortest observed time with a functional path verified by the producer and a suitable connected plan. High: the next useful test when constraints interact. Ultra: a separate delegation experiment with deeper specialized research. Sol: lowest processed usage, but one concrete failure in the conditional output.
There is no universal model ranking or conclusion about pricing. No participant needed a reminder from the producer to return artifacts; this does not prove that all tasks are free from early interruptions.