PTENES
← Astra Effort Library
On this page

I tested every GPT-6 Astra effort level. Here’s what I’d use — YouTube

https://www.youtube.com/watch?v=OQipTxv9Qv0

Translated transcript in Portuguese. Opinions and numerical approximations are the narrator’s. The original automatic transcript is at original-en/docs/transcricao-video.txt; obvious names such as “GBT6,” “Astro,” “Soul,” and “Excel draw” were normalized to GPT-6, Astra, Sol, and Excalidraw. For exact numbers and methodological limitations, see SUPER-GUIDE.md and evidence/results.json.

(00:00) GPT-6 Astra has been available for a while, and one of the biggest discussions on YouTube and X is this: which effort level gives you the best value or return for what you spend on this model? I’ve also seen people say it’s inherently lazy.

(00:17) It does what you tell it to do, but nothing beyond that, and sometimes not even that. So I ran some experiments to see whether changing the effort level also changed this behavior. I tested every Astra level on the same research and build project and included Sol for comparison. In this video, I’ll show you what each one delivered, how long it took, how many tokens it used, and whether effort made any difference.

(00:42) By the end, you’ll be able to judge which level suits your everyday projects. Let’s get started. This post on X was what made me want to run the test. Tibo leads OpenAI’s core products and is also known as someone who seems to reset his usage allowance every hour. He made a bold claim.

(01:00) GPT-6 Astra in Low performs better than GPT-6 Sol in High. Sol is the previous model and is being forgotten because of this new release. But for many goals and tasks, it’s still more than enough. When you look into real experiences and reports about this model, especially in Low, you hear that it stops too soon and asks permission every five minutes.

(01:26) I wanted to evaluate two things: does extra effort necessarily improve results? And does it offer more autonomy, making the model stop asking permission and seeming lazy? In OpenAI’s official documentation, we find the following.

(01:42) The model was designed to collaborate better, so it tends to ask the user more questions. The documentation also says this can make it stop when the user expected reasonable assumptions. Something like Fable 5.1 is very good when you give it an overview, a summary of the goal, and it not only extrapolates, but also adds steps it considers appropriate.

(02:06) It seems that Astra, at least at some effort levels, may not do this as well. To make the experiment as transparent and balanced as possible, I controlled six variables: the same context, the same access to tools, MCPs, and skills, the same build scope, and the same conditions. And I separated Ultra because it often creates a series of subagents.

(02:29) I’ll show you a little trick for opening several conversations with different effort levels using the same prompt. And finally, we recorded how many tokens and how much time each conversation used. This was the prompt sent to all seven conversations; I’ll explain it piece by piece.

(02:47) The gist was to let each agent research everything it could on X and Reddit, two great places to hear user feedback. The target customer was this: find a worthwhile SaaS opportunity for small business owners in the service industry with 2 to 20 employees.

(03:07) I didn’t choose the category, problem, or product. Start with the problems, consider three opportunities, and choose one. For the evidence, we asked for research on these platforms and an analysis of three existing alternatives to the initial idea. The goal was to gather recurring frustrations, makeshift solutions, signs of spending, and counterevidence.

(03:29) This is the most important sentence in the prompt for evaluating the opportunity: assume that customers and competitors can use Astra to quickly reproduce competent software. In other words, don’t build a SaaS that could be made with vibe coding in a weekend. Ideally, the business model should have some defensible advantage.

(03:47) For the planning, I wanted it to create its own Excalidraw files using AI to generate the corresponding JSONs. With more advanced models, I no longer need a skill: I can ask for an Excalidraw file. It’s a format that lets you annotate and create diagrams as if you had drawn them by hand.

(04:08) Instead of a skill, I wanted each level to demonstrate creating the entire file so I could observe how it organized its reasoning. For the main product, I wanted a polished website, similar to the sites I prepare to explain these concepts on YouTube.

(04:26) I wanted a ChatGPT site that explained the full build plan after the research. In addition to the site and the Excalidraw file, I wanted a working SaaS prototype. Splitting the task into these parts would give each level room to show real differences, if any existed.

(04:48) The central questions were: can we trace the reasoning that this is a good idea? Why would anyone pay? Do the pieces connect—the chosen profile, the stated pain point and its evidence, the product? Does the workflow work? Does it make sense as a SaaS? And for the conclusion, did it finish and verify its work? Did it reflect on what happened? These are all the conversations we ran.

(05:17) I’d be lying if I said I opened them all manually. Just ask Codex for a series of conversations and tell it to rename them by effort level, change the effort, and link them to a specific folder. You can ask something like this.

(05:38) I want you to use the same prompt; you can create it. It should be a very simple task. Create three separate conversations with easy-to-read names in the sidebar, each with an Astra level: Low, Medium, and High. Run this prompt in all three at the same time, now.

(05:58) And track each one’s progress. This is a recent feature that works in both Claude and Codex. When you open Recent items, you can see the new test conversations starting to appear separately. There they are.

(06:17) Now we have Astra Medium, Low, and High, all running. When you open them, you see the same task. That’s how you start testing at scale. The result of each conversation was a board like this, a site, a full analysis, and statistics on tokens, time, and the main outcome.

(06:39) Astra in Low took 37 minutes, used almost 15 million tokens, went through eight Reddit conversations and none on X, and created 161 elements on the board. This doesn’t necessarily measure quality; it shows how much work it produced. The product tracks cleaning tasks that haven’t been completed through to their correction and reinspection, for commercial cleaning business owners who regularly service offices.

(07:03) The proposed advantage is customer-specific patterns and a history of fixes that worked. Interesting. This is what the site looked like. We’re not evaluating beauty, but the core concept. This is the draft for this SaaS. It’s not very impressive, but it might be a viable business.

(07:23) The big gap is that there’s no evidence from X, and it didn’t go beyond what was requested. Astra Medium took just 30 minutes, 10 million tokens, and six Reddit conversations. The product calculates prices for extra cleaning services and records customer acceptance, for companies with written contracts.

(07:42) The proposed advantage is accurate scope records and estimates improved by actual services. I don’t know how defensible an advantage that is. What’s interesting is that we’re once again in the cleaning industry, even though I didn’t specify it. Maybe, given the profile and scope provided and all the conditions, this seems like the best industry.

(08:02) The site is similar. When you explore the plan and open the board, the breakdown seems clearer: customer, alternatives, opportunity, and owner journey. It has a good analysis of the implications. I’d need to read everything to assess the business model in detail, but the quality and variety of the analyses seem better than Low, with less time and fewer tokens.

(08:31) Astra High took close to 40 minutes and 14 million tokens. It researched Reddit and finally found a post or group of posts on X, and created the most detailed board. The idea was to reschedule visits when a cleaning crew can’t work.

(08:50) Maybe this is the simplest description so far: for residential cleaning business owners who coordinate crews and schedules. The proposed advantage is company-specific constraints and outcomes that improve scheduling. So far, the site is the clearest about the vision and the pain point. The pain point appears here.

(09:09) At 7:42, Cedar’s van won’t start. There are two recurring visits that need coverage; we need to fill time slots that would otherwise be canceled. The idea tries to reduce lost revenue. But was it worth spending 40 minutes and 14 million tokens, or could you use Medium’s 30 minutes and a more precise prompt, or a follow-up, to get to the same point? Extra High took 45 minutes, but used a little less than High: 12 million.

(09:40) What’s interesting, and why I say this isn’t AGI and consider many of these claims exaggerated, is that effort level should somehow correspond to efficiency, but it doesn’t. High used close to 14 million tokens.

(09:59) Low used close to 15 million and Medium close to 10 million. There’s variation and unpredictability. If I repeated this, I’d say the times would probably be similar, but token use is increasingly unpredictable.

(10:16) This time, the business isn’t cleaning. It takes renovation add-ons from approval to invoicing, for small business owners who lose track of what to charge. The proposed advantage is scope checks tailored to approval, invoicing, and payment history. A long description. The site looks on par with High, but gets worse when you examine it closely.

(10:39) It has a similar layout and four similar tabs. The connected board opens in a new tab, a nice touch. There are filters and buttons for navigating. But in scope and appearance, the board seems a little worse than High’s. This was the High version.

(11:01) It also has buttons, in a different position and a little more tightly spaced, and more evidence for each claim. When you click on customer evidence, you go straight to the associated conversation. There’s a better breakdown and more sources linked to the exact conversations. Side by side, the Extra High version is simpler and less detailed.

(11:29) It has fewer references and slightly better navigation. Other than that, I don’t see any added benefit for the extra five or ten minutes and possible tokens over Medium and High. Max confused me: it took 46 minutes but used only 10 million tokens.

(11:50) It was token-efficient, but a little time-inefficient to reach that usage. The idea keeps the scope, price, and approval of extra services in one record. It seems focused on databases or data organization, for residential renovation business owners with teams of 2 to 10 people.

(12:09) The proposed advantage is an on-site workflow, assisted setup, and referrals from accounting professionals. Interesting: it considers a business supported by referrals. I want to point out that even in Max, there are no references to X, although they appeared in High.

(12:27) More evidence that using more tokens—or, here, taking more time to process them—doesn’t necessarily improve the result. The site is cleaner and more polished. It looks like a real product compared with High and Extra High. It conveys the vision and the problem we’re trying to solve.

(12:49) But would you spend an extra ten minutes instead of adding one or two prompts to a lower level? We’ve reached Astra Ultra: 42 minutes, less than Max, but an impressive 22 million tokens for this result. This is what you get with 22 million tokens.

(13:10) It records and combines renovation add-ons before work begins. Basically, it automates service change requests for small residential renovation business owners who handle frequent extras. The proposed advantage is specialty-specific setup and a proven workflow improved by real outcomes. The site isn’t a Da Vinci masterpiece.

(13:29) It looks a lot like the previous ones, with minor changes to spacing and polish that don’t add much to the experience. When you explore the opportunity, there’s a raised button and a board almost identical to what we’ve already seen.

(13:47) There are two sources in an odd position: an extra set of sources per section. Is this worth the extra time and tokens? Twenty million tokens and subagents to put this together. As a control comparison, we also tested Sol High, because based on Tibo’s post, Astra Low should offer an equivalent experience. Sol took just 32 minutes.

(14:12) Just 6 million tokens. If you’ve forgotten Sol, here’s a reason to remember it. Astra shouldn’t be used for everything. You don’t need to bring a nuclear bomb to every fistfight. Surprisingly, it researched six Reddit conversations and two on X. Quantity doesn’t necessarily mean quality, but at least it consulted those sources.

(14:36) The business plan organizes recovery when delays affect a service route, very similar to the previous one. The audience is owners who also coordinate operations at recurring service businesses with 2 to 20 employees. The advantage is recovery outcomes, maintained rules, and responsible support. I hope the site won’t be as pretty, but will be functional.

(14:56) Okay, this is incredibly ugly; it’s so bad it makes you want to close your eyes or squint. I think the first image is trying to show a business model diagram. I can easily say Astra is much better at websites and quickly creating this kind of ChatGPT site.

(15:17) But the real question is: could you use a skill that creates a beautiful site with proper spacing, and stick with Sol High or Medium? Based on the experiment and my everyday use, I love Astra Medium in Fast. It’s my favorite setting across the Codex model versions. It’s fast.

(15:37) It doesn’t consume tokens at an alarming rate and is very efficient. I found the same thing in this experiment: the site, structure, business plan, business model, and board seemed very similar between Medium and High.

(15:53) If you want a little more quality, you can use High and spend a few more tokens. But for most everyday tasks, if you want Astra—though you shouldn’t use it for everything—Medium is more than enough. I hope this helps you understand which effort level might work for you.

(16:12) Based on this test and others I’ve run, I see few cases where Max, Extra High, or Ultra are worth the extra time, effort, or tokens. I ran many other tests and documented findings, additional information, and tips on when to use different levels for certain types of tasks.

(16:34) I’ll make that available at the second link in the description. You can get it for free, give it to Codex or Claude, and ask for an explanation. And as always, if you want to stay ahead in AI, especially in practical applications for work, business, learning, or making money, check out the first link below.

(16:57) I go into topics like this in depth in the Early AI Dopters community. If you enjoyed this and found it useful, I’d appreciate a like on the video—it helps it reach more people—and a comment, if you’d like. See you in the next one.