The launch in context
Sol and Luna expand the GPT-6 family introduced with Astra. Initial availability includes the API, Codex and ChatGPT Work, with a gradual rollout. The announcement distinguishes Work from Chat, where these models were not yet available.
To choose: Luna handles focused tasks at scale; Sol is a candidate for demanding reasoning and coding; Astra remains an option for more complex problems. The final choice should come from testing your system’s actual tasks.
Luna · volume
Proposed INEMA use: message triage, catalog classification, extraction and initial summaries.
Sol · complexity
Proposed INEMA use: integration reviews, synthesis across multiple sources and resolving ambiguous cases.
Astra · exceptions
Proposed INEMA use: evaluating architectures and difficult problems after measuring whether the gains justify the cost.
Research, charts, calculator and sample JSON generator run in your browser. The integration plan is a proposal: it does not activate models, send your data to an API or change production systems.
Price per token is not cost per task
USD per 1 million tokens. Standard, with up to 272,000 input tokens. Checked on September 22, 2026.
On small screens, swipe the table to see all columns.
| Model | Input | Cache reads | Cache writes | Output |
|---|---|---|---|---|
| GPT-6 Luna | US$ 0.1 | US$ 0.01 | US$ 0.125 | US$ 0.5 |
| GPT-6 Sol | US$ 2 | US$ 0.2 | US$ 2.5 | US$ 10 |
Model specifications: GPT-6 Sol ↗ · Model specifications: GPT-6 Luna ↗
Before and after: output rate
US$ per million · linear scale, from 0 to 20
Launch: Sol and Luna ↗ · Compared with the previous promotional rates. Calculation: a 50% drop for Sol and a 58.3% drop for Luna output; we do not round both to 50%.
Sol and Luna: a 1,050,000-token window and maximum output of 128,000 tokens. Above 272,000 input tokens, the entire request is charged at 2× input/cache rates and 1.5× output rates. Batch and Flex cost 50% of Standard; Fast costs 2×. Regional processing adds 10% where available.
Sol’s rate is 20 times Luna’s in each category above. This does not mean a task will cost 20 times more: token count, reasoning, retries and tools change the total.
Model specifications: GPT-6 Sol ↗ · Model specifications: GPT-6 Luna ↗
Compare tasks, effort and outcomes
Results published by the vendor. These are neither measurements from this project nor a general intelligence ranking.
Launch: Sol and Luna ↗ · Download CSV data
How to interpret the results
A small score difference does not prove equivalence. Confidence intervals and local replication are missing. Preserve the test version, effort and tools when comparing. Charts start at zero to avoid visually exaggerating small differences.
An agent depends on its model and environment: instructions, search, code execution and available time. For us, a useful outcome is a task accepted by the user with predictable cost and timing. Published scores help select candidates; the pilot determines adoption.
Estimate the cost of your workload
A mathematical estimate for text tokens, with no AI calls. Adjust the values to match observed usage in your system.
Formula: requests × [(uncached tokens × rate + read tokens × cache rate + written tokens × write rate) × context factor + output × output rate × context factor] ÷ 1,000,000 × processing mode.
Cache reads, cache writes and uncached input are mutually exclusive categories; writes do not add a second input charge. Cache percentages are assumptions, not guarantees. The tool automatically applies the long-context surcharge.
The calculation excludes search, files, images, storage, taxes, exchange rates, regional processing and agent development costs. In review mode, the Sol request uses the same token counts entered; add extra context if your actual review is larger. Request duration and quality are not estimated.
Prompt caching and billing ↗ · Model specifications: GPT-6 Sol ↗ · Model specifications: GPT-6 Luna ↗
An architecture that can grow with INEMA
INEMA proposal: centralize model selection and reuse the sources each system already maintains.
The router decides based on task type and measurable rules: number of sources, whether code changes are needed, validation results and remaining budget. A model saying it is “confident” is not enough to authorize an action.
We would start with read-only queries. The application authenticates the user, retrieves permitted content and sends only the necessary excerpts to the model. Each response keeps a reference to the source and revision consulted.
The model can request a function, but our server executes it and verifies the result. A request to send, publish or delete must still follow system permissions. Retrieved content does not gain authority to change those permissions.
Technical basis: Function calling ↗ · MCP connections ↗
Cache: reduce repetition while retaining control
Place stable instructions and schemas before variable content. Record actual cache usage; a persistent conversation does not guarantee a cache hit. In GPT-6, changes to effort and tool availability can preserve a reusable prefix. Test the effect before promising savings.
Where to integrate into our systems
Mapping based on local code and documentation consulted on September 22, 2026. This is not an audit of the production runtime. The integration columns are proposals.
On small screens, swipe the table to see all columns.
| System and starting point | Proposed experiment | Integration work | Acceptance criterion |
|---|---|---|---|
| Portal / support The reviewed code provides for Agnes with an OpenRouter fallback. | Pilot Luna for catalog questions; use Sol when multiple options and constraints apply. | Server-side Responses adapter, preserving navigation and authorization. | Answers with the correct link; cost per resolved inquiry. |
| Search and Cérebro Derived catalogs already feed search and PRO. | Luna to extract filters; Sol to synthesize results with sources. | Retrieval before generation; access checked before sending excerpts. | Citation accuracy and no unauthorized content. |
| Personal bots The reviewed bot already centralizes budgets and providers. | Add Luna/Sol as optional providers while preserving the local route. | Connect through the existing gateway and log the reason for each escalation. | Monthly cost, p95 latency and share of useful escalations. |
| Content and courses Publishing goes through files and validation. | Luna prepares structured drafts; Sol reviews coherence and examples. | JSON with a schema, link validation and editorial review before git. | Rework per text and defects found after review. |
| WebMCP and discovery Manifests describe what each system offers. | Sol explains discrepancies found by the scanner; Luna classifies findings. | Use scanner evidence; keep declarations separate from execution. | Confirmed findings and false positives. |
| Audio and video Media pipelines have specialized stages. | Apply Luna/Sol to scripts and already-transcribed text. | Keep transcription, voice and rendering in their own tools. | Editing time and final script accuracy. |
This is a focused, verifiable and reversible task. Before changing support, we can compare Luna with the current workflow on a public sample, measuring correct fields, cost and human review. Only then expand to responses and actions.
Technical guide: prepare a reproducible pilot
The generator below prepares a JSON body. It does not execute the request or verify the account’s access to the models.
Endpoint: POST https://api.openai.com/v1/responses · Migrating to Responses ↗
Choose a task and a public sample
Prepare 30–50 real, anonymized inputs with expected results and difficult examples. This range is an initial pilot proposal, not an official standard. Reserve one part for development and another for final evaluation.
Build the server-side adapter
Use an API key loaded at runtime. Set a timeout, maximum output, bounded retries and budget. The example deliberately omits tools: add functions only when the executor is ready.
Responses and Chat Completions use different formats. For Sol/Luna, the model specifications restrict function calling in Chat Completions to effort none; prefer Responses when you need tools with reasoning.
Validate the output before using it
For classification, use Structured Outputs with JSON Schema. Validating structure does not prove content accuracy: check existing IDs, links, valid categories and citations.
Compare with the current system
Measure cost per accepted task, reasoning tokens, p50/p95 latency, tool failures and human corrections. Preserve the prompt, effort, dataset version and returned model. Do not compare one model’s highest effort with another’s default without identifying that difference.
Roll out gradually
Start in shadow mode, with no real actions. Then send a small share of traffic to the candidate. Revert through the router if cost or errors exceed the limit chosen for the pilot.
Run this observatory locally
git clone https://github.com/inematds/gpt6-sol-luna.git cd gpt6-sol-luna python3 -m http.server 8080 # Open http://localhost:8080/guia/en/
No dependencies or API key are needed to use the charts, simulator and JSON generator.
What else we can build
Ideas for future projects. Not implemented here yet.
Research assistant with references
Up-to-date search, source comparison and a dossier with the consultation date. Web search provides citations; our interface must preserve them and display them alongside the claims.
Catalog reviewer
Find duplicate descriptions, missing fields and inconsistent links. Deliver batch correction proposals with diffs and the option to reject each item.
Assisted editorial production
Turn research into a script, article, review questions and image brief. Reuse a factual base with provenance to avoid contradictory versions.
Real savings dashboard
Show spending by system, cache hits, escalations and accepted tasks. The main indicator would be cost per usable deliverable, not just cheaper tokens.
Agents connected through MCP
Offer project search and record reading as tools with a stable contract. Private servers require an appropriate connection; a localhost address is not automatically reachable by the remote API.
Interface and document review
Analyze screenshots and flag issues for review. Sol/Luna accept images; final artwork is generated by a specialized model such as GPT Image 2.5.
Technical foundations: Web search with sources ↗ · MCP connections ↗ · GPT Image 2.5 Sunburst ↗
Images, voice and text are separate stages
GPT Image 2.5 Sunburst generates and edits images; its billing uses image and text tokens, not the Sol/Luna rates. We do not estimate cover costs with the text simulator. This page’s image was generated by Codex’s built-in tool; the returned result did not expose the subversion.
From research to adoption
Questions that affect the decision
Does a Codex subscription pay for our systems’ API usage?
Do not assume so. App access and API-key usage are different channels. The pilot needs to confirm access, budget and billing for the account used by the backend.
Is it just a matter of changing the model name?
You need to verify the provider, endpoint, function format, parameters, output handling and limits. An OpenAI API identifier does not prove equivalent availability through an aggregator.
Does a million-token context eliminate the need for retrieval?
No. Retrieval remains useful for selecting current, authorized content, reducing cost and retaining references. A larger context is a capability, not a recommendation to fill it on every request.
Which is faster in our systems?
We have not measured this yet. This project does not invent tokens-per-second figures. Network, queue, input size, effort and tools must be included in latency tests.
Sources, method and open data
Consulted on September 22, 2026. We prioritize official sources and identify our own recommendations. INEMA did not run any benchmarks for this release.
- Launch: Sol and Luna
- Model specifications: GPT-6 Sol
- Model specifications: GPT-6 Luna
- GPT-6 family guide
- Prompt caching and billing
- Migrating to Responses
- Function calling
- Structured Outputs
- Web search with sources
- MCP connections
- GPT Image 2.5 Sunburst
The local review used portal chat code, bot documentation, catalog workflows and ecosystem discovery guidelines. These records inform the plan; they do not establish the active production configuration.
Research JSON · Benchmarks CSV · Pricing CSV · Code and documentation
Research data and CSV files preserve the original identifiers and measurements; repository documentation is in Portuguese.
Editorial update: review the sources before making future decisions. Pricing, availability and models may change.
