Lilith.
⌕
Editorial illustration: GPT-6.1 Sol gets close to Astra at one-fifth of the token price
Lilith illustration · editorial remix

OpenAI has released GPT-6.1 Sol at $2 per million input tokens, $0.10 per million cached input tokens and $10 per million output tokens. Standard input and output cost one-fifth of GPT-6 Astra's rates, while independent testing shows a much smaller performance gap.

An independent index leaves less than one point between Sol and Astra

Artificial Analysis measured 51.8 for GPT-6.1 Sol and 52.7 for GPT-6 Astra at maximum reasoning effort. A task in its Intelligence Index cost $0.72 with Sol and $3.26 with Astra. Those are costs for a particular evaluation mix, not universal prices for completing a task.

The model replaces GPT-6 Sol only 7 days after its launch. Standard input and output rates remain $2 and $10 per million tokens, but cache reads fell from $0.20 to $0.10. For agents that repeatedly read a long context at every step, that line item may matter more than the new version number.

A cheaper near-flagship changes the economics of agent routing

Teams have often sent the hardest tasks to Astra and routed the rest to a weaker model. Sol 6.1 narrows the territory where Astra's fivefold token rates pay for themselves. The practical result may be simpler routing: the cheaper model handles most coding, document work and tool use, while Astra remains for tasks where internal evals show a measurable advantage.

The rate card also rewards repeated access to the same context. That matters in long agent runs because the bill is built not only from the final answer, but from every return to code, documents and tool history.

One-fifth of the token price does not guarantee one-fifth of the bill

Models use different numbers of tokens and tool calls to finish the same work. Artificial Analysis measured a lower cost per task, but its mix may not resemble a company's workflow. Astra also still leads the index, and fewer failed attempts on a difficult task can outweigh a higher rate.

Prompts above 272,000 input tokens enter a higher tier for the entire request: input and cache rates double and output rises by 1.5 times. Teams using million-token contexts therefore need to model their actual traffic rather than multiply a marketing ratio by request count.

The bill for a completed task in internal evals will decide

The useful test runs both models on the same repositories, documents and tools. It should measure success rate, retries, latency, human interventions and total cost per completed task. A price per million tokens captures only part of that cost.

If Sol 6.1 keeps its near-Astra performance outside public benchmarks, Astra will move from default choice to specialist. A second independent evaluation and sustained production use will show whether this is a durable cost shift or an advantage on a few test suites.

Lilith's verdict

Sol 6.1 pushes Astra off the main road and into the heavy-load lane. The winner will be decided by the receipt after the same job is finished, not the price sign at the pump.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗ ↗