ServicesSolutionsWorkProcessInsightsCompanyLife at SeikoContactFree AI audit

Home / Blog

GPT-6.1 Sol: OpenAI's $2-Per-Million-Token Model Changes the AI Automation Cost Equation

OpenAI's GPT-6.1 Sol offers near-flagship AI at one-fifth the token price. What cost-per-task economics mean for your automation. Free AI audit.

Rows of server racks with green status lights in a data center, the infrastructure where AI models run

Image: Pexels

At its DevDay event on September 29, OpenAI introduced GPT-6.1 Sol and positioned it as a lower-cost model alongside its higher-capability offering. OpenAI states Sol "nearly matches GPT‑6 Astra's intelligence on agentic coding, computer use, and professional work at one-fifth of Astra's standard input and output token prices." (9to5Mac)

The release followed reported changes around the Astra model line: the planned GPT‑6.1 Astra upgrade was cancelled on the eve of the event after it failed internal safety testing, according to reports. (MadRobot, Android Headlines) For businesses running AI agents against real workloads, this release could materially change automation cost calculations for some workloads.

What OpenAI actually shipped

The headline numbers are straightforward. In the API, GPT‑6.1 Sol (`gpt-6.1-sol`) costs $2 per million input tokens and $10 per million output tokens — a fifth of Astra's reported $10/$50 rates. A major pricing difference is cached input: $0.10 per million tokens, which is 95% below Sol's own standard input-token price and half of what the week-old GPT‑6 Sol charged for cache reads. (Wccftech)

Sol is available through the API and, in ChatGPT, to Plus, Pro, Business, Enterprise, and Edu users inside ChatGPT Work and Codex — but not yet in regular Chat, per DevDay coverage. (MadRobot)

Why does the cache price matter more than the sticker price? Agent workloads can create higher token usage than chat-style workloads because they often involve repeated tool calls, resubmitted context, and intermediate steps. An agent resends its system prompt, tool definitions, and conversation history on every turn of the loop; in a long task, that repeated context can dominate the token bill. At $0.10 per million cached tokens, the cost of repeated context reuse can decrease substantially — and context reuse is exactly what agent sessions do constantly. As reported from the DevDay keynote, Sam Altman framed it as giving developers "more room to build and run capable agents that reuse context across requests." (Wccftech)

Reported benchmark and cost comparisons

These are OpenAI's own benchmark claims, reported by the tech press — treat them as the vendor's best case.

On DeepSWE v1.1, a benchmark for complex software engineering in real codebases, OpenAI reports Sol matching Astra at roughly one-fifth of the cost and beating its predecessor GPT‑6 Sol by 6.4 percentage points. On OSWorld 2.0, which tests agents operating a real computer, Sol lands within 2.1 percentage points of Astra at roughly one-seventh the cost per task. (Android Headlines)

On document-heavy professional work, OpenAI reports Sol scoring higher than Anthropic's Opus 5.5 on GDP.pdf — answering questions from complex PDFs with tables, charts, and fine print — at less than half the cost per task. On AutomationBench, which measures whether agents complete multi-step business workflows correctly, the company reports Sol finishing 2.2 points above Opus 5.5 at medium reasoning effort for roughly a third of the cost. And on Terminal-Bench Science, scientific workflows cost $5.47 per task on Sol against $23.21 for Opus 5.5 and $23.80 for Astra, though Astra still holds the highest score of the group at 68.1%. (Wccftech)

Factuality improved as well: OpenAI reports the share of responses containing a factual error on difficult prompts falling from 11.4% to 7.7% at low reasoning effort, while staying within 1.9 percentage points of Astra's error rate overall.

An independent comparison came from Artificial Analysis on launch night, as reported by dev.to: their Intelligence Index scores GPT‑6.1 Sol at 52 versus 53 for Astra, at a cost per index task of $0.72 against Astra's $3.26. Opus 5.5 still leads the index at 58, but at $5.98 per task. (dev.to)

Why cost-per-task beats benchmark scores

For many production workloads, model selection involves balancing capability, reliability, latency, and cost — not chasing the top benchmark number. Agentic workloads consume tokens in long loops: tool calls, retries, self-correction, context replays. The cost that shows up on your invoice is cost per completed task, not cost per token and not a benchmark number. A model that is slightly weaker on paper but several times cheaper — and uses cached context at a steep discount — wins on any workload where that gap doesn't change the outcome.

Cached-input pricing also has structural implications for agent architecture. Systems that carry big tool lists and long memories across dozens of turns — previously too expensive to run at scale — become economical. It also raises a question worth asking: has anyone recently checked whether your highest-cost model is doing anything the cheaper tier couldn't?

What this means for your business

If you run agents in production — coding assistants, document processing, customer-support automation, multi-step back-office workflows — this release hands you a practical exercise:

Route, don't replace. Take a representative sample of your current flagship-model traffic and replay it through Sol. Measure task success end to end, not benchmark scores. Move every workload where the gap is invisible to the cheaper model. Keep the flagship for the hard minority and route the volume — use replayed representative workloads to find the candidates.

Test before you trust. A cheaper model is not a drop-in replacement — evaluate on your own data, monitor for quality regression, and set approval thresholds for the swap. Measure latency and failure modes alongside cost: a model that needs two retries to match the flagship's first-pass success rate may not be cheaper in practice.

Price your automation per task, not per token. Build a cost-per-completed-task metric into your agent observability now. It should include retries, tool-loop turns, and cache hits. The moment you can see cost per task by workflow, model-routing decisions stop being debates and become arithmetic. Useful segmentation looks like this: code generation measured by test pass rate, document extraction measured by extraction accuracy, customer workflows measured by completion rate — then route per workload, not per company.

Use the cache deliberately. If your agents resend long system prompts and tool definitions every turn, Sol's $0.10 cached-input rate is where the savings compound. Session design — what you keep in context versus what you re-fetch — is now a cost lever, not just an engineering choice.

Put budget controls behind the router. As agent volume grows, inference spend gets harder to predict — especially when agents coordinate multi-step tasks with minimal human oversight. The same week Sol launched, Cribl launched StreamAI, an AI gateway that routes each prompt to the right model with built-in budget controls — a signal that vendors are packaging cost routing as infrastructure. Whether you buy or build, a routing layer with spending controls can help organizations manage agent costs at scale.

One caveat: all the capability numbers above come from OpenAI's own tests. The third-party index confirms the price-to-performance story broadly, but your workloads are the only benchmark that counts. Verify on your own data before you commit.

Cost per task comparison: GPT-6.1 Sol versus flagship and competing models across agent benchmarks

The broader pattern: intelligence is getting cheaper faster than it is getting smarter

Sol replaces GPT‑6 Sol just a week after that model's release, according to the cited reporting — model releases now behave like software patches rather than generational events.

The takeaway is durable: build your pipelines model-agnostic, measure cost per completed task, and let the workhorse models carry the volume. Pricing and model changes will keep shifting automation economics — the teams that benefit most are the ones whose architecture can switch models in an afternoon.

Want to know what your agent workloads actually cost — and which ones could move to a model like Sol today? Start with our free AI audit: we'll map your current AI spend against task-level outcomes and show you where the routing wins are hiding.

Sources

Want to know what this means for your stack? A free AI audit maps your workflows and shows where automation pays off — in your numbers, not ours.