Google announced Gemini 4 Argon on September 30, introducing its new Gemini 4 frontier model — which Google describes as its most powerful model yet. Argon is built to "sustain deep reasoning across complex, long-horizon workflows," with three stated targets: real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. (Google, TechCrunch)
Two details matter for enterprise buyers. First, Argon raises its maximum output from 64K to 1M tokens — the largest output ceiling Google has shipped. Second, Google has introduced it at $2 per million input tokens and $10 per million output tokens. That introductory rate will later rise to $4/$20. (Google, DX Today)
GPT-6.1 Sol also launched at $2/$10 per million input/output tokens, giving buyers two newly announced models at the same initial price point. The launch-week pricing looks aggressive — but the fine print is where the business analysis starts. (Geeky Gadgets)
What Google actually shipped — and who gets it first
Argon is beginning with a restricted rollout. Rather than a public launch, Google is releasing the model to a set of trusted cyber defenders through its Fairwind Program. Google says broader access will start with paid API customers and Google AI Ultra subscribers before expanding further to developers, enterprises, and consumers. There is no firm public release date, and Google says it is engaged in the US government's voluntary pre-release model access process while it iterates on safeguards. (Google)
The sequencing is deliberate. Argon can autonomously find, validate, and patch software vulnerabilities, and Google describes these capabilities as dual-use — which is why it is restricting early access and strengthening safeguards against cyber misuse. For the initial cohort, trusted defenders and Google's own teams receive a version released without cyber guardrails so they can use its full defensive capability. Wiz is already using Argon through its Scan for Good program, and Google says the model uncovered a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals — a flaw previous frontier models had missed. (Google)
On benchmarks, Google reports state of the art on DeepSWE v1.1 at 77.9% (long-horizon software engineering), the top spot on the Vals Index across finance, coding, legal, and tax work, number one on AutomationBench at 51.3% (Zapier's end-to-end business workflow benchmark), and 91.7% on LVBench for long video understanding. On the vulnerability-remediation benchmark CWE-bench v1, it ties for first at 68%. These are vendor-reported results. They are useful for comparison, but enterprises should validate performance on their own workloads before treating them as purchasing evidence.

Why the 1M-token output limit matters
Previous Gemini models capped output at 64,000 tokens. Argon expands that to one million tokens in a single response — which Google calls industry-leading. Output length rarely makes headlines, but for some long-running workflows, the output-limit increase may be more operationally significant than another benchmark gain.
Some high-value enterprise workloads require long, multi-step execution: the software-engineering and professional knowledge-work tasks Google highlights in its announcement. Long workflows are often orchestrated across multiple model calls, which requires the application to manage state and intermediate outputs. A higher output limit allows substantially more generation to remain within a single model trajectory. (Google)
One necessary qualification: a 1M-token maximum is a capacity limit, not proof that million-token generations will be reliable, economical, low-latency, or desirable on real workloads. Enterprises still need to test those properties themselves.
Google's own internal examples make the case concrete. Argon agents analyzed fleet-wide profiling telemetry and applied memory optimizations across Google's data centers — 300 TiB freed so far, with an estimated 500 TiB to 1 PiB in total savings. Separate Argon agents are migrating C/C++ codebases to Rust at scale, up to the Fuchsia Zircon kernel at 800K+ lines. For the libgav1 video decoder, agents replaced 32K lines of SIMD code with safe Rust the compiler could auto-vectorize — a memory-safe decoder running 2.7x faster than the existing Rust port, with identical video output. And in quantum research, Argon beat a published baseline by 40% on algorithmic optimization in minutes. (Google)
Whether those numbers hold outside Google's walls is unverified. But the pattern is the thing: long, autonomous, multi-step work with more of the reasoning kept inside a single trajectory.
Argon launches at $2/$10 — before moving to $4/$20
Argon's introductory API price is $2/$10 per million input/output tokens, with cached input priced 95% below standard input. (Google)
Here is the comparison with the post-introductory rate included, because the steady-state price is what long-term budgets should be built on:
- Argon (introductory): $2/$10 per million input/output tokens
- Argon (post-introductory): $4/$20, per reporting on the announcement
- GPT-6.1 Sol: $2/$10
- Anthropic cheaper option: $4/$20
- Anthropic high-end option: $10/$50
(Google, DX Today, Geeky Gadgets, CoinCentral)
Once the introductory period ends, Argon's steady-state price matches Anthropic's cheaper option — not the dramatic undercut the launch-week headlines suggest. Argon and GPT-6.1 Sol now share the same launch price, increasing price competition at this part of the market. That is worth watching, but it is not the same as permanent price convergence.
There are two caveats. First, these are introductory prices on a model with a staged rollout — Argon is not broadly available yet. Second, token prices are only part of the economics: what matters for a buyer is the total cost per successfully completed workflow, not the sticker price per token.
What this means for your business
Price is only one differentiator. Buyers should also compare workflow fit, availability, reliability, and operational controls. When two vendors launch at the same price, the decision shifts from "which model can we afford" to "which capability actually removes a bottleneck in our pipeline." Audit your automation workloads for measurable bottlenecks — the tasks where long-horizon reasoning or long outputs are the constraint.
Design longer trajectories with guardrails, not just capacity. A higher output limit creates the option to keep more of a long task within one trajectory, but it concentrates risk in a single run. If one agent performs more steps before returning control, you need approval boundaries, checkpointing, resumability, execution logs, and rollback/recovery strategies. Present these as design requirements from the start — they are not features the model provides for you.
Evaluate the defensive tooling, not the headlines. Google reports that Argon identified a critical vulnerability that previous frontier models had missed, and Wiz is already using it to protect public infrastructure. Security teams should evaluate whether stronger AI-assisted vulnerability discovery and remediation tools warrant changes to their existing processes. That is the defensible takeaway — not that anyone's current patch cycle is suddenly obsolete.
Don't switch on launch price alone. A practical decision rule: do not migrate because of an introductory price or a benchmark table. Test only if the workload has a measurable bottleneck the new capability could remove. Keep model-specific dependencies behind an abstraction layer, and build against the capability, not the brand.
Measure caching on repeatable workloads. Google's 95% cached-input discount is directly supported, but it is most relevant to workflows that repeatedly reuse eligible context. At Argon's introductory rate, the discount would reduce eligible cached-input pricing from $2 to $0.10 per million tokens. Before projecting savings, validate cache-hit rates and eligible context on your actual traffic. (Google)
For every production workflow, measure input tokens per completed task, output tokens per completed task, cache-hit rate, retries, failed runs, and human-review time. Lower token prices only help if they translate into a lower cost per successfully completed task.
For buyers, the practical question is not which launch generated the strongest benchmark table, but whether a new model can lower the measured cost or complexity of a specific production workflow.
Want to know which of your workflows are actually ready for long-horizon AI agents — and what each one costs per completed task? Get a free AI audit and we'll map it out with you.
Sources
- Introducing Gemini 4 Argonblog.google
- Google releases Gemini 4 Argon, called its most powerful model yettechcrunch.com
- Google's New AI Model Just Undercut OpenAI and Anthropic on Pricecoincentral.com
- OpenAI DevDay 2026 Reveals GPT-6.1 Sol and Dots AIgeeky-gadgets.com
- DX Today AI — October 1, 2026: Google Unveils Gemini 4 Argonnews.dxtoday.com
Want to know what this means for your stack? A free AI audit maps your workflows and shows where automation pays off — in your numbers, not ours.
