Back to Technology briefing Technology

Google unveils Gemini 4 Argon with 1M-token output; staff split on coding

Google announced Gemini 4 Argon with a 1 million-token output limit and strong scores on several of its cited benchmarks, rolling it out first to trusted cyber defenders. Bloomberg reporting, reprinted by the Northwest Arkansas Democrat-Gazette, says some employees doubt its real-world coding; Google said that claim is inaccurate.

By US Brief desk · Updated 2026-10-04T05:38:00-07:00

AI-assisted · US Brief desk · Sources listed below

Exterior of Google's Googleplex headquarters buildings in Mountain View, California

What happened

Google announced Gemini 4 Argon, a new frontier model it says is aimed at long software, legal, finance, and cyber-defense jobs. Access starts with trusted cyber defenders in Google's Fairwind Program while the company expands safety testing.

What Google claims: The company raised the output limit to 1 million tokens, up from 64,000 on prior Gemini models, so the model can keep working on long tasks in one pass. It cited a state-of-the-art 77.9% on DeepSWE v1.1 for long-horizon software engineering and said Argon leads the Vals Index across finance, coding, legal, and tax work. Introductory API pricing is listed at $2 per million input tokens and $10 per million output tokens, rising later to $4 and $20.

Limited rollout: Google said it is in the U.S. government's voluntary pre-release access process and will gather feedback before a broader developer and consumer launch. InfoWorld noted the limited start follows the delay and cancellation of a planned Gemini 3.5 Pro release.

Why it matters

Gemini underpins Search, Gmail, Maps, and Chrome for more than a billion users each, Bloomberg noted. A strong or weak coding model affects Google's fight with OpenAI and Anthropic for developers and businesses.

What’s next

Broader API and Google AI Ultra availability after Fairwind feedback, Google said. Watch independent coding evals once access widens.

More context

Internal doubts: Bloomberg reported, in a piece carried by the Northwest Arkansas Democrat-Gazette, that some people with access say Argon does less well on messy coding and front-end design than its benchmark scores suggest. Google told Bloomberg it would be inaccurate to say the model is underperforming on coding and pointed back to DeepMind chief Koray Kavukcuoglu's positive remarks. Bloomberg also described a split: some staff think rivals are improving faster; others say Argon has caught up.

Outside read: InfoWorld reported Argon at 68.9% on the Vals Index, ahead of Claude Opus 5.5 at 67.0%, but trailing that rival on PostTrainBench (45.3% vs. 49.3%). Analyst Pareekh Jain told InfoWorld everyday coding looks more average and that enterprises should test on their own workloads.

4 listed sources

References listed by US Brief; a source count is not a verification score.

Editorial sourcing notes

Primary company claims: Google DeepMind blog (Koray Kavukcuoglu), Sept 30, 2026 — Argon, Fairwind cyber defenders, 1M output tokens (from 64K), DeepSWE v1.1 77.9%, Vals Index lead, AutomationBench 51.3%, LVBench 91.7%, CWE-bench v1 68% tie, intro $2/$10 per million tokens then $4/$20. Balance: Bloomberg (Julia Love, Davey Alba) via NWA Democrat-Gazette Oct 4 — anonymous sources on coding/front-end skepticism, abandoned 3.5 Pro, Google denial pointing to Kavukcuoglu; spectrum of internal views. InfoWorld (Nidhi Singal), Oct 1 — Vals 68.9% vs Opus 5.5 67.0%; PostTrainBench Argon 45.3% vs Opus 49.3%; analyst Pareekh Jain on mixed everyday coding. Do not invent rival model names beyond what sources print (Fable/Astra appear in Bloomberg reprint). Re-check ~5:20 AM PT Oct 4.

Like US Brief? You can support it with a tip.