Gemini 4 Argon Targets Million-Token Agent Work—but Starts Behind a Safety Gate

October 2, 2026

A prismatic AI core operates inside a governed containment chamber while a continuous workflow extends through coding, documents, analysis, and security stations.
Gemini 4 Argon is built for unusually long agent trajectories, but Google is beginning with controlled access while it tests frontier safeguards.

Google has announced Gemini 4 Argon, a frontier model designed for complex, long-running work across software engineering, finance, legal research, multimodal analysis, and defensive cybersecurity.

This is not yet a general release. Argon is initially rolling out to trusted cyber defenders through Google DeepMind’s Fairwind Program. Google says access will expand to developers, enterprises, and consumers after further safeguard testing, beginning with paid API customers and Google AI Ultra subscribers.

Built for unusually long workflows

Google says Argon’s output limit will reach one million tokens—up from 64,000—giving the model room to sustain much longer reasoning and execution trajectories.

The company is already using Argon internally. Google reports that teams have deployed Argon agents to identify data-center memory optimizations, assist with large C/C++-to-Rust migrations, and improve quantum-computing algorithms.

Its planned introductory pricing is $2 per million input tokens and $10 per million output tokens, with cached input priced at a 95% discount. After the introductory period, pricing is scheduled to rise to $4 per million input tokens and $20 per million output tokens.

Strong results—with important caveats

Google reports that Gemini 4 Argon scored 77.9% on DeepSWE v1.1 for long-horizon software engineering, 51.3% on AutomationBench for end-to-end business execution, 91.7% on LVBench for long-video understanding, and 68% on CWE-bench v1, tying for first in vulnerability remediation. These are vendor-reported results and should be validated against each product’s own workloads.

Independent testing from Vals placed Argon first among 41 models on its September 30 Vals Index, scoring 68.90% at an estimated $15.68 per test. It led Vals’ finance evaluation and performed strongly in coding, tax, legal, and cybersecurity tasks, although it was weaker on some computer-use evaluations.

Vals lists a 262,144-token maximum output for the configuration it tested, rather than Google’s announced one-million-token limit. Builders should therefore verify the production limits once broader API access begins.

Why it matters

The important shift is the length and autonomy of the work Google is targeting. A model capable of generating and evaluating hundreds of thousands of tokens in one trajectory could handle migrations, investigations, research projects, and optimization loops that currently require repeated orchestration and context handoffs.

The restricted rollout is equally significant. Google is treating frontier cybersecurity capability, prompt-injection resistance, action monitoring, and sandbox hardening as release requirements—not features to add after deployment. That complements the broader move toward AI security systems that combine model investigation with controlled validation.

Gemini 4 Argon could become a serious platform for long-horizon agents. For now, however, it remains an announced frontier model behind a controlled-access safety gate rather than a generally available developer product.

Relevant links

← Back to stories