Franklin AI News Brief

Google announces Gemini 4 Argon, with cyber-defense access first

Key Takeaways

  • Google has announced Gemini 4 Argon with a one-million-token output limit and introductory API pricing.
  • Selected cyber defenders get access first; the broader rollout has no announced date.
  • Google announced Gemini 4 Argon on September 30, opening its new frontier model to selected cybersecurity partners before a broader launch.
  • The announcement sets out a model intended for sustained coding and professional work, alongside a staged release that leaves most developers waiting for access.
  • In Google's launch announcement, the company says paid API customers and Google AI Ultra subscribers will be first in line for the wider rollout.

Google announced Gemini 4 Argon on September 30, opening its new frontier model to selected cybersecurity partners before a broader launch. The announcement sets out a model intended for sustained coding and professional work, alongside a staged release that leaves most developers waiting for access.

In Google's launch announcement, the company says paid API customers and Google AI Ultra subscribers will be first in line for the wider rollout. Google has not given a date for that stage. Readers looking for a model they can use today should check the official Gemini API model catalog, rather than assume the announcement means unrestricted availability.

Who can use Gemini 4 Argon now

Early access runs through Google's Fairwind Program, which prioritizes vetted defenders working on critical systems, including healthcare and telecommunications. Selected partners can use Argon directly or through CodeMender, Google's code-security agent. Eligibility for Fairwind does not automatically mean every partner receives Argon.

The program puts conditions on access: participating organizations must control which employees can use the model and restrict dual-use work to authorized defensive or research purposes. Partners cannot resell or redistribute access. That matters when evaluating offers from third-party sites claiming immediate Gemini 4 access; a name on a landing page is not proof of an authorized integration.

Google says trusted defenders and its internal teams will receive a version without cyber guardrails. That is a description of controlled defensive access, not permission for unrestricted public use. This continues the specialized rollout approach behind Gemini 3.8 Flash Cyber, while bringing a new model to the same security-focused audience.

A larger output budget and introductory pricing

Argon's announced output limit is one million tokens, increased from 64,000. This is an output allowance, not a newly announced one-million-token input context window. It gives a task more room for generated work; it does not guarantee that a long response is correct or that a lengthy agent run reaches a useful result.

Google lists introductory API prices of $2 per million input tokens and $10 per million output tokens, with a 95% discount on cached input. The announcement lists later rates of $4 and $20 respectively, without specifying when the introductory period ends.

At the introductory output rate, generating the full million-token allowance would cost $10 in output charges alone. Input, tools and repeated attempts would add to the bill. Teams planning lengthy runs should set spending limits and measure the total cost of a completed task, including retries, rather than budget from the price of one short response.

Reading the benchmark claims

Google reports 77.9% on DeepSWE v1.1, an evaluation of long-horizon software engineering, and says Argon leads the Vals Index. Those are specific evaluation results, not evidence that the model will solve the same percentage of a company's development backlog.

The Vals Index methodology combines professional tasks across fields such as finance, legal work and coding, weighting sectors by their economic contribution. That makes its overall score different from a coding-only test. A legal team and a software team should look at the relevant task results before treating the combined ranking as a purchasing decision.

For defensive security, Google reports a 68% result on CWE-bench v1, tied for first. CWE-bench evaluates vulnerability remediation. Fixing a known benchmark vulnerability is a narrower test than finding and safely repairing an unknown flaw in a live system; neither score removes the need to review patches and run regression tests.

Gemini 4 Argon intelligence comparison

The latest Franklin benchmark snapshot records these Artificial Analysis Intelligence Index scores. Higher is better. These are index points, not percentages, and the reasoning configurations differ.

Model and tested configuration Intelligence Index
Claude Opus 5.5 (Adaptive Reasoning, Max Effort, Default Fallback) 57.6
GPT-6 Astra (Max) 52.7
Gemini 4 Argon (High) 52.6
Grok 4.7 (Xhigh) 46.4

Snapshot: September 30, 2026, 21:08 UTC. Source: Artificial Analysis Intelligence Index methodology. This selected comparison is separate from Google's vendor-reported benchmarks above. The 0.1-point gap between Astra and Argon is too small to establish a meaningful practical advantage without uncertainty estimates and workload testing.

The deployment questions that remain

The announcement does not provide a broad-release date or a public API model identifier. Google says it is gathering feedback and refining safeguards before expanding access. Until that changes, general users cannot plan a production migration around a confirmed availability window.

Once access expands, useful tests should include an existing bug backlog, permission boundaries and the amount of human correction each task needs. Systems that can write substantial code also need limits on what they can execute. The containment problem addressed by Nvidia's Open Agent Safety Platform remains relevant regardless of which model supplies the reasoning.

Franklin AI has not independently tested Argon. The launch specifications and performance claims above are attributed to Google and the linked benchmark descriptions, with rollout details checked against Google's published access program.

Our read

Franklin AI Take

Argon deserves attention for demanding coding and security work, but availability is the first constraint. When access expands, judge it by verified fixes, required human intervention and cost per completed task. A larger output allowance can support longer work; it can also make a failed run more expensive. Keep code execution permissioned and review security patches before deployment.