Gemini 4 Argon reached 965 points and 657 comments on Hacker News while an ordinary developer still could not send it a single API request. Google announced a price, a one-million-token output limit and a thick table of scores on September 30. The model is initially going to a much narrower group: vetted cybersecurity defenders. That gap between announcement and access is the useful part of this launch.
Google says Argon is rolling out first through its Fairwind Program, while a wider release will begin later with paid API customers and Google AI Ultra subscribers. It gives no date for that second stage beyond "as soon as possible." So this is a model debut without the usual developer ritual. There is no public model ID to paste into an SDK, no latency to measure and no production behavior to compare with the launch charts.
A price without a public endpoint
Google has already set introductory API rates at $2 per million input tokens and $10 per million output tokens. Cached input will cost 95 percent less than fresh input. After the introductory period, those rates double to $4 and $20, although the announcement does not say when the lower rates end. A team can sketch a budget from those numbers, but it cannot test how quickly Argon consumes that budget or how often a task needs another attempt.
The most unusual specification is the output allowance. Google raised it from 64,000 tokens on the prior model to one million tokens for Argon. Google specifically describes this as an output limit, distinct from the input context window. At the introductory rate, a full one-million-token response would carry $10 in output charges before counting the prompt, tool calls or retries. Google presents that headroom as a way to sustain one long reasoning trajectory, including jobs that generate hundreds of thousands of tokens.
That design moves a familiar engineering problem downstream. If an agent can run for that long, applications need sensible cancellation, checkpoints and spending limits. They also need to decide what should survive when a tool fails near the end of a run. The million-token figure tells developers how far a trajectory may go. Until the API arrives, it says nothing about streaming stability, rate limits or whether the useful part of a job survives a timeout.
Publishing the future price before the endpoint also changes what developers can compare. Argon's introductory input price looks approachable in isolation, but cost per token does not reveal cost per finished task. Google has supplied token rates, without public measurements for time to completion, retry rates or tool-call overhead. Those missing observations matter more for a long-running coding agent than another percentage point on a static test.
Why cybersecurity gets the first copies
Google's explanation is straightforward: the same model that finds a serious bug can help exploit one. The company says Argon can inspect code, validate a vulnerability and produce a patch. Trusted defenders and Google's internal teams will receive a version without cyber guardrails, while the broader model is being prepared to refuse harmful cyber and chemical, biological, radiological and nuclear requests. The public Argon page reports a 68 percent score on CWE-bench v1, tied with GPT-6 Astra and one point above Claude Opus 5.5 in Google's table.
Fairwind is a controlled distribution channel with organization-level admission rules. Google DeepMind says the program works with more than 650 partners and gives priority to governments, critical infrastructure operators and core technology platforms. Participating organizations must use individual accounts, phishing-resistant multifactor authentication and access controls. They must track employee use, cannot resell or redistribute access, and face an organizational background check.
The permitted work includes authorized threat simulation, reverse engineering and malware analysis for defense or academic research. Creating malware is outside the rules. Academic labs may apply for defensive benchmarking, but Google says student users should use CodeMender with public models instead. Fairwind partners can run Argon by itself or through CodeMender, Google's agent for automating software fixes.
This rollout makes access control one of Argon's product features. A public API normally relies on policy filters at request time. Fairwind adds eligibility checks, organizational controls and audit obligations before the request exists. Whether that structure keeps its shape when Argon reaches paid API accounts will determine how much of the cyber model regular security teams can use. Google has not described the exact capability difference between the Fairwind build and the later public one.
The benchmark table refuses a simple verdict
Argon leads several of the tests Google selected. It scores 77.9 percent on DeepSWE v1.1, ahead of the comparison models in the published table, and 51.3 percent on AutomationBench. Its 84.2 percent GraphWalks result for inputs between 256,000 and one million tokens also tops the listed rivals. Google's full results page puts those wins beside less favorable outcomes rather than showing a clean sweep.
On FrontierSWE v2, Argon records 55.0 percent, below GPT-6 Astra at 65.5 percent and Opus 5.5 at 62.3 percent. Terminal-bench 4.0 gives Argon 57.4 percent, while every named rival in the table scores slightly or substantially higher. Opus also leads Argon on PostTrainBench, and Astra leads it on the listed science terminal test. For developers, the table describes a model with a particular profile: strong on some long jobs and enterprise tasks, less convincing in other coding harnesses.
The comparison has methodological seams. Google's evaluation document says all Argon results use the Gemini API at the highest thinking setting unless noted. Some Argon scores were computed by Google, while competing numbers often came from providers or public leaderboards. The OSWorld score is the best of three runs, and the video benchmark used different frame counts because the APIs had different limits. These are disclosed choices, but they are not one controlled race on identical infrastructure.
The cyber evidence needs the same care. CWE-bench is public, but Google's real-world vulnerability set and Wiz's black-box penetration test are internal. Google reports that Argon scored 85.8 percent on its vulnerability-discovery set and 70.9 percent on Wiz's test. The methodology describes what those sets measure, yet outside teams cannot reproduce the private results from the published material.
Google's own code is the stronger sales pitch
The launch includes internal work that is more concrete than a general claim of better reasoning. Google says Argon agents found memory changes that have freed more than 300 TiB across its data centers, with projected total savings between 500 TiB and 1 PiB. It also used agents on C and C++ to Rust migrations ranging from smaller libraries to more than 800,000 lines in the Fuchsia Zircon kernel. Google says these migrations still go through automated checks, emulation and human review.
One case involved libgav1, Google's open source AV1 decoder. According to the company, Argon replaced 32,000 lines of SIMD code in an existing Rust port, repeatedly checked compiler output and produced a version that ran 2.7 times faster than that Rust port with identical video output. This is vendor-reported work. It shows the intended unit of work as a sustained migration with profiling and review, well beyond a single generated function.
Those internal migration examples also explain the oversized output ceiling better than the benchmark table does. A code migration across hundreds of thousands of lines creates long tool traces, compiler feedback and patches. Giving a model room to keep one trajectory alive could reduce the need to compress that state. The unanswered question is whether outside developers get the same harness, tool access and recovery machinery that Google used internally. A bare model endpoint would reproduce only one part of the system.
What to watch before writing integration code
There is no useful Argon code snippet yet because Google has not published the public API identifier or availability documentation. Inventing one would hide the central fact of the release. The practical signals will be an entry in the Gemini API or Vertex AI model catalog, documented regional access, rate limits, safety behavior and a firm date for the pricing change. Google currently promises only that paid API users and AI Ultra subscribers come first in the broader rollout.
Independent evaluations will matter after access opens. Developers can then test whether DeepSWE strength carries into their repositories, whether the million-token output limit produces recoverable work, and how the guarded public model handles legitimate security tasks. The 965-point Hacker News launch shows that people want to inspect Argon. The next meaningful number is how many can run the same job twice and get a result they would merge.
Google has already told developers what Argon may cost and how long it may work. The next evidence that matters is the first public endpoint, along with its restrictions. The release becomes testable when developers can see what Argon finishes, what its guardrails stop and what remains after a long run fails.