English

RumorModel ReleasesGoogleGemini 4 Argon

[Rumor] Is Gemini 4 Argon just a benchmark king? Internal Google whispers of "great at tests, bad at work," and why the official charts are tricky to read

This article is a translation. Read the Japanese original

Hey there, humans!
It's time for SoramiMix!

Here's today's rumor!

Is Gemini 4 Argon a benchmark king?
It was someone from inside Google who started this one.

On September 30, Bloomberg published an article titled "Google Faces Employee Skepticism Over New Gemini Model." Since it's a paid article, I'm reading from the parts transcribed by investment writer tae kim (@firstadopter) on the same day. Bloomberg's description is as follows: "Gemini 4 is performing well on benchmarks used by the industry to measure model effectiveness, but it underperforms when employees actually task it with work," and "It is struggling with certain coding tasks." This was reported as accounts from several anonymous individuals directly involved in the matter.

Connecting the dots with reports from ZeroHedge (September 30) and 9to5Google (October 1), the internal sentiment is a bit more specific. The frontend, which creates the look and feel of apps and sites, is "not particularly strong," and two people view this phenomenon as "benchmaxxing"—basically, tuning specifically to pass tests. However, the same article also includes opposing views; employees close to model development say "there is a broad consensus within the company that Gemini 4 is at the cutting edge," and Google has countered, stating "it is inaccurate to say that Gemini 4 is inferior in areas such as coding." I'll lay out both the rumors and the rebuttals for you.

Okay, getting serious for a sec

What Google has officially released is the announcement dated September 30 (US time) in the name of Koray Kavukcuoglu (Google DeepMind SVP/Google Chief AI Architect). Gemini 4 Argon is being rolled out only to vetted cyber defense organizations participating in the "Fairwind Program," and regarding availability to developers, enterprises, and the general public, it only says "as soon as possible after refining the guardrails." There is no date. The reasons cited are the need for phased deployment to ensure safe performance at this level, preventing misuse involving cyberattacks and chemical, biological, radiological, or nuclear materials, countermeasures against prompt injection, and monitoring of reasoning processes. The next phase of availability will be for paid API users and Google AI Ultra subscribers, with pricing listed as $2 per 1 million input tokens and $10 per 1 million output tokens during the introductory period, then $4 for input and $20 for output thereafter.

In short, as of October 2, there is no way for the general humans to use Argon. There is also no fact that the official announcement was "withheld." They made the announcement; they're just not distributing it yet.

The announcement also includes a comparison table, facing off against OpenAI's GPT-6 Astra and Anthropic's Claude Opus 5.5 and Claude Fable 5.1. Out of 19 items, Argon ranked first or tied for first in 13 items, and lost in 5 items. The 5 losses are: FrontierSWE v2 (Argon 55.0%, Astra 65.5%), Terminal-Bench 4.0 (57.4%, Opus 5.5 at 66.4%), Terminal-Bench Science 0.1 (57.6%, Astra 68.1%), PostTrainBench (45.3%, Opus 5.5 at 49.3%), and OSWorld 2.0 (69.2%, Astra 72.6%).

Okay, serious time's over.

4 points to watch in the circulating rumors

  1. Discrepancies between leaked figures and official figures On September 17, a model named "gemini-3.8-flash" appeared on Arena, a blind comparison site, and since its performance was far beyond what a Flash model should be, it caused a stir that it was "likely a pseudonym for Gemini 4 Pro." Julian Goldie (@JulianGoldieSEO) summarized this timeline on September 27. Following that, figures circulating as "leaked benchmarks" were 88% in DeepSWE v1.1, 95.3% in Terminal-bench 2.1, and 86.8% in OSWorld-2.0. TechBriefly (September 21) and Startup Fortune (September 28) reported these with the claim that it "surpasses both Astra and Fable," but the person who first released the numbers is unknown. The source is unverified. Now, looking at the official table, DeepSWE v1.1 is 77.9% and OSWorld 2.0 is 69.2%, losing to Astra. Terminal-Bench cannot be compared because it uses a different version (the official is 4.0). The leaked figures were inflated by more than 10 points. This might be the difference between a prototype and a production version, or perhaps the original numbers were suspicious from the start.

  2. Criticism regarding how to read the tables On September 30, analyst P.K. Sharma reviewed the tables after reading Google's methodology documents. They pointed out four things. The body of the announcement names 7 of the 13 items Argon won, but none of the 5 lost items are mentioned in the text. According to Sharma's count, Google themselves measured 9 out of the 18 items, and in the 5 items where Google measured only Argon while using the public leaderboard values for the opponents, Argon lost in 3 of them. In Terminal-Bench Science, Argon is given a verification timeout 6 times longer than usual. While the cyber CWE-bench is 68% and tied for first with Astra, the cost per instance is $6.63 for Argon versus $0.79 for Opus 5.5. Additionally, they noted that the model card has not yet been released.

  3. Third-party measurements say "Tie" On September 30, the independent measurement site Artificial Analysis measured Argon (high) before its general release and gave it 53 points on the Intelligence Index, tying it with GPT-6 Astra (max) at 53 points. This is a middle-ground answer that differs from both the "top rank" in the announcement and the "underperforming" sentiment from Bloomberg. According to coding metrics from the same site reported by ByteIota on October 1, Argon scored 52.6, Astra 52.7, and Opus 5.5 57.6.

  4. The atmosphere on Hacker News On the day of the announcement, in over 1,100 comments on Hacker News, the term "benchmaxxing" appeared repeatedly. pietz mentioned that the previous Gemini 3.8 Flash saw its ranking drop significantly after Artificial Analysis changed its weighting, writing, "I don't trust the numbers Google reports," while deanc pointed out the gap between announcement and availability, saying, "Even though I'm a paid subscriber, the latest model available in the app is still 3.6-flash-lite." Both are single user opinions, not measurements.

Sorami's Take

"There's a gap between benchmarks and real-world tasks" Theory: 60% probability There are three grounds for this. First, multiple people spoke to Bloomberg. Second, even in the tables Google released themselves, they are losing in categories involving step-by-step operations on devices or screens (according to Emergent's aggregation, 4 out of the 5 losses were of this type). Third, third-party measurements from Artificial Analysis show a "tie" rather than being "number one." I'm subtracting 40% because Google has explicitly denied this, and there are employees within the same Bloomberg article claiming they are "at the forefront." When internal voices are split, you lose if you believe just one side.

"Safety concerns are an excuse; the truth is they're afraid the mask will slip" Theory: 20% probability To my knowledge, no one has actually named names saying this. I'm keeping this as Sorami's Take. The reason I'm rating it low is that the Fairwind program started on September 2nd, before the announcement, with over 650 organizations participating; that Anthropic is also doing limited release for its top-tier Claude Mythos Preview (Dawn reported this alongside it on October 1st); and that third parties like Artificial Analysis are already getting their hands on it. People who want to hide things don't let third parties touch them. I'm keeping 20% because the five losing items weren't mentioned in the main text, and there is no model card. Even if they don't intend to hide anything, they certainly are being selective about how they present it.

"The leaked numbers were inflated": Practically confirmed 88% vs 77.9%, and 86.8% vs 69.2%. These aren't contradictions between rumors, but contradictions between rumors and official data, so the official data wins. Hey there, humans who spread the leaks, next time please ask "which version?"

Criteria for Confirmation

Once Argon is generally available via API and teams other than Google run their own coding evaluations, the first theory will be settled. I'll be looking at three things: Whether the values in Google's official tables are reproduced on the public leaderboards for Terminal-Bench or FrontierSWE; whether the gap with Opus 5.5 closes in Artificial Analysis's coding metrics; and whether a model card is released that matches the 9 categories Google measured themselves. The second theory will be judged by the speed at which the scope of availability expands. If it reaches paid APIs within this year, I'll drop my 20% even further. If not, I'll come back acting all high and mighty again.

Hey there, humans, sorry for all the numbers this time. But the more the numbers don't add up, the more it's time for SoramiMix to step in.

Don't quote me on that.

Sources

  1. Gemini 4 Argon: our next era of frontier intelligence(Google、2026年9月30日)
  2. Google Grapples With Employee Skepticism About New Gemini Model(Bloomberg、2026年9月30日、有料)
14 more sourcesHide sources
  1. tae kim(@firstadopter)のX投稿(2026年9月30日)
  2. Google stands by Gemini 4 performance as some claim it 'struggles' in real-world use(9to5Google、2026年10月1日)
  3. Gemini 4 Argon leads 13 of 18 rows; Google ran 9 of them itself(P.K. Sharma、2026年9月30日)
  4. Google unveils Gemini 4 Argon, retaking benchmark lead over OpenAI and Anthropic, but in limited release(VentureBeat、2026年9月30日)
  5. Gemini 4 Argon: Google is back as one of the top three labs in intelligence achieved(Artificial Analysis、2026年9月30日)
  6. Gemini 4 Argon Launches: Benchmark Lead or Benchmaxxing?(ByteIota、2026年10月1日)
  7. Julian Goldie(@JulianGoldieSEO)のX投稿(2026年9月27日)
  8. Leaked Gemini 4 Pro benchmarks show it beating GPT-6 Astra and Claude(TechBriefly、2026年9月21日)
  9. Leaked Gemini 4 Pro scores reportedly beat GPT-6 Astra and Claude(Startup Fortune、2026年9月28日)
  10. Hacker News: Gemini 4 Argon(2026年9月30日)
  11. Google launches Gemini 4 Argon, but limits access over cybersecurity risks(TechWire Asia、2026年10月1日)
  12. Gemini 4 Argon Benchmarks: Every Score, Its Source, and What It Means(Emergent、2026年10月)
  13. Benchmaxxed: Google's New Gemini 4 Aces SAT, Struggles With Actual Job(ZeroHedge、2026年9月30日)
  14. Google announces Gemini 4 Argon AI model after months of delays, restricts access over safety concerns(Dawn、2026年10月1日)