Gemini 4 Argon spotted in A/B testing on Antigravity

Gemini 4 Argon appears to be quietly reaching some users through A/B tests, and the early outputs shared on X look like a clear step up from Gemini 3.8 Flash. Whether it lives up to Google's benchmark claims is another matter.

· 4 min read
Gemini

Google hasn't opened Gemini 4 Argon to the public yet, but some users seem to be getting an early taste anyway. Over the past couple of days, a handful of people on X have shared outputs they believe are from Argon, served to them through A/B tests and quiet routing behind other model names.

One of the first was @leo114119, who posted a one-shot render of a classical Chinese palace complex.

The account was a Jio-bundled Gemini Pro plan, and the result looks properly detailed, not the flat, generic output Gemini models were often criticized for. In a follow-up post, the same user shared their theory about what might be happening. They claimed the main account never got the test, while a sub-account added to a family group did. They noted a similar pattern with Opus 5.5 earlier, and attached a fresh Pelican test animation for good measure.

Meanwhile, another X user @auricxofficial posted that Gemini 3.8 Flash was routing to Argon, and that Gemini 3.1 Pro in Antigravity now points to 3.8 Flash under the model ID "gemini-3.8-flash-exp-a". They also shared a screen recording of a 3D shooter game that they made with Argon.

That said, @leo114119 shared a more detailed comparison between Gemini 4 Argon and Gemini 3.8 Flash, where the two models built three UI screens for a weather app in HTML for iPhones. Argon's version looks noticeably more polished.

In case you've been following our coverage, you might recall that we covered early Gemini 4 Pro outputs last month, when leakers described an internal "argon" checkpoint with much better frontend taste. Honestly, we don't know whether Google has made any major improvements since then.

Business Insider has since reported that Barium-B is the checkpoint picked for the public Argon release, while employees are also trying a separate checkpoint called Carbon that reportedly feels close to Opus 5.5 for coding, as we noted when Argon hints emerged alongside the Carbon checkpoint. So what people are seeing on X could be any of these.

If you've been under a rock, Google initially announced Gemini 4 Argon on September 30. Since then, the company has been restricting access to a select group of users. The company claims a state-of-the-art 77.9% on DeepSWE, a test of long software engineering tasks, along with first place on the Vals Index for finance, legal, and tax work and on Zapier's AutomationBench. Its own comparison table also shows Argon trailing Claude Opus 5.5 and GPT-6 Astra on some coding and terminal tests.

Gemini

These are Google's numbers, so take them with a grain of salt. We'll only know how good Argon really is once people can use it every day.

While the test results are impressive, the company is also making big behind-the-scenes changes. We've seen Antigravity getting ready for Argon with Ultra upgrade prompts, and newer builds now include Argon-only context options of 256K, 512K, and 900K. Google is also working on an Ultra mode for AI Studio Build, though any link to Argon is unconfirmed. Google says paid API customers and AI Ultra subscribers come next, with no date given.

If Google's benchmarks hold up, Argon should compete with Anthropic's Fable 5.1 and OpenAI's GPT-6 Astra. Argon's A/B outputs look promising, but a few screenshots won't settle that fight. For now, it's anyone's guess.