Google secretly tests Gemini 4 Carbon checkpoint to rival Claude Opus 5.5

Google staff are internally testing a hidden Gemini 4 build called Carbon that rivals Claude Opus 5.5 for coding.

Models · Source: TestingCatalog

What happened

Google is preparing to launch Gemini 4 Argon, but employees are secretly testing a different beast. A Business Insider report reveals internal testing of a checkpoint called Carbon. Early employee feedback says Carbon feels comparable to Anthropic Claude Opus 5.5 for coding tasks. Staff are testing it through an internal platform called Jetski, which is associated with Google Antigravity.

The public release of Argon will reportedly use a different checkpoint called Barium-B. Google announced Argon on September 30 for cybersecurity partners, with paid API customers and Google AI Ultra subscribers expected to follow. Meanwhile, Carbon represents a separate and potentially more capable iteration. It remains unclear if Carbon will eventually become an Argon update or launch as a completely separate Gemini 4 model.

Hints of the upcoming public Argon release are already leaking through the Antigravity platform. The selector shows context options of 256K, 512K, and 900K. The platform also added an agent called Chief of Staff, previously known as Concierge. Gemini Web is also testing low, medium, and high thinking effort settings to prepare for the new models, pointing to a broader focus on autonomous coding agents.

Key facts

Why it matters

Anthropic currently dominates agentic coding with the Claude family. Google knows this. By testing Carbon internally against the Opus 5.5 benchmark, Google is signaling a direct assault on the developer ecosystem. If Carbon matches Opus 5.5, builders will have a serious alternative for autonomous coding agents. The context window pricing structure also shows Google is getting aggressive on inference costs, forcing developers to be precise with their token budgets.

The split between internal and public models creates a fragmented rollout. Founders building on the Google stack need to know they might not be getting the best model on day one. Barium-B is for the public, but Carbon is the real prize. This means your initial tests with Argon might not reflect the actual ceiling for Google coding capabilities. You have to plan your product roadmap around delayed access to the top tier reasoning models.

For builders

Plan for tiered context window costs

Argon introduces variable quota costs based on context size. The 512K option consumes 1.3x quota, while 900K takes 1.8x. Founders who fail to optimize token usage will pay significantly higher API bills.

Prepare for autonomous coding agents

Google is heavily focusing on agentic capabilities with features like Chief of Staff. Builders should architect their apps to leverage these native autonomous agents. Startups building thin agent wrappers will lose their competitive edge.

Wait for the Carbon checkpoint

Do not judge Gemini 4 solely on the initial Barium-B release. The Carbon checkpoint is where the Opus 5.5 level coding performance lives. Enterprise customers paying for early access lose out if they commit before Carbon drops.

My take

I think Google is playing games by holding back its best model. Launching Barium-B while employees hoard Carbon shows they are terrified of shipping a coding model that falls short of Anthropic. They need to ship the real thing or do not ship at all.

Original reporting: TestingCatalog. This is my rewrite and opinion.

More AI news for builders