Anthropic Ships Sonnet 5.5: 30 Percent Faster and Cheaper

Anthropic just dropped Claude Sonnet 5.5, delivering near-Opus coding performance at a massive discount for everyday agentic tasks.

Models · Source: Hacker News

What happened

Anthropic released Claude Sonnet 5.5 today. It is the second model in the new 5.5 family. The model runs 30 percent faster than its predecessor, Sonnet 5. It also costs up to 30 percent less for most practical work. The base pricing remains identical at two dollars per million input tokens and ten dollars per million output tokens. The actual cost savings come directly from the model using far fewer tokens to complete the exact same tasks.

The performance jump in coding and agentic workflows is massive. Sonnet 5.5 scores 70.6 percent on the Terminal-Bench 4.0 agentic coding evaluation. For context, Sonnet 5 only managed a 10.3 percent score on that exact same test. It also scores 55.5 percent on CursorBench, putting it right behind the much heavier Opus 5.5 model. It even beats the game Pokémon Red using nothing but visual screenshots.

Anthropic is also introducing adjustable effort levels for inference. Builders can scale compute from Low to Max effort depending on the task complexity. At lower settings, the model answers faster and uses fewer tokens. At higher settings, it reasons longer and checks its own work. Claude Haiku 5.5 is also slated to join the model family in the coming weeks for highly cost-sensitive applications.

Key facts

Why it matters

This completely changes the unit economics for agentic workflows. Builders no longer have to pay premium Opus prices for complex, multi-step coding tasks. Sonnet 5.5 batches tool calls much more efficiently and requires significantly fewer shell runs to finish a job. You get faster iteration loops and lower API bills without sacrificing reasoning capability. Startups building autonomous coding agents or complex data extraction pipelines will see immediate margin improvements just by swapping the model endpoint.

The second-order effect is a fundamental shift in how we purchase and manage model compute. Anthropic is pushing adjustable inference where you scale effort up or down per individual task. Lower effort settings on Sonnet 5.5 now beat maximum effort on Sonnet 5 for a fraction of the cost. This will force competitors to offer similar sliding scales for inference compute. Builders will need to start building dynamic routing systems that adjust effort levels based on the specific difficulty of each user prompt.

For builders

Cheaper automated code reviews

CodeRabbit is already moving simple and moderate code reviews to Sonnet 5.5. The new model uses significantly fewer output tokens and stops unnecessary web searches. You save money on API costs while keeping review quality exceptionally high.

Faster customer support agents

Zendesk reports that customer support tickets are processed 20 percent faster with fewer wrong decisions. If you build AI support tools, upgrading to Sonnet 5.5 reduces your latency immediately. Your customers get accurate answers quicker, and you spend less money on compute.

Dynamic inference cost control

You can now dial effort levels from Low to Max on the Claude Platform. Running Sonnet 5.5 on Low effort beats Sonnet 5 for less than a tenth of the cost per task. This lets founders optimize profit margins based on the exact difficulty of the user request.

My take

I love seeing Anthropic focus on practical efficiency instead of just raw parameter size. Giving builders a slider for inference effort is exactly what we need to build profitable AI products. Sonnet 5.5 makes Opus look like an overpriced luxury for most real-world tasks.

Original reporting: Hacker News. This is my rewrite and opinion.

More AI news for builders