SpaceXAI drops Grok 4.7 for coding and knowledge work at $2 per million tokens
Grok 4.7 matches frontier models on long coding tasks but keeps the $2 input price. This puts massive pressure on GPT-5.6 Sol and Opus 5.
Models · Source: Hacker News
What happened
SpaceXAI just released Grok 4.7. It is a new foundation model built specifically for software engineering and long-context knowledge work. The model uses a larger base than Grok 4.6. It was trained with longer reinforcement learning runs on a harder mix of tasks. The training weighted heavily toward problems that take many hours to complete. This makes it better at verifying its own work and managing long contexts.
The pricing strategy is aggressive. Grok 4.7 costs two dollars per million input tokens. Output tokens cost six dollars per million. This matches the exact price and speed of Grok 4.6 while delivering significantly better performance. SpaceXAI also introduced a fast variant. This version doubles the output speed but costs twice as much. The model is available today in Cursor, Grok Build, and through the Grok API.
Benchmark results show strong gains in specialized professional fields. Grok 4.7 hit 71 percent on DeepSWE using high effort. It scored 19.6 percent on the Harvey Legal Agent Benchmark. It also features a completely rewritten safeguard stack. This new stack blocks malicious cyber tasks while allowing legitimate security research. It tops the LatchBio biosafety benchmark at 62.4 percent. SpaceXAI is also giving select cybersecurity partners invite-only access to the red-team capabilities of Grok 4.7 for defense research. This shows a clear push into enterprise security markets.
Key facts
- $2 — Input token price per million
- $6 — Output token price per million
- 71.0% — Score on DeepSWE v1.1 with high effort
- 3.3% — Risky dual-use prompts allowed on HackerBench v0.3
Why it matters
The cost of running autonomous agents just dropped significantly. Builders running multi-hour terminal tasks or complex coding loops face a tough choice. You usually have to choose between cheap models that fail at reasoning or expensive models that destroy your profit margins. Grok 4.7 sits right in the middle of this spectrum. It gives you frontier-level reasoning on long tasks without the massive price tag of Fable 5.1 or GPT-5.6 Sol. When you build products that require models to think for hours, token costs compound rapidly. Grok 4.7 changes that math entirely. You can now run extensive agent loops without burning through your runway.
This release shifts the economics of vertical AI applications. Models are getting much better at specific professional workflows. We see this in the electrical engineering and clinical reasoning benchmarks. Grok 4.7 natively understands the Grok Bot harness. This means founders can build specialized agents for lawyers, nurses, and financial analysts much faster. The GDPval benchmark shows Grok 4.7 reaching an Elo score of 1695, beating GPT-6 Astra. This proves that you do not need to pay maximum prices to get maximum performance in professional knowledge work. The barrier to entry for complex enterprise AI is falling rapidly. Startups can now deliver professional-grade automated work at a fraction of the historical compute cost.
For builders
Cheaper long-running coding agents
Grok 4.7 is optimized for problems that take hours to complete. You can run extensive coding loops in Cursor or Grok Build without burning cash. Startups building autonomous developer tools win big here. Your compute costs stay low while capability rises.
Safer cybersecurity applications
The new safeguard stack allows benign security work while blocking malicious prompts. This is a massive unlock for founders building AI defense tools. You get the reasoning power without the constant false-positive refusals. Security teams can finally automate defense research.
Double speed for premium users
SpaceXAI offers a fast variant that doubles output speed for twice the price. If your product relies on real-time user experience, you can pay a premium to cut latency. This lets you tier your own pricing based on speed. Enterprise customers who need instant answers will pay for this.
My take
I see SpaceXAI commoditizing the middle layer of AI reasoning here. By keeping prices flat while boosting performance on multi-hour tasks, they are forcing everyone else to adapt. If you build agents, I suggest making Grok 4.7 your new default workhorse.
Original reporting: Hacker News. This is my rewrite and opinion.