xAI released Grok 4.7 on September 21, 2026, calling it the company’s most capable model for coding and knowledge work. The model is served at the same price and speed as Grok 4.6.
The Big Picture
- xAI priced Grok 4.7 at $2 per million input tokens and $6 per million output tokens, matching Grok 4.6.
- The model scored 46.3% on CursorBench 4.0, a benchmark for longer-running coding tasks, against 40.4% for Grok 4.6.
- Grok 4.7 posted 71.0% on DeepSWE v1.1 at high effort, up from 65.2% for its predecessor.
- xAI built an entirely new safeguard stack for the release, citing a 3.3% pass-through rate for risky prompts on its internal HackerBench v0.3 test.
- The model is available now in Cursor, Grok Build, the Grok API, and through third-party coding harnesses and cloud platforms.
xAI Ships a Larger Base Model
Grok 4.7 runs on a new, larger base model than Grok 4.6, according to xAI. The company trained it with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take multiple hours to complete. xAI said the model verifies its own work more carefully and manages longer context than the prior version, and it trained Grok 4.7 to natively understand the Grok Bot harness for conversational tasks and general knowledge work.
xAI’s statistics on Grok’s user base and pricing tiers show the company has scaled quickly since the SpaceX merger earlier in 2026, giving this release a larger built-in audience than prior Grok launches.
Grok 4.7 is here.
— SpaceXAI (@SpaceXAI) September 21, 2026
It's a notable improvement over Grok 4.6 at the same price and speed. pic.twitter.com/H3OTBbXyvO
Benchmark Scores Against Rivals
xAI published head-to-head scores for Grok 4.7 against Grok 4.6, GPT-5.6 Sol Max, and Fable 5.1 Max across seven benchmarks. On CursorBench 4.0, Grok 4.7 scored 46.3%, trailing Fable 5.1 Max’s 51.8% but ahead of GPT-5.6 Sol Max’s 41.7%, at a lower token price than either rival. Gains were largest in specialized domains: 64.0% on EEBench, an electrical engineering benchmark, up from 53.0% for Grok 4.6, and 19.6% on the Harvey Legal Agent Benchmark, more than triple GPT-5.6 Sol Max’s 2.5%. On Terminal-Bench 4.0, Grok 4.7 posted 38.0%, behind Fable 5.1 Max’s 57.9%.
A New Safeguard Stack
xAI said Grok 4.7 is the strongest model it has tested on refusals and jailbreak resistance, built on what it calls an entirely new safeguard stack. In dual-use domains such as cybersecurity and biological research, the company said the model leads on both utility for legitimate tasks and refusal of dangerous ones.
The model topped LatchBio’s biosafety benchmark at 62.4%, according to xAI. On HackerBench v0.3, xAI’s internal benchmark for risky and malicious cyber tasks, the company said Grok 4.7 blocked all but 3.3% of risky dual-use prompts while rarely blocking legitimate security work. xAI has also begun giving select cybersecurity partners invite-only access to Grok 4.7’s red-team capabilities for defense research.
Why It Matters?
Grok 4.7 lands as a targeted upgrade rather than a wholesale reset: same price, same speed tier, a larger base model, and a longer reinforcement learning run aimed at the multi-hour coding and office work that shorter-context models struggle to sustain. The safety framing matters as much as the scores. By pairing a refusal-resistance claim with a low false-block rate on legitimate security work, xAI is answering the standard criticism of dual-use models: tightening safeguards usually costs researchers real functionality.
Teams already routing tasks through Cursor or the Grok API get Grok 4.7 without a migration cost, since pricing did not move. Security teams evaluating red-team access should watch for xAI’s invite process rather than assume general availability. Anyone who benchmarked coding assistants against Grok 4.6 this quarter should rerun those comparisons, since the CursorBench and DeepSWE gains are large enough to shift a recent procurement call.
