Elon Musk's SpaceXAI Launches Grok 4.7 With 71% DeepSWE Benchmark Score
Elon Musk's SpaceXAI launched Grok 4.7 on Sept 21: 71.0% on DeepSWE v1.1, 38.0% on Terminal-Bench 4.0, pricing held at $2/$6 per million tokens.
AI SummaryAI
- SpaceXAI launched Grok 4.7 on September 21 as its strongest coding and knowledge-work model.
- Grok 4.7 scored 71.0% on DeepSWE v1.1, ahead of Fable 5.1 max at 70.0%.
- Terminal-Bench 4.0 score rose to 38.0% from Grok 4.6's 20.3%.
- Grok 4.7 pricing holds at $2 per million input tokens and $6 per million output tokens.
Grok 4.7 Targets Multi-Hour Workloads
SpaceXAI, Elon Musk's AI unit, unveiled Grok 4.7 on September 21, billing it as the company's most capable model to date for coding and knowledge work. Per the official announcement, the system is built on a larger base model than Grok 4.6 and went through longer reinforcement-learning cycles on harder task mixes, with a deliberate focus on problems that take several hours to complete. SpaceXAI frames the release around endurance: where earlier models answered in minutes, this one is tuned to grind through work that would occupy a human engineer or analyst for hours. The company says the model is now markedly better at verifying its own output over long runs, at handling very long context windows, and at operating the Grok Bot agent framework natively — upgrades it says are visible in conversation, general knowledge tasks and the generation of professional documents and presentations. The published benchmarks put the xhigh configuration close to the frontier: on DeepSWE v1.1, a software-engineering suite, Grok 4.7 posted 71.0%, edging past Fable 5.1 max at 70.0% and trailing only GPT-5.6 Sol max at 72.7%. On Terminal-Bench 4.0, which scores multi-hour terminal work, the model jumped to 38.0% from its predecessor's 20.3%. On suites simulating multi-hour professional jobs for lawyers, nurses and financial analysts — AA Briefcase and HealthBench among them — SpaceXAI says Grok 4.7 now runs level with the industry's best. Safety is the other headline claim: a rebuilt safety stack makes this the company's toughest version yet at refusing improper requests and resisting jailbreaks. In HackerBench v0.3, which tests malicious cyber tasks, only 3.3% of high-risk dual-use prompts got through, with few false rejections of legitimate defensive security work. Select cybersecurity partners have been granted invitation-only red-team access to support defense research. The model is live now through Cursor, Grok Build, the Grok API, major cloud platforms and model routers.
$2 Pricing Held While Brood War Embarrasses Grok 4.6
Pricing is the quiet stunner: Grok 4.7 launches at $2 per million input tokens and $6 per million output tokens — identical to Grok 4.6 — plus a fast variant that doubles output speed at double the price. The tokenomics undercut the competition sharply: GPT-5.6 Sol lists at $4/$20 per million tokens and Fable 5.1 at $10/$50, putting Grok's rates at roughly a third to a half of rival pricing with faster delivery. Against those numbers, SpaceXAI's own comparison table shows Grok 4.7 leading on electrical engineering (EEBench: 64.0% versus 39.4% for GPT-5.6 Sol and 56.4% for Fable 5.1) and on the Harvey legal-agent benchmark (19.6% versus 2.5% and 6.7%), while trailing Fable 5.1 on software engineering (CursorBench 4.0: 46.3% versus 51.8%) and on clinical reasoning (HealthBench Professional: 56.7% versus 62.1%). Every figure above comes from SpaceXAI's published data. The corporate backdrop matters: SpaceXAI took its current shape in February, when xAI merged into SpaceX, followed by the June acquisition of Cursor — part of Musk's empire that spans Tesla (TSLA) and, per our earlier reporting, a push to consolidate AI subscription plans within weeks. Hours before launch, though, a crowd-sourced benchmark dinged the predecessor. Developer Ben Swerdlow's Brood War Bench, in which AI agents play StarCraft: Brood War, circulated on X on September 20: 19 configurations from Codex, Claude and Grok fought 171 head-to-head matches, and Codex Astra swept all 18 of its games at maximum reasoning. Grok 4.6's three setups lost to every Codex and Claude configuration, with a best result of 2 wins in 18. Swerdlow logged one match in which top-reasoning Grok burned 11,138 reasoning tokens over 43 minutes yet issued only six command batches — never fielding a single combat unit — and wrote that Grok is not yet smart enough for the game, with older models treating real-time strategy as turn-based. He added that no entrant surpassed novice level, and Grok 4.7 has not been ranked. Readers tracking the market in real time can follow live spot and futures prices on Bitget.
Pricing Pressure Is the Real Signal
The pairing of launch-day numbers with the Brood War fiasco reads as one story: frontier labs now sell reasoning at commodity prices, and agentic reliability — not static benchmark scores — decides who wins. COINOTAG's read: the $2/$6 schedule gives Grok 4.7 tokenomics rivals must match, much the way futures pricing forces markets to re-anchor expectations, while labs without comparable compute risk becoming exit liquidity in this price war. Historically, Musk compute headlines have moved sentiment in Musk-linked assets such as Dogecoin (DOGE) — watch Grok 4.7's agent results, not its launch claims.
Related Tags

AI-generated, AI-reviewed, under COINOTAG editorial oversight.


