Elon Musk's company xAI released Grok 4.7 on Monday afternoon. The company called the release its best model to date. They described it as a notable improvement over Grok 4.6 at the same price and speed.
The rollout followed several apparent delays. Musk had walked the timeline back at least five times since late July. He mentioned four weeks out, then a few weeks, then 3 to 4 weeks, then 10 days on September 1, and finally that it needs a few more days to cook on September 11.
According to xAI, the model spends longer working through hard problems. It also double checks its own answers more often than Grok 4.6 did. The company stated it includes their strongest safety guardrails yet.
Musk followed up on X by calling Grok 4.7 a strong combination of intelligence, speed, and low cost. There is no waitlist this time. The model is live now in the Grok app, Cursor, Grok Build, and the xAI API.
Grok 4.7 packs 2.1 trillion parameters. This is up 40% from the 1.5 trillion parameters in Grok 4.6. Grok 4.6 was itself a refinement of Grok 4.5. The costs are set at $2 per million input tokens and $6 per million output tokens.
Parameters are internal knobs a model tunes during training. More parameters generally mean more capacity to learn patterns. Tokens are the basic amount of information an AI model can either register or generate.
xAI also folded in supplemental training data pulled from SpaceX, Musk's rocket company. This data includes Starlink satellite telemetry, manufacturing records, and engineering failure logs. The pitch is a model that reasons better about hardware and physical systems than anything trained purely on internet text.
Benchmark scores present a familiar story. GDPval measures how a model performs on real, economically valuable knowledge work using tasks vetted by working professionals. It scores models as an Elo rating, which is the same head-to-head ranking system chess uses.
Grok 4.7 hit 1695 on GDPval. Claude Fable 5.1 topped the chart at 1735. AA-Briefcase, built by Artificial Analysis, tests multi-hour office work. Grok 4.7 posted 1657 against Fable 5.1 at 1678.
CursorBench 4.0 plots accuracy against cost and token count. Grok 4.7 lands in the middle. It is pricier per task than GPT-6 Astra and Claude Sonnet 5, but still short of Fable 5.1.
Musk had already set expectations lower days before launch. He wrote that Grok 4.7 should land roughly on par with Anthropic's Claude Opus 5.0, not the newer Opus 5.1, noting that multimodal performance still needs work.
In that same post, he sketched the next three models. He outlined Grok 4.8 as a meaningful step up, Grok 4.9 in the Astra and Fable class, and Grok 5 as a possible frontier leader. Those upcoming upgrades do not have a release date yet, and Musk wrote, We shall see.



