xAI launches Grok 4.5, its strongest model, with a 500,000-token context window

xAI released Grok 4.5 on July 8, calling it the company's strongest model to date. Trained on tens of thousands of Nvidia GB300 GPUs and built on a new 1.5-trillion-parameter foundation model xAI calls V9, the model is now live in Grok Build, on all Cursor plans, and through the xAI API console.
What the model can do
Grok 4.5 supports a 500,000-token context window, along with text and image input, function calling, structured outputs, web and X search, code execution, and configurable reasoning depth. xAI is serving it at roughly 80 tokens per second, and the company says it delivers twice the token efficiency of competing frontier models on comparable tasks — meaning it can produce equivalent results while consuming fewer tokens, which lowers the effective cost of a given task even before accounting for the sticker price.
Pricing is set at $2 per million input tokens and $6 per million output tokens, with a discounted $0.50 per million for cached input and higher rates once a request exceeds 200,000 tokens.
Benchmark results, with some caveats
xAI is positioning Grok 4.5 primarily as a coding and agentic-work model. On the SWE Marathon benchmark — a test of sustained, multi-step software engineering tasks — Grok 4.5 posted a 29.0% pass@1 resolution rate, ahead of Anthropic's Opus 4.8 at 26.0% and xAI's own prior Fable model at 24.0%. The model also leads on DeepSWE 1.0 (62.0%) and Terminal-Bench 2.1 (83.3%).
That leadership isn't universal, though. Of the four benchmarks xAI chose to publish alongside the launch, Grok 4.5 wins two and loses two to Opus 4.8 — a split that suggests the model is a genuine step forward in specific coding workflows rather than a blanket improvement over every rival on every task.
Availability gap in the EU
Grok 4.5 is not yet available to European users, either through xAI's own products or its API console. xAI says EU access is expected in mid-July, without giving a firm date. That delay places Grok 4.5 alongside a pattern of AI model rollouts in 2026 where regulatory review — whether EU AI Act compliance checks or similar processes — has separated a model's initial launch date from its availability in Europe by days to weeks.
Competitive context
The release lands the same week OpenAI is completing the phased rollout of its GPT-5.6 family (Sol, Terra, and Luna) and puts pressure on Anthropic, whose Opus 4.8 remains competitive on two of the four benchmarks xAI selected. For developers choosing between frontier coding models, the practical takeaway is that no single model currently leads across the board — Grok 4.5's SWE Marathon and Terminal-Bench results make it a strong pick for certain agentic coding tasks, while Opus 4.8 remains ahead on the benchmarks xAI didn't win.
Originally reported by xAI. Read the original article for additional details.
View original source