Grok 4.5: SpaceXAI’s most capable reasoning model released
On 8 July 2026, SpaceXAI (the company formerly known as xAI, rebranded following its US$60 billion acquisition of AI coding startup Cursor) released Grok 4.5, its most capable reasoning model to date. The launch marks the first significant product release since the Cursor acquisition was completed, and it signals a deliberate strategic shift away from competing purely on academic benchmark scores toward optimising for what SpaceXAI describes as “intelligence per unit of time and cost.” For enterprise professionals and technical consultants, that framing matters more than any leaderboard position.
The model was co-trained with Cursor directly on SpaceX’s Colossus supercomputer facility in Memphis, Tennessee, a cluster housing 200,000 Nvidia GPUs. This co-training approach is notable because it means the model was not simply fine-tuned for software engineering tasks after the fact. The architectural and training decisions were made with agentic coding workflows in mind from the outset. The result is a model that resolves complex, multi-repository coding tasks in fewer steps than comparable frontier models, which has direct consequences for the economics of running autonomous agents at any meaningful scale.
For Australian professional services firms, including engineering consultancies, law firms, planning businesses, and environmental practices, the commercial implications of Grok 4.5 deserve serious attention. The combination of a sharply competitive price point, high inference speed, and demonstrated efficiency on complex knowledge work tasks positions this model as a credible backend for production-grade autonomous workflows. Businesses that have been watching agentic AI from the sidelines due to cost unpredictability now have a more concrete basis for evaluating deployment.
Key details of Grok 4.5 pricing, speed, and benchmark performance
Grok 4.5 is priced at US$2 per million input tokens and US$6 per million output tokens. These figures represent roughly half the cost of comparable frontier-class models from OpenAI and Anthropic at equivalent capability tiers. For organisations running long-horizon agentic tasks where output token counts compound quickly across many steps, this pricing differential translates into a material reduction in operational AI expenditure. The model is immediately available through Grok Build, the SpaceXAI developer console, and is natively integrated across all subscription tiers of the Cursor integrated development environment.
On the SWE-Bench Pro benchmark, which tests a model’s ability to resolve real-world software engineering tasks across complex, multi-repository codebases, Grok 4.5 achieves 4.2 times greater token efficiency than Claude Opus 4.8 (max). This is not a marginal improvement. Token efficiency at this magnitude means that for every equivalent task completed, the model consumes approximately one-quarter of the output tokens that a competing model would generate. Because pricing for large language model inference is primarily driven by output token count, this efficiency ratio is the single most important commercial characteristic of the release for businesses deploying autonomous agents at scale.
Inference speed on the fast-model tier is reported at 80 tokens per second. For context, many frontier-class models operate at significantly lower throughput on complex reasoning tasks, which introduces latency that is tolerable for batch processing but problematic for interactive debugging, real-time agent execution, or any workflow where a human is waiting on a response loop. At 80 tokens per second, Grok 4.5 is positioned for workflows that require both depth of reasoning and low enough latency to remain practically interactive.
Beyond software engineering, SpaceXAI has positioned Grok 4.5 as an enterprise knowledge work model. It currently ranks first on Harvey’s Legal Agent Benchmark, a third-party evaluation specifically designed to measure performance on legal reasoning, document analysis, and structured legal workflow tasks. The model is also demonstrated to autonomously construct complex, multi-sheet financial models drawing on live web research, and to generate formatted presentation slide decks from text outlines. These capabilities extend the model’s relevance well beyond the software development context in which it was primarily trained.

Australian business and professional services context for Grok 4.5
Australian professional services firms operate in a market where AI adoption has moved from experimental to operational across many business functions, but where cost unpredictability has remained a genuine barrier to deploying autonomous agents in production. The phenomenon commonly referred to as “token blowout” occurs when an agentic workflow, particularly one involving iterative reasoning, document review, or multi-step research tasks, consumes far more output tokens than anticipated because the model reasons verbosely or takes redundant intermediate steps. This has made budgeting for AI-assisted workflows difficult and has led some firms to cap usage in ways that undermine the productivity gains the technology was intended to deliver.
Grok 4.5’s 4.2 times token efficiency advantage on complex task resolution directly addresses this risk. For Australian legal practices, engineering consultancies, and environmental firms that have begun trialling agentic workflows for document review, report drafting, regulatory research, or data analysis, the shift from a high-output-token model to one that resolves tasks in materially fewer steps changes the financial model for these deployments. At Australian enterprise consumption volumes, the difference between US$6 per million output tokens at 4.2 times efficiency versus a competing model at equivalent capability is not theoretical. It is a budget line item that affects whether autonomous agent deployment remains financially viable at production scale or retreats back to controlled pilot conditions.
References and related sources
- Primary source: siliconangle.com
- x.ai
- gizmodo.com
- venturebeat.com
- marktechpost.com
How iEnvi can help
iEnvi integrates technology and data-driven approaches into environmental consulting. We monitor AI and technology developments that affect how environmental professionals deliver services to clients.
This is an iEnvi Machete news summary. Prepared by iEnvi to summarise the source article for environmental professionals tracking AI, data, and technology developments that affect consulting and project delivery.
Published: 09 Jul 2026
Need advice on this topic? Speak to an iEnvi expert at info@ienvi.com.au or 1300 043 684, or contact us online.
Need advice on this issue? iEnvi provides practical, senior-led environmental consulting across contaminated land, remediation, ecology and environmental risk.