THE LAUNCH: A MODEL BUILT ON CURSOR'S DATA
Grok 4.5 reached developers on July 8 through the SpaceXAI console, the Grok Build agent, and inside Cursor itself, then opened to the public on grok.com and the X app on July 9. It's built on xAI's V9 foundation model at roughly 1.5 trillion parameters — about triple the size of the "v8-small" model xAI had in production — and it's the first SpaceXAI release built with meaningful input from Cursor, the AI coding editor SpaceX agreed to acquire for $60 billion in June. Cursor's interaction data, drawn from how millions of engineers actually write, review, and debug code inside the editor, fed directly into training, which xAI is positioning as the reason Grok 4.5 performs disproportionately well on coding and agentic tasks relative to its size. What xAI hasn't disclosed is where the boundary sits between "interaction signals used to improve the product" and "proprietary customer code used to train a foundation model" — a distinction that matters more now that Cursor is a wholly-owned SpaceX asset feeding a commercial model sold back to Cursor's own former competitors.
THE PRICE CUT: WHY $2 PER MILLION TOKENS IS THE WHOLE STORY
Every other detail about Grok 4.5 is downstream of its pricing. At $2 per million input tokens and $6 per million output, it undercuts Claude Opus 4.8's $5 and $25 by more than 60% on the headline rate, and it undercuts GPT-5.6 Sol's $5 and $30 — the model this site covered going fully public just one day earlier — by an even wider margin. Artificial Analysis measured the practical effect directly: Grok 4.5 completes an average coding-agent task for $2.49, against $11.80 for the same class of task in Claude Code, a roughly 80% reduction. Part of that gap is architectural rather than purely economic — Grok 4.5 averages 1.9 million tokens per task, against 6.2 million for GPT-5.5 and 7.2 million for Fable 5, meaning it's also simply more token-efficient at reaching an answer, not just cheaper per token once it gets there. For any team running high-volume agentic workloads where token spend is the dominant cost line, that combination is the entire pitch, independent of how the model ranks on any single benchmark.
WHAT MUSK CLAIMED VERSUS WHAT HE LATER ADMITTED
The launch messaging called Grok 4.5 "an Opus-class model, but faster, more token-efficient and lower cost" — language pitched squarely against Anthropic's current flagship. Musk's own follow-up walked that back without anyone forcing him to: asked to place it more precisely, he described Grok 4.5 as "roughly comparable to Opus 4.7, but much faster." Opus 4.7 is not Anthropic's current flagship. Opus 4.8 superseded it, and Fable 5 — Anthropic's first Mythos-class model, released June 9 and the subject of the export-control fight this site covered through most of June — sits above both. Musk's revised framing is, in effect, an admission that Grok 4.5 competes with a Claude generation that is itself two releases behind the model Anthropic is currently selling, even as the initial marketing reached for the name of the one still on the market today.
WHAT THE BENCHMARKS ACTUALLY SHOW
Independent scoring lines up closer to Musk's walked-back claim than his opening one. Artificial Analysis's Intelligence Index puts Grok 4.5 fourth among frontier models, behind Fable 5, GPT-5.5, and Opus 4.8. On SWE-Bench Pro, the gap is concrete: Fable 5 scores 80.4%, Opus 4.8 scores 69.2%, and Grok 4.5 scores 64.7% — nearly sixteen points behind the model it was marketed against. On DeepSWE 1.1, the order holds: Fable 5 at 70%, GPT-5.5 at 67%, Opus 4.8 at 59%, Grok 4.5 at 53%, though xAI's own release notes highlight friendlier numbers on benchmarks it selected itself — 62.0% on DeepSWE 1.0, 83.3% on Terminal Bench 2.1. None of those scores are disputed as fabricated; they're simply the more favorable subset of a wider picture in which Grok 4.5 is a real, usable, near-frontier model that is not, on the evidence so far, an Opus-class one.
THE TRADE-OFF NOBODY PUT IN THE HEADLINE
The cost and speed gains come with a cost of their own. Independent testing found Grok 4.5's hallucination rate roughly doubled relative to its predecessor, rising to about 54% on the evaluation in question, up from around 25%. Knowledge accuracy improved over the same comparison, from roughly 35% to 52% — but a model that knows more while also fabricating more of what it states with equal confidence is a specific, familiar failure mode: broader competence paired with less reliable self-assessment of when it's wrong. For coding-agent workloads, where a hallucinated API call or a fabricated dependency fails fast and gets caught in a test run, that trade-off is more tolerable. For the "knowledge work" tasks xAI is also marketing Grok 4.5 against — the same broader category Anthropic highlighted when it took Claude Cowork to mobile and web this week — a near-doubled hallucination rate is a materially different risk profile than the coding benchmarks alone would suggest.
THE MISSING SAFETY CARD
xAI published model cards for both Grok 4 and Grok 4.1. Grok 4.5 shipped without one. That gap isn't cosmetic: a model card is typically the document an enterprise compliance or procurement team uses to clear a new model for use in a regulated environment, and its absence is the specific reason cited for Grok 4.5's delayed availability in the EU, now expected around mid-July rather than at launch. It leaves independent evaluators — and any enterprise legal team doing its own review — working from third-party benchmarks and Musk's own public statements rather than an architecture summary, training-data disclosure, or safety-evaluation table from xAI itself. For a model trained in part on interaction data from a coding tool SpaceX now owns outright, and being pitched hardest at exactly the cost-sensitive, high-volume enterprise workloads a compliance review exists to protect, shipping ahead of that documentation is a choice, not an oversight — speed to market, priced ahead of the paperwork that would normally accompany it.
THE REBRAND BEHIND THE LAUNCH
Grok 4.5 is also the first model to ship entirely under the SpaceXAI name. The rebrand went live July 6 with a single post on X — "We are now @SpaceXAI" — completing a process Musk announced in May, when he said the xAI brand would dissolve into SpaceX rather than continue as a separate entity. The new identity folds AI model development into the same organizational umbrella as SpaceX's rockets, Starlink's satellite network, and its compute infrastructure, with Musk's stated ambition extending to orbital data centers. Grok 4.5's training data lineage — a foundation model built in part on a coding tool SpaceX bought outright two months prior — is the first concrete example of that structure actually functioning as an integrated business rather than a branding exercise, for better or worse depending on which side of the model-card gap a given customer is standing on.
THREE LABS, THREE BETS, ONE WEEK
Grok 4.5's launch doesn't sit in isolation. In the span of four days, three frontier labs each changed how their flagship model reaches customers, and each chose a different lever to pull. Anthropic moved Fable 5 to credits-only pricing for every Claude subscriber on July 7, letting the promotional 50% weekly-usage inclusion it offered when the model returned from its export-control hold expire on schedule — a bet that capability alone justifies asking users to pay directly for access. OpenAI's GPT-5.6 Sol, Terra, and Luna finished a two-week, government-requested staged rollout and went fully public on July 9, covered on this site the day it happened — a bet that a government-vetted trust signal, however contested afterward, was worth the delay. Grok 4.5 went public the same week at a fraction of either rival's price, on a model that trails both on independent benchmarks — a bet that in a market where all three labs are converging on similar coding-agent capability, the tiebreaker is what the invoice says at the end of the month. None of the three bets has settled yet, and this is the first week all three have been placed at once.
WHAT THIS MEANS FOR TEAMS EVALUATING GROK 4.5
For teams weighing Grok 4.5 against Opus 4.8 or Fable 5 for coding-agent work, the practical calculus is genuinely close, not one-sided: a roughly 80% cost reduction on measured task spend is a real, defensible number for high-volume workloads where the model's output gets checked by a test suite or a human reviewer regardless of which vendor produced it. That calculus changes for anything closer to unsupervised knowledge work, where a near-doubled hallucination rate and the absence of a model card both cut against the same use cases — tasks where nobody is independently verifying the output and where a compliance team would normally want documentation before sign-off. The sensible default for now is scoping Grok 4.5 to the workloads its own numbers actually support — cost-sensitive, test-covered, high-volume coding tasks — while treating anything closer to autonomous or compliance-adjacent work as unproven until xAI publishes the model card its two predecessors both had at launch.