On a Saturday evening in late June, engineers at SpaceX and Tesla began receiving access to a new artificial-intelligence model, an early signal of the release cadence xAI has promised for the rest of the year. Elon Musk said on June 28 that Grok 4.5 had entered internal testing at both companies, describing it as the successor to xAI’s flagship chatbot built on a foundation of the company’s V9 base model, which he said contains 1.5 trillion parameters.
The model’s training included an unusual addition: data drawn from Cursor, the coding tool that has become one of the fastest-growing developer products of the past two years. Musk said early evaluations showed Grok 4.5 performing close to, and in some tests above, Claude Opus, the flagship model from Anthropic. He added that reinforcement learning is still improving the model and that xAI’s internal “GrokBuild” benchmark suite is being refined alongside it.
The internal test matters because of where it is running. SpaceX and Tesla are not typical software shops. Tesla’s full-self-driving stack, its Optimus humanoid program and the manufacturing lines that build both are increasingly shaped by large models. SpaceX operates rockets and satellite constellations under hard real-time constraints. Putting a new model in front of those engineers, rather than releasing it directly to consumers, gives xAI real-world feedback that a public chatbot launch cannot provide, according to people familiar with the company’s process.
The V9 lineage places the model among the largest that any lab has described publicly. xAI has scaled its base models aggressively since the company’s founding, and a 1.5-trillion-parameter foundation would put Grok 4.5 at the top of that progression. Musk has long argued that raw scale, combined with training data and reinforcement learning, is what separates frontier models from the rest, and the internal-testing announcement is the first time the company has put that claim in front of working engineers rather than benchmark watchers.
The Cursor connection is the more telling detail. Coding has become the most contested front in the AI race, with Cursor, GitHub Copilot and Claude Code fighting for developers, and model makers measuring themselves against coding benchmarks the way automakers once quoted horsepower. By training on Cursor’s data, xAI is signaling that it intends to compete where the usage is densest, and where developer loyalty translates into subscription revenue. Analysts who follow the market said the move suggests xAI is chasing leadership in software-engineering tasks ahead of a broader release, a category in which OpenAI and Anthropic have set the current standard.
Musk’s post also carried a broader promise: SpaceX, he said, will release a fully new, from-scratch model every month for the remainder of the year. That cadence would be a departure even for an industry that has grown accustomed to rapid iteration. Most frontier labs refine a single architecture for months, adding data and compute, before unveiling a successor. Training entirely new models monthly implies enormous compute and an aggressive willingness to retire work that does not survive evaluation.
The computing requirements are the least hidden part of xAI’s strategy. The company has built out its own data-center footprint, a project that has consumed billions of dollars and drawn attention to the power and water demands of AI infrastructure. A monthly release schedule would strain that capacity further, though Musk has repeatedly said compute is the constraint xAI is most willing to spend against. Skeptics note that a from-scratch model every month raises quality risk: some months will produce models that fail internal evaluation, and the company has not said what it will do when that happens.
For Tesla shareholders, the announcement carries a familiar subtext. Musk has argued that Tesla’s value rests partly on its AI capabilities, from self-driving software to the Optimus robot. A model trained partly on coding data is not the same as a model that can drive a car, but the company’s executives have described a shared foundation underneath both. The internal test gives Tesla engineers early access to capabilities that might eventually appear in the vehicles’ on-board systems.
Anthropic and OpenAI have not responded publicly to the claims about comparative performance. Benchmark results posted by model makers have a history of being contested, and independent evaluations will determine whether Grok 4.5’s early numbers hold up. Analysts noted that xAI’s approach, benchmarking against a rival’s flagship while still in internal testing, is as much marketing as measurement.
The race’s pace is the story underneath the announcement. A year ago, flagship models arrived roughly twice a year. Now xAI is promising twelve from-scratch releases in twelve months, and its competitors are compressing their own cycles. The internal test at SpaceX and Tesla is the first step in that schedule, and the first chance for the wider world to see whether the model behind it is as fast as the release plan implies.


