Microsoft-Backed Chip Startup D-Matrix Begins Shipping AI Inference Chips

SANTA CLARA, Calif. — The benchmark that decided d-Matrix’s future ran in a company lab this spring. Engineers paired a pair of the startup’s Corsair accelerators with an Nvidia Blackwell GPU and watched a response time drop from 24 seconds to under two — a roughly tenfold improvement over a GPU-only setup, according to testing by Gimlet Labs that the company now cites in customer meetings.

On June 9, d-Matrix said the chip had entered full production, with volume shipments to priority customers beginning this month. The announcement moves the company from sampling to scale at a moment when the economics of artificial intelligence are shifting from training models to running them — the inference market that Nvidia has dominated by default.

d-Matrix was founded in 2019 by Sid Sheth, its chief executive, and Sudeep Bhoja, veterans of Broadcom’s networking silicon group. The company has raised roughly $500 million at a valuation near $2 billion, with Microsoft among its investors through the M12 venture arm. The Microsoft connection is notable given the software maker’s own silicon ambitions: Maia 200, its inference chip; PC processors built with Nvidia; and an in-house quantum chip announced last week. D-Matrix is the external bet in a portfolio of internal ones.

Corsair’s design departs from the industry’s standard playbook. Instead of HBM memory stacked on CoWoS packaging, the chip uses an SRAM-based in-memory compute architecture on organic substrates, paired with LP-DDR5 memory. The choice was deliberate, Mr. Sheth said: it avoids the packaging bottlenecks that have slowed other accelerators and keeps the supply chain predictable. Chips are manufactured with TSMC on its N6 node, with Alchip Technologies handling packaging, and they slide into standard PCIe slots in existing data-center racks. Four chips come packaged on a card that sells for tens of thousands of dollars.

The performance claims are specific. Paired with a Blackwell GPU in a disaggregated pipeline, Corsair runs inference about ten times faster, at roughly a third of the cost and up to five times the energy efficiency of a GPU-only setup, according to Gimlet Labs testing cited by the company. For agentic workloads, where a model must reason, call tools and compose answers across many steps, the latency difference compounds: a slow inference step repeated dozens of times turns into a visibly slow assistant.

d-Matrix is not selling chips alone. It built SquadRack, a rack-scale system developed with Arista, Broadcom and Super Micro, so customers can deploy Corsair without redesigning their data centers. Mr. Sheth declined to name buyers but said he has commitments from hyperscalers, neoclouds and frontier AI labs, with about 90% of them in the United States and the rest in the Middle East and Southeast Asia.

The timing reflects a market in motion. Training models remains the industry’s biggest expense, but the cost of serving them has become the fastest-growing line on AI data-center budgets as agentic software pushes inference demand far beyond what general-purpose GPUs were designed for. That has opened the door for a wave of purpose-built silicon — OpenAI and Broadcom’s Jalapeño inference chip, IBM’s sub-nanometer stack, Qualcomm’s software expansion — all chasing the same realization: GPUs win at training, but inference is where differentiation lives.

Analysts who follow the sector said d-Matrix’s production announcement is early evidence that the niche is real, but the proof will come in adoption. Nvidia’s CUDA ecosystem remains the default for developers, and any challenger must convince customers that a second stack is worth operating. The company’s answer is economics: for inference-heavy workloads, the power and cost savings pay for the integration work within months, Mr. Sheth said.

Both founders came from Broadcom’s networking group, where they spent years building the switching silicon that moves data inside data centers. The in-memory approach they chose — placing compute next to SRAM instead of shuttling data to and from HBM — is the same broad strategy pursued by Cerebras and Groq, which also bet that memory bandwidth, not raw compute, is the binding constraint on inference speed. What distinguishes Corsair is the packaging: by building on organic substrates and standard PCIe cards, d-Matrix avoids the CoWoS advanced-packaging lines that have been a bottleneck for GPU suppliers, and the chips can slot into racks that already exist. The supply-chain argument has become a sales argument; customers who waited months for accelerators want chips that arrive on schedule. Mr. Sheth said the company designed Corsair with supply-chain predictability as a core requirement, not an afterthought. The inference market it is chasing is large and growing: analysts estimate the cost of serving AI models is becoming the dominant line on data-center budgets as agentic workloads multiply the number of inference calls per task.

The roadmap extends the bet. Raptor, the next chip, is planned for 2027 on TSMC’s 4-nanometer node, and Mr. Sheth said it could run out of the foundry’s new factory in Arizona, which would make d-Matrix an early customer of U.S.-based advanced manufacturing. For a company that built its pitch around supply-chain predictability, the line from Taiwan to Arizona is the next chapter of the same story.

Related Posts

  • September 6, 2026
  • 10 views
Anthropic Moves Its IPO Filing to Late September

The bankers and lawyers running Anthropic’s initial public offering had told investors to expect the company’s registration documents as soon as this week. The calendar has moved. Anthropic now plans…

  • September 6, 2026
  • 11 views
OpenAI Quietly Revises GPT-6 Astra Scores After Launch

When OpenAI released GPT-6 Astra on Sept. 3, the launch post carried the usual furniture of a modern model debut: coding results, speed comparisons and a figure for how often…