Alex Ratner spent years telling companies the bottleneck in artificial intelligence was not the model but the data. On September 22, the market put a number on that idea.
Snorkel AI, the startup Ratner co-founded, said it raised $350 million at a valuation of $3.5 billion. The round, led by Insight Partners and S32, values the company at nearly three times the $1.3 billion it was worth when it last raised money in May 2025.
The speed of the jump is the story. Seventeen months ago, Snorkel sold software that helped enterprises label their own training data. Since then it has turned itself into what it calls a “data factory,” producing finished datasets and reinforcement-learning environments and selling them directly to the AI labs building frontier models.
Ratner and his co-founders came out of Stanford’s AI lab in 2019, where they built the open-source Snorkel project around a technique called weak supervision, using rules and patterns to label data cheaply rather than paying annotators for every example. The company’s early pitch was that data, not model architecture, determined how well AI systems performed.
That argument has aged well. As frontier labs converged on similar models, the variable that still separates them is the quality of the data they train on. Ratner has framed the shift in blunt terms: most data that labs will find valuable will need human input, and nearly all of it will have to be produced with synthetic and automated methods to keep up.
The company’s customers now include frontier labs, hyperscalers, enterprises and the U.S. federal government. Coding data is one of its largest areas of demand, as labs race to train models that write better software. Growth has been sharp enough that the company says it expects to reach profitability this year, even as it keeps spending to hire specialists and build tooling.
The round pulled in a broad syndicate. Addition, Greylock and Wells Fargo returned, joined by new investors including March Capital, Third Point Ventures and Blumberg Capital, along with Lightspeed and GV. The list reflects how mainstream the business of supplying training data has become.
The sector was transformed by a single deal in 2025, when Meta bought a 49 percent stake in Scale AI for roughly $14.3 billion. That price told investors that data was not a service layered on top of AI but a scarce input with its own scarcity value, and it drew capital toward rivals such as Mercor and Surge AI.
Scale AI is the reference point the market now uses. Alexandr Wang founded the company in 2016 at age 19, and it became the dominant supplier of human-annotated data before Meta’s investment. Snorkel’s path has been different: it started with software and a research pedigree, and only recently began selling finished data. The two companies now overlap more than either founder would have conceded a year ago.
Snorkel’s distinction is its blend of people and software. It pairs AI tooling with a network of specialists in coding, law and medicine who design tasks and grading rubrics, while models handle quality control. The pitch is that the most complex training data needs both human judgment and automated scale.
Underneath the fundraising is a shift in what the industry needs. The public web, once an effectively free source of text, has been largely mined, and the easiest labeling tasks have been automated. What remains are the hard problems: data that requires a doctor, a lawyer or an engineer to judge, and simulated environments where models can practice tasks that have no ready-made dataset. That is the scarcity investors are paying for.
The bet carries risk. The valuation has tripled on the expectation that demand for specialized data keeps rising, and that AI labs will keep buying rather than build the capability in-house. If model progress slows or labs consolidate their suppliers, the premium now baked into the price would be tested.
Ratner has argued the opposite: that data becomes more scarce and specialized as models improve, and that no lab can produce everything it needs internally. The more capable the models become, the more demanding the datasets and simulated environments required to train and evaluate them.
The shift from software to services widened Snorkel’s market at the same moment the industry’s appetite for data exploded, and the round gives it capital to hire specialists and build the tooling to keep pace. Customers appear to agree, at least for now.
What the valuation buys, beyond the money, is a claim on a simple thesis: that in an AI market where models are starting to look alike, the company that controls the data holds an edge. Snorkel has spent seven years arguing that point. Investors just paid $3.5 billion to say they believe it.


