a16z Leads a $40 Million Round in AI Evaluation Startup Vals

Vals, a startup that sells private tests for measuring the abilities of artificial-intelligence models, has raised $40 million in a Series A round led by Andreessen Horowitz, valuing the company at about $400 million. The round follows a seed investment led by 8VC and Bloomberg Beta, and it positions Vals as a bet on a question the AI industry has not yet answered: how to know, reliably, what a model can actually do.

The company’s premise is that the public benchmarks the industry has used to rank models have been compromised. When a model’s score on a widely known test becomes a marketing number, the incentive to train against that test grows, and the score stops measuring general ability and starts measuring preparation. Vals instead builds private test sets in professional fields such as law, finance, engineering, and medicine, pairing domain-specific tasks with automated scoring to produce an evaluation that a model cannot have studied for.

The private, domain-specific approach is a direct response to a specific problem. Public benchmarks have repeatedly been found to have leaked into training data, inflating scores in ways that flatter the model and mislead buyers. By keeping its test sets confidential and refreshing them, Vals is selling the thing the industry lost when its measuring sticks became targets: a signal that has not been gamed.

Vals was co-founded by Rayan Krishnan, who is 25 and whose résumé runs through Palantir as an intern, Microsoft, and Stanford’s AI lab before he started the company. The company’s pitch is that the faster models improve, the more valuable independent measurement becomes, because every capability claim a lab makes needs a third party to verify it and every enterprise buyer needs a reason to trust the claim.

Alongside the funding, Vals announced a set of new products: Vals Smith, a frontier risk benchmark, and Vals Index 2.0. The additions signal an ambition that reaches beyond basic capability testing and toward the questions that regulators and safety researchers have begun to ask, such as whether a model will behave dangerously under pressure and how its performance holds up as it is updated.

The evaluation business has grown into its own layer of the AI economy. As frontier labs have become more secretive about their models’ abilities and more aggressive in their marketing, the demand for independent assessment has come from every direction at once: enterprises deciding what to buy, regulators deciding what to allow, and investors deciding what a company is worth. Vals is one of a cluster of startups trying to become the auditor for an industry that has outgrown its old report cards.

Andreessen Horowitz’s involvement matters because the firm is a large investor in the very labs whose models Vals would evaluate. The round is a wager that the measurement layer can be a profitable business even when it sits adjacent to the firms that make the models, and that neutrality can be maintained in a market where many of the parties paying for evaluations are also parties with a stake in the results.

The challenge for Vals is the one that faces every evaluator in a fast-moving field: the tests must stay ahead of the models. A private test set is valuable only while it remains unpublished and unstudied, which means Vals must continuously produce new professional-grade problems, a costly and labor-intensive undertaking that does not scale the way model training does. The company’s domain focus is an attempt to make that production disciplined rather than improvised.

The 25-year-old founder’s age has drawn attention, but the structure of the company’s argument is older than he is. Every market that produces claims eventually produces people paid to check the claims, and AI is now producing claims faster than any technology in memory. The models have become better at a rate that has outpaced the tools used to measure them, and that gap is the business Vals is trying to occupy.

The company’s bet is that evaluation will become a standing cost of doing business for the labs, the way auditing is for public companies. If that proves true, the evaluators sit in a durable position, selling a recurring service rather than a one-time test. If it does not, Vals is a specialized consultancy with a clever product, and the market for its services will stay small. The funding round is, in effect, a wager on which of those futures arrives first.

The round was oversubscribed, and the terms reflect a market that has come to regard measurement as a real business rather than a side effect of research. Whether Vals can hold the position the round implies depends on execution more than on the idea, which is now widely shared: the models have outrun their report cards, and someone has to write new ones.

Related Posts

  • September 23, 2026
  • 15 views
Hack VC’s Former Partner Found Dead in the California Desert

Hsin-Ju Chuang spent nearly a decade inside the crypto industry’s fastest-growing companies, including a stretch running growth at Solana. In the final weeks of her life, she had turned against…

  • September 23, 2026
  • 22 views
SpaceX Stops Taking Falcon 9 Bookings Beyond 2028

For years, satellite operators planning a launch a few years out could simply call SpaceX and secure a spot on a Falcon 9. That option is now closing. SpaceX has…