Vals, a benchmarking startup formed in 2024, has raised $40M in a Series A led by Andreessen Horowitz, after a seed round led by 8VC and Bloomberg Beta.
Benchmarks drive model marketing, but legacy tests are old and widely published, which lets labs train against them. Vals keeps its materials private and, instead of quizzing models on general knowledge, scores them on tasks drawn from law, finance and coding.
“We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks were not keeping up with that frontier advance,” co-founder Rayan Krishnan said.
The 25-year-old, who interned at Palantir and worked at Microsoft and Stanford’s AI lab as an undergraduate, argues evaluation should verify whether a model can produce human-quality work inside a domain, rather than whether it can pass a bar-style exam. Vals also probes downside behaviour, asking what happens if a model is let loose.
As corporate buyers increasingly demand proof before they sign, independent evaluation is turning into infrastructure of its own.