Andreessen Horowitz-backed Vals aims to set the benchmark of excellence in AI evaluation

Business 0 views Source: autosite

Vals, a 2024-founded AI evaluation startup, raised $40 million in Series A funding led by Andreessen Horowitz to replace outdated academic benchmarks with private tests of real-world, domain-specific tasks in law, finance, and coding. Revenue has grown 8x year over year, the team tripled to 25, and the company now serves federal agencies as AI firms prepare to go public.

AI benchmarking has become the standard way companies prove their models work — and, when the numbers look good, a powerful marketing tool. But many widely used benchmarks were designed years ago and struggle to measure what today's frontier models can actually do. Vals, a startup founded in 2024, wants to replace that aging system, and investors are buying in: after a seed round led by 8VC and Bloomberg Beta, the company raised $40 million in a Series A last month, led by Andreessen Horowitz.

A Founder Who Saw the Gap Coming

Rayan Krishnan, Vals' 25-year-old co-founder, interned at Palantir and, while an undergraduate at Stanford, worked for Microsoft and the university's well-known AI lab. He says the company grew out of his frustration watching benchmarks fall behind the very models they were meant to measure.

"We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks [were] not keeping up with that frontier advance," Krishnan said. With AI now woven into nearly every sector, he argues, benchmarks should exist to confirm that models genuinely deliver what vendors promise.

Private Tests, Real-World Tasks

Vals sets itself apart in two key ways. First, unlike systems with publicly available test questions — which companies can effectively train their models on, amounting to studying for the answer key — Vals keeps its test materials private. Second, rather than testing general knowledge, the startup evaluates how models handle complex, domain-specific tasks in fields like law, finance, and coding.

"Historically, I think evaluation has been done to evaluate intelligence in a very abstract way," Krishnan explained, comparing it to asking whether a model knows enough to pass a bar exam. Instead, he said, "Can they do work that produces a product of the same quality as a human within every domain?"

The startup also looks for failure modes, not just successes — trying to understand what the negative consequences would be "if these models ran wild in the world." Its scope keeps expanding beyond traditional industries: Vals is building benchmarks around recursive self-improvement, mental health, cybersecurity, biosecurity, and even the law of armed conflict, exploring how models might apply the Geneva Convention.

Fast Growth and a New Federal Focus

Clients pay Vals to evaluate their models — an unusual arrangement, since it means paying to hear your product fall short. But Krishnan likens the revenue model to students paying the College Board to sit the SAT: reliable measurement helps companies diagnose weaknesses and improve. The results, in turn, increasingly inform which AI models enterprises choose to adopt.

The business is scaling quickly. Revenue is currently eight times what it was last year, and the team has tripled from eight people at the start of the year to 25. Krishnan plans to hire another 10 to 15 employees and move to a much larger office. The company has also launched a program offering model evaluations to federal agencies.

Krishnan believes his company's approach will shape how AI companies build trust as they mature into public firms. "AI companies are starting to go public," he said, noting that Anthropic is slated for later this year and predicting an OpenAI listing soon. As models become central to the economy, he expects the kind of evaluations Vals performs to drive adoption — and to feature in public filings and investment disclosures.

Whether traditional benchmarks can catch up, or Vals becomes the industry's de facto referee, is one of the next big questions in AI accountability.

Meta description: Vals raised $40 million led by Andreessen Horowitz to replace outdated AI benchmarks with private, real-world model evaluations. Inside its fast growth.

Tags: Vals, AI benchmarking, Andreessen Horowitz, AI evaluation, AI startups

Featured image: Abstract technology style — a stylized digital gauge or data-score visualization in dark tones, no people, logos, or text.

Tags: AI startupsValsAI benchmarkingAndreessen HorowitzAI evaluation