Editorial still-life photograph of scattered benchmark scorecards and AI evaluation reports with red pencil marks beside an orange clipboard and magnifying glass on a warm ivory surface

You've Been Buying AI Tools Based on Rigged Report Cards

September 21, 2026

Last month I sat with a client who'd just spent $2,400 a month on an AI writing tool. Their sales team picked it because it ranked #1 on a popular benchmark. After six weeks, the tool couldn't handle their industry's compliance language. It hallucinated regulatory terms that don't exist.

That benchmark tested general knowledge and creative writing. Not a single question about regulated industries.

This is the problem Vals just raised $40 million to fix.

The testing system is broken

Vals, a two-year-old startup backed by Andreessen Horowitz, closed its Series A this week. The company builds domain-specific benchmarks that test whether AI models can actually do real work in fields like law, finance, and software engineering.

Co-founder Rayan Krishnan puts it bluntly: most public benchmarks measure trivia, not job performance. AI companies know which tests matter for PR, and they train specifically to ace them. It's the equivalent of a student memorizing the answer key instead of learning the material.

The numbers back this up. Models that score within a point of each other on popular leaderboards can perform 30% differently on real-world tasks. That gap is where your money disappears.

Why this matters if you buy AI tools

If you run a business and you're comparing AI products, you've probably looked at benchmark scores. Maybe a vendor sent you a comparison chart showing their model beating competitors on MMLU or HumanEval. Those scores aren't fake, but they're incomplete in ways that cost you money.

Think about what's really being tested. Public benchmarks ask questions like "What's the capital of France?" or "Write a Python function to sort a list." Your business needs the AI to understand your specific workflows, terminology, and edge cases. No public benchmark tests that.

Vals builds custom evaluation suites for specific industries. A legal benchmark tests contract analysis, not trivia. A finance benchmark tests earnings call interpretation, not math puzzles. The difference between a model that scores well on general tests and one that actually handles your industry's complexity is often the difference between an AI budget that pays for itself and one that doesn't.

The bigger signal

This funding round tells you something about where the AI market is heading. When investors put $40 million behind "testing AI properly," it means the market has matured past the phase where impressive demos close deals.

Buyers are getting burned. The honeymoon period where every AI tool seemed magical is ending. Companies are asking harder questions: Does this actually work for my use case? Can you prove it?

If you're an AI agency or consultant, this shift changes your pitch. Your clients will start asking for evidence that goes beyond vendor marketing. Having your own evaluation process, even a simple one, gives you an edge.

What you can do right now

You don't need to wait for Vals to test your specific industry. You can build a rough version of this yourself.

Pick 20 real tasks from your actual workflow. The messy ones. The ones with edge cases and weird formatting. Run them through whatever AI tool you're evaluating. Score the outputs honestly. That homegrown test will tell you more than any leaderboard.

I've done this with clients at Laimen AI, and the results always surprise people. The model that looks best on paper rarely wins the practical test. Sometimes the second or third option handles your specific needs better because it was trained on data closer to your domain.

The $40 million question

Vals is betting that neutral, verifiable benchmarks become the standard for AI procurement. Think of it like Underwriters Laboratories for AI. Before you buy a toaster, someone tested it so it won't burn your house down. Before you buy an AI tool, someone should test whether it actually does what the vendor claims.

We're not there yet. But the fact that a16z just wrote a big check says the industry is moving in that direction. In the meantime, be your own benchmark. Test with your own data, your own edge cases, your own definition of "good enough."

The vendors who welcome that scrutiny are the ones worth paying.

— Mark Garza, Laimen AI

Mark Garza

Mark Garza

Mark is an automation and AI growth strategist and the founder of Laimen AI.

LinkedIn logo icon
Back to Blog