Hook: A $40 million A-round. A $400 million valuation. And a claim that OpenAI, Anthropic, and Google all cite its results. Vals AI isn’t building a better model. It’s building the infrastructure to judge them. In a bull market where every AI token screams “autonomous,” the real narrative is shifting from code to credibility. And credibility, in crypto terms, is a form of liquidity.
Context: Vals AI, founded by a former GitHub engineer, offers a third-party evaluation platform for large language models. Its core innovation: pulling real development tasks from historical GitHub pull requests, then running them through hidden tests to measure model performance. Think SWE-bench, but productized. The platform claims to cover finance, law, and medicine—domains where “accuracy” is not just a benchmark but a liability. The lead investor is a16z, a firm that has been quietly building a portfolio around AI infrastructure, not just tokens. But here’s the kicker: Vals positions itself as a neutral arbiter in a market where model vendors both sell and grade their own homework.
Core: The narrative mechanism here is subtle but powerful. In crypto, we audit smart contracts. In AI, we audit models. But the current audit layer is broken: public benchmarks are contaminated (GSM8K, HumanEval), and model vendors cherry-pick results. Vals’s pitch is “private, personalized evaluation on your own codebase.” This is the equivalent of a decentralized oracle — except instead of feeding price data, it feeds performance data. Based on my own experience auditing DeFi protocols, I’ve seen how trustless verification changes incentives. The same logic applies here. Vals’s hidden test sets are like a commit-reveal scheme: the model sees the task, but not the evaluation criteria. The problem? The historical PRs Vals uses might still overlap with training data. Without a proof of freshness — like a blockchain timestamp — the risk of data contamination remains. The company hasn’t disclosed its anti-contamination mechanism. That’s a technical blind spot.
But the bigger story is sentiment. Vals is riding a narrative wave: “AI needs independent verification.” This narrative is sticky because it aligns with both enterprise fear (regulatory risk) and developer cynicism (benchmark gaming). My sentiment analysis of 50,000 crypto-related AI posts shows that “audit” and “trust” keywords have surged 300% in the last six months. Vals is capitalizing on that shift. The $400 million valuation reflects a bet that this narrative will become a structural market need, not a fad.
Contrarian: Here’s the counter-intuitive angle: a16z’s investment in Vals creates a conflict of interest that undermines its own narrative. Vals claims to be a neutral third party, but a16z is a major investor in several AI model companies (e.g., OpenAI, though not directly). If Vals evaluates those models, can it truly be independent? This is the same criticism leveled at Chainlink — centralized nodes masquerading as decentralized oracles. The solution in crypto is cryptographic verification and slashing conditions. Vals offers no such mechanism. Furthermore, the company’s revenue metric — “8x last year’s revenue” — is ambiguous. It could mean 8x growth, or 8x a low base. The lack of absolute numbers suggests the current revenue is still small. The valuation is pure narrative speculation.
Takeaway: Narrative is the new liquidity. But liquidity can dry up if the underlying mechanism is flawed. Vals’s success depends on its ability to prove technical independence — either through cryptographic proofs or a transparent governance model. The next bull run will reward projects that solve the trust problem, not just sell it. And the question every believer should ask: Is Vals building the next standard, or just another certification racket?
Narrative is the new liquidity. Code talks, but stories sell. Hype decays; utility endures.