Public benchmarks are saturating and models are getting better at optimising for the tests themselves, so a16z brings on Vals founder Rayan Krishnan to argue for independent evaluations that keep evolving. The conversation covers why self-reported model scores mislead, how Vals tests a model in the hours before release, and what changes when token spend starts to rival salaries.