The Tests Kept Passing
The descriptions keep sounding familiar. Someone writes about an eval set that has stopped catching regressions. Someone else describes a judge model that agrees with itself on every run and is wrong in the same direction on every run. The two problems have nothing to do with each other, and both read like something already lived through. There was no particular article where that landed, no conversation that set it off. It accumulated, the way these things do, until the recognition got frequent enough to be worth distrusting. ...