Why Traditional Benchmarks Fail Modern AI Models with OpenAI Research Scientist Noam Brown
In a Nutshell
Traditional benchmarks understate model capabilities by not accounting for test-time compute scaling. Modern models like o1.5 show substantial gains when evaluated across compute budgets, but these improvements are invisible in standard benchmark grids that lack a tokens/cost/time axis. Safety evaluations and responsible scaling policies remain misaligned with this reality, as they fail to test what models can achieve with large inference budgets, creating evaluation gaps that grow worse with each release cycle.
These notes were generated by AI and may contain inaccuracies.
With GPT-3, you couldn't scale test time compute. If you gave it a budget of $10 million and asked what GPT-3 could do, it really couldn't do that much more than what you could do with $10 or $1. The current frameworks and responsible scaling policies don't really account for the amount of test time compute. They just say what's the capability of the model. The problem is we're in a world now where the capability of the model is a function of how much money you put into it. If you give it a budget of $10,000, it can do a lot more than what it can do with a budget of $10. Give it a budget of $10 million, you can do even more. At what budget should you evaluate these models? The policies that exist today don't really address that question.
The motivation was the release of o1.5 and the initial reaction was skepticism that it was a substantially better model. That only lasted for a few hours before people had time to play around with it and saw that it was actually substantially better. A lot of the skepticism came from the benchmark grid that was published. Whenever a new model is released, there's this benchmark grid where they show all these different benchmarks on the x-axis and the performance of different models on the y-axis. If you look on paper at the difference between o1.5 and o1.4 or other models, it was an improvement, but it wasn't a huge improvement. It was only a few percentage points in some benchmarks.
Sign in to read the full notes
Get access to AI-generated notes, topic timestamps, and more.