How exactly does scaling inference compute affect the performance and reliability of LM agents in agentic benchmarks? We believe that most current evaluations underestimate performance because they do not account harness.
How exactly does scaling inference compute affect the performance and reliability of LM agents in agentic benchmarks? We believe that most current evaluations underestimate performance because they do not account harness.
Project Details
Updated 07/26/26 · Edited by orgThe project idea is taken from the Apollo list. Modern models demonstrate significant performance gains through scaling inference computations, but for agent systems, this process is much more complex than for simple QA tasks. In this project, we aim to empirically investigate how scaling computations through inference time limits or best-of-N methods affects agent performance on complex tasks.
Theory of Impact
Updated 07/26/26 · By grantmaking.aiUnderstanding how scaling inference computations affects model performance is critical for predicting the performance of future frontier models. Understanding whether there is a log-linear relationship between computation and performance for agents will help us more accurately predict the pace of R&D automation and other risks associated with autonomous systems. Furthermore, most modern EVALs make little use of harnessing, resulting in somewhat underestimated estimates of model performance on real-world tasks, and we are interested in quantifying this bias.
People
Updated 07/26/26 · Edited by orgTeam Member
Funding Details
- May 1, 2026
- Nov 1, 2026
- 4 month
- -
- -
- -
- -
- -
- Seeking first grant
- -
Track Record
You can check our prototype!
And some previous work in animal safety and evaluations.
Discussion
No comments yet. Be the first to share your thoughts.