OpenAI has detailed how two inference-time configuration changes tripled its ARC-AGI-3 benchmark scores without modifying the underlying model.
The post explains that carefully tuning reasoning behavior and evaluation settings enabled substantially better performance on the interactive reasoning benchmark, highlighting how model configuration can be as important as model size for difficult agentic tasks.
OpenAI argues that benchmark results should be interpreted alongside the inference setup, since small changes in reasoning parameters can produce large differences in performance. The findings reinforce the importance of standardized evaluations and transparent reporting as frontier AI systems become increasingly configurable.

