How enabling two settings tripled our scores on the ARC-AGI-3 benchmark | OpenAI
Developers can improve AI efficiency and accuracy by using production settings rather than generic harnesses. It warns that some benchmarks may underrepresent a model's actual capabilities.
- GPT-5.6 Sol scored 13.3% with the official generic harness
- Score increased to 38.3% on the public task set when using Responses API harness with retained reasoning and compaction
- Output token usage was reduced by 6x
- GPT-5.5 scored 0.4% on the benchmark