News
Separating signal from noise in coding evaluations
OpenAI says a new analysis points to reliability and accuracy issues in SWE-Bench Pro, a widely used coding benchmark. The findings raise questions about how well current evaluations measure real model performance and whether benchmark results are as stable as they appear.
OpenAI has conducted a detailed audit of SWE-Bench Pro, a coding benchmark designed to evaluate agentic software development capabilities through realistic tasks from public and private repositories. The investigation reveals that approximately 30% of the benchmark's tasks are broken, meaning they do not accurately measure model performance.
The audit utilized a datapoint analysis pipeline combined with human-supervised agent reviews and a campaign involving experienced software engineers. Specifically, the automated pipeline flagged 27.4% of tasks as problematic, while the human annotation process identified 34.1% of the dataset as flawed.
Identified issues fall into four primary categories: tests that are overly strict regarding implementation details, prompts that omit necessary requirements, tests with insufficient coverage that allow incomplete fixes to pass, and prompts that are misleading or contradictory.
OpenAI warns model developers that these flaws can misrepresent a model's actual capabilities and affect research priorities. The findings highlight the challenges of creating fair, high-difficulty benchmarks and suggest that AI agents are increasingly useful for scalable data quality assurance.
Readers’ Forum
No contributions yet — open the debate.
Open the discussion
No account or password needed — just enter your e-mail and we’ll send you a one-time sign-in link. First time here? You’re set up automatically.
Your rating will be applied automatically after you sign in.
Check your inbox
We’ve sent a sign-in link to …. Open it on this device — this tab will sign you in automatically.
Waiting for your click …
·