OpenAI has conducted a detailed audit of SWE-Bench Pro, a coding benchmark designed to evaluate agentic software development capabilities through realistic tasks from public and private repositories. The investigation reveals that approximately 30% of the benchmark's tasks are broken, meaning they do not accurately measure model performance.
The audit utilized a datapoint analysis pipeline combined with human-supervised agent reviews and a campaign involving experienced software engineers. Specifically, the automated pipeline flagged 27.4% of tasks as problematic, while the human annotation process identified 34.1% of the dataset as flawed.
Identified issues fall into four primary categories: tests that are overly strict regarding implementation details, prompts that omit necessary requirements, tests with insufficient coverage that allow incomplete fixes to pass, and prompts that are misleading or contradictory.
OpenAI warns model developers that these flaws can misrepresent a model's actual capabilities and affect research priorities. The findings highlight the challenges of creating fair, high-difficulty benchmarks and suggest that AI agents are increasingly useful for scalable data quality assurance.
Comments
No comments yet — be the first.
Open the discussion
No account or password needed — just enter your e-mail and we’ll send you a one-time sign-in link. First time here? You’re set up automatically.
Your rating will be applied automatically after you sign in.
Check your inbox
We’ve sent a sign-in link to …. Open it on this device — this tab will sign you in automatically.
Waiting for your click …
·