Parsimony

Parsimony

How much code do coding agents write to fix real issues? Research preview.

Parsimony measures published SWE-bench Verified patches in coding units: each statement, call, comparison, name and literal counts as one unit. Formatting, comments and punctuation are not counted. Patches that pass the tests score higher when they are smaller. Failed patches score zero or below.

Leaderboard

Model

Score: mean over all tasks, from −25 to 100. Per solve: mean score over the tasks the model solved, where 50.5 is average. Median churn: units added plus removed per solved task.

Click a column heading to sort; click again to reverse. Score uses the midpoint when shown as bounds; 95% CI sorts by its lower bound. Missing values stay last. Per solve and median churn describe successful patches only; sorting them does not change the all-task score.

95% CI shows score uncertainty from resampling tasks, widened for missing measurements. It helps avoid overinterpreting small score gaps; it is not a measure of code complexity and does not cover every source of bias.

Method

Tasks

Enter a task ID. The page opens on the task where solutions differ most in size.


Maintainers' pull request

ModelResultNet unitsChurnTask score

Limits