Parsimony
How much code do coding agents write to fix real issues? Research preview.
Parsimony measures published SWE-bench Verified patches in coding units: each statement, call, comparison, name and literal counts as one unit. Formatting, comments and punctuation are not counted. Patches that pass the tests score higher when they are smaller. Failed patches score zero or below.
Leaderboard
Population coverage
| Rank | Model |
|---|
All-task score rank (unchanged when sorting other columns): the range covering 95% of each model's ranks when the tasks are resampled 2,000 times; overlapping ranges mean the order is uncertain (sensitivity report).
Score: mean over all score-panel attempts, from −25 to 100; higher is better. Passing patches earn 1–100 credit relative to reference patches (80% net units added, 20% units changed). Failed patches receive a footprint penalty from −25 to 0. Missing measurements contribute bounds, not zero. This combines correctness and footprint; it is not a raw complexity count. The 95% CI is computed separately and never changes Score.
Net units added: coding units added minus removed—our proxy for net added complexity, not semantic complexity. In the all-task view this is the mean across measured, in-scope passing and failed attempts; missing measurements are excluded, never treated as zero. For a selected task it is that attempt’s net change. Click headings to sort; click again to reverse. Score bounds sort by midpoint; 95% CI sorts by its lower bound. Missing values stay last.
95% CI is an uncertainty display, not an input to Score. It comes from resampling tasks, widened for missing measurements. It helps avoid overinterpreting small score gaps; it is not a measure of code complexity and does not cover every source of bias.
Compare models
Choose any two numeric table values. One point per model, colored by company. Score uses the midpoint when bounded; interval and rank endpoints are labeled explicitly. Models missing either selected value have no point. Hover, focus or tap a point to see its model name instantly. Choosing the other axis’s value swaps the axes. Numeric axes fit the visible points with padding; Solved stays at 0–100%.
Each task was run four times; each attempt is scored against every passing patch of the task.
Method
- Each patch is applied to the task's base commit and both versions of each changed file are parsed. Nothing is executed.
- On each task, a passing patch is ranked against all passing patches from the reference models: 80% on net units added, 20% on units changed.
- A failed patch receives a footprint penalty from −25 to 0. The scoring formula is unchanged by which columns or graph axes are displayed.
- Test, docs and generated files are excluded. Only Python is measured.
Limits
- Smaller is not always better. Units measure size, not readability or design.
- Scores are relative to the reference models, which all use the same agent harness (mini-SWE-agent).
- Pass/fail comes from SWE-bench's published results; tests are not rerun.