Parsimony
How much code do coding agents write to fix real issues? Research preview.
Parsimony measures published SWE-bench Verified patches in coding units: each statement, call, comparison, name and literal counts as one unit. Formatting, comments and punctuation are not counted. Patches that pass the tests score higher when they are smaller. Failed patches score zero or below.
Leaderboard
Population coverage
| Rank | Model |
|---|
Rank is by mean net units added, lowest first, across measured in-scope attempts—including failures. Equal means tie; missing means are unranked. Partial coverage can change the order. This is smallest measured footprint, not best coding ability.
Archived combined-score diagnostics
The previous combined score and its 95% CI remain available as optional graph diagnostics and in the data, but do not determine footprint rank. Its 80% net / 20% units-changed formula is unchanged.
Combined-score CI endpoints are not net-unit uncertainty or inputs to footprint rank.
Net units added means coding units added minus removed, not semantic complexity. Measured shows eligible footprints / published population. Missing or entirely excluded edits are not zero. Solved comes from the upstream benchmark and never affects footprint rank. A failed no-op or large deletion can rank first. Click headings to sort; missing values stay last.
Compare models
Choose any two numeric metrics, including uncertainty endpoints. One point per model, colored by company. Archived combined-score diagnostics use the midpoint when bounded; their CI and rank endpoints are labelled explicitly and do not affect footprint ranking. Models missing either selected value have no point. Hover, focus or tap a point to see its model name instantly. Choosing the other axis’s value swaps the axes. Numeric axes fit the visible points with padding; Solved stays at 0–100%.
Each task was run four times; every measured attempt contributes to its model’s footprint mean.
Method
- Each patch is applied to the task's base commit and both versions of each changed file are parsed. Nothing is executed.
- Footprint rank uses the arithmetic mean of net coding units added across all measured, in-scope attempts. Successful and failed attempts both count; passing references are not required.
- Correctness is reported separately by the upstream benchmark. Smaller failed or no-op attempts are not better solutions.
- Test, docs and generated files are excluded. Only Python is measured.
Limits
- Smaller is not always better. Units measure size, not readability or design.
- Scores are relative to the reference models, which all use the same agent harness (mini-SWE-agent).
- Pass/fail comes from SWE-bench's published results; tests are not rerun.