Parsimony

Parsimony

How much code do coding agents write to fix real issues? Research preview.

Parsimony measures published SWE-bench Verified patches in coding units: each statement, call, comparison, name and literal counts as one unit. Formatting, comments and punctuation are not counted. Among patches that pass the tests, smaller footprints earn higher credit. Resolve rate is reported separately.

Leaderboard

Model Footprint credit ↑ Resolve rate Measured solves Median churn

Footprint credit: mean per-task credit over measured, in-scope passing patches (1–100, higher means smaller relative to the reference panel; 50.5 is the panel midpoint). Measured solves: scored / published passing attempts in this panel. Median churn: units added plus removed per measured solve.

Models solve different task subsets: this ordering is descriptive, not a same-task comparison or an overall capability ranking. Missing and out-of-scope passing patches are excluded, not assigned zero. No footprint confidence intervals or rank ranges are shown.

All-task score diagnostics (includes failures)

The original score averages all tasks, from −25 to 100. Its intervals and rank ranges describe that score, not footprint credit. Unscored tasks contribute bounds; ranks use their midpoint. Attempts of the same task are resampled together.

ModelAll-task score95% CI

Method

Tasks

Enter a task ID. The page opens on the task where solutions differ most in size.


Maintainers' pull request

ModelResultNet unitsChurnTask score

Limits