Parsimony

Parsimony

How much code do coding agents write to fix real issues? Research preview.

Parsimony measures published SWE-bench Verified patches in coding units: each statement, call, comparison, name and literal counts as one unit. Formatting, comments and punctuation are not counted. Patches that pass the tests score higher when they are smaller. Failed patches score zero or below.

Leaderboard


Model

Rank is by mean net units added, lowest first, across measured in-scope attempts—including failures. Equal means tie; missing means are unranked. Partial coverage can change the order. This is smallest measured footprint, not best coding ability.

Archived combined-score diagnostics

The previous combined score and its 95% CI remain available as optional graph diagnostics and in the data, but do not determine footprint rank. Its 80% net / 20% units-changed formula is unchanged.

Combined-score CI endpoints are not net-unit uncertainty or inputs to footprint rank.

Net units added means coding units added minus removed, not semantic complexity. Measured shows eligible footprints / published population. Missing or entirely excluded edits are not zero. Solved comes from the upstream benchmark and never affects footprint rank. A failed no-op or large deletion can rank first. Click headings to sort; missing values stay last.

Compare models

Choose any two numeric metrics, including uncertainty endpoints. One point per model, colored by company. Archived combined-score diagnostics use the midpoint when bounded; their CI and rank endpoints are labelled explicitly and do not affect footprint ranking. Models missing either selected value have no point. Hover, focus or tap a point to see its model name instantly. Choosing the other axis’s value swaps the axes. Numeric axes fit the visible points with padding; Solved stays at 0–100%.

Method

Limits