Parsimony

Parsimony

How much code do coding agents write to fix real issues? Research preview.

Parsimony measures published SWE-bench Verified patches in coding units: each statement, call, comparison, name and literal counts as one unit. Formatting, comments and punctuation are not counted. Patches that pass the tests score higher when they are smaller. Failed patches score zero or below.

Leaderboard


Model

Rank is by mean net units added, lowest first, across measured in-scope attempts—including failures. The solved column averages only attempts the upstream benchmark marks solved; click it to sort by it instead. Equal means tie; models with no measurements are unranked. Partial coverage can change the order, and a few very large deletions can dominate a mean. This is smallest measured footprint, not best coding ability.

Archived combined-score diagnostics

The previous combined score and its 95% CI remain available as optional graph diagnostics and in the data, but do not determine footprint rank. Its 80% net / 20% units-changed formula is unchanged.

Combined-score CI endpoints are not net-unit uncertainty or inputs to footprint rank.

Net units added means coding units added minus removed, not semantic complexity. Measured shows eligible footprints / published population. Missing or entirely excluded edits are not zero. Solved comes from the upstream benchmark and never affects footprint rank. A failed no-op or large deletion can rank first. Click headings to sort; missing values stay last.

Compare models

Choose any two numeric metrics, including uncertainty endpoints. One point per model, colored by company. Each model line’s newest generation is labelled and fully colored; earlier generations are faded. Archived combined-score diagnostics use the midpoint when bounded; their CI and rank endpoints are labelled explicitly and do not affect footprint ranking. Models missing either selected value have no point. Hover, focus or tap a point to see its model name instantly. Choosing the other axis’s value swaps the axes. Numeric axes fit the visible points with padding; Solved stays at 0–100%.

Method

Limits

Add your own benchmark

Any benchmark, any language: if you know the amount of code used, you can explore it here. No patches or Parsimony analyzer required.

Download the example JSON, replace its benchmark name, unit and results, then open it below. Use total for code used, net for added minus removed, or churn for added plus removed. Units can be lines, tokens, bytes or your own clearly defined measure. Keep different units and measurement methods in separate files.


Your file stays in this browser tab: nothing is uploaded, published or saved. Up to 5 MB / 10,000 results. Imported counts are unverified and never mixed into the published leaderboard.

Want to share it publicly? See the format and contribution guide.

Same answer, less code

Task: does a finite list of numbers contain a negative value? Both solutions return the same answer—including False for an empty list.

More bookkeeping · 6 lines

def has_negative(values):
    found = False
    for value in values:
        if value < 0:
            found = True
    return found

Same result · 2 lines

def has_negative(values):
    return any(value < 0 for value in values)

For [3, -1, 2] both return True; for [0, 2] and [] both return False. The second removes the mutable flag and explicit nested if, using Python’s built-in any.

Illustrative example, not a benchmark result. The shorter version stops at the first match; equivalence here assumes a plain finite list of numbers, not side-effecting iterators. Less bookkeeping can be easier to read, but line count and coding units are footprint measures—not proof of lower semantic or runtime complexity. Both have O(n) worst-case time.