More bookkeeping · 6 lines
def has_negative(values):
found = False
for value in values:
if value < 0:
found = True
return foundHow much code do coding agents write to fix real issues? Research preview.
Coding units count syntax elements—statements, calls, comparisons, names and literals—not lines or runtime complexity. Net units = added − removed.
We parse published SWE-bench Verified patches and count syntax elements plus block-end markers. Formatting, comments and punctuation do not count. Python, JavaScript/TypeScript and Go use separate parser-specific tracks; raw unit counts are not comparable across them.
Rank is by mean net units added, lowest first, across measured in-scope attempts—including failures. The solved column averages only attempts the upstream benchmark marks solved; click it to sort by it instead. Equal means tie; models with no measurements are unranked. Partial coverage can change the order, and a few very large deletions can dominate a mean. This is smallest measured footprint, not best coding ability.
Net units added means coding units added minus removed, not semantic complexity. Measured shows eligible footprints / published population. Missing or entirely excluded edits are not zero. Solved comes from the upstream benchmark and never affects footprint rank. A failed no-op or large deletion can rank first. Click headings to sort; missing values stay last.
Archived combined-score diagnostics use the midpoint when bounded; their CI and rank endpoints are labelled explicitly and do not affect footprint ranking. Models missing either selected value have no point. Choosing the other axis’s value swaps the axes. Numeric axes fit the visible points with padding; percentage axes normally stay at 0–100%. Points outside fitted axes remain at their actual, true-scale positions.
Solved % = upstream resolved attempts / frozen attempt population × 100. Tests are not rerun by Parsimony. Footprint = sum of net coding units / N measured in-scope attempts: each attempt has equal weight 1/N, including failures. Missing measurements are not zero; solved-only footprint is separate. Benchmarks and languages are not pooled and have no benchmark weights. Default footprint uses no correctness weight.
Archived score only: passing results blend 80% net / 20% churn per-task percentiles; failed results use the same 80/20 weights in a bounded growth/churn penalty. It never affects footprint rank. Full formula.
Combined-score CI endpoints are not net-unit uncertainty or inputs to footprint rank.
Solved %: upstream pass rate. Footprint: mean net coding units, not a %. No correctness weighting.
Hover, focus or tap for values; all models are labelled.
Each task was run four times; every measured attempt contributes to its model’s footprint.
| Rank | Model |
|---|
Mean net units added, including failures. Smaller footprint ≠ better coding.
Separate boards—not ingredients in a combined ranking.
Bars show frozen task counts relative to the largest source, not percentages or weights. Counts describe planned populations, not measurement eligibility.
Each leaderboard uses only its own benchmark and language: 100% from that board, 0% from the others. These are source shares, not combined-score weights.
Any benchmark, any language: if you know the amount of code used, you can explore it here. No patches or Parsimony analyzer required.
Download the example JSON, replace its benchmark name, unit and results, then open it below. Use total for code used, net for added minus removed, or churn for added plus removed. Units can be lines, tokens, bytes or your own clearly defined measure. Keep different units and measurement methods in separate files.
Your file stays in this browser tab: nothing is uploaded, published or saved. Up to 5 MB / 10,000 results. Imported counts are unverified and never mixed into the published leaderboard.
Want to share it publicly? See the format and contribution guide.
Task: does a finite list of numbers contain a negative value? Both solutions return the same answer—including False for an empty list.
def has_negative(values):
found = False
for value in values:
if value < 0:
found = True
return founddef has_negative(values):
return any(value < 0 for value in values)For [3, -1, 2] both return True; for [0, 2] and [] both return False. The second removes the mutable flag and explicit nested if, using Python’s built-in any.
Illustrative example, not a benchmark result. The shorter version stops at the first match; equivalence here assumes a plain finite list of numbers, not side-effecting iterators. Less bookkeeping can be easier to read, but line count and coding units are footprint measures—not proof of lower semantic or runtime complexity. Both have O(n) worst-case time.