More bookkeeping · 6 lines
def has_negative(values):
found = False
for value in values:
if value < 0:
found = True
return foundHow much code do coding agents write to fix real issues? Research preview.
Coding units count syntax elements—statements, calls, comparisons, names and literals—not lines or runtime complexity. Net units = added − removed.
We parse published SWE-bench Verified patches and count syntax elements plus block-end markers. Formatting, comments and punctuation do not count. Python, JavaScript/TypeScript and Go use separate parser-specific tracks; raw unit counts are not comparable across them.
Rank is by mean net units added, lowest first, across measured in-scope attempts—including failures. The solved column averages only attempts the upstream benchmark marks solved; click it to sort by it instead. Equal means tie; models with no measurements are unranked. Partial coverage can change the order, and a few very large deletions can dominate a mean. This is smallest measured footprint, not best coding ability.
Net units added means coding units added minus removed, not semantic complexity. Measured shows eligible footprints / published population. Missing or entirely excluded edits are not zero. Solved comes from the upstream benchmark and never affects footprint rank. A failed no-op or large deletion can rank first. Click headings to sort; missing values stay last.
Archived combined-score diagnostics use the midpoint when bounded; their CI and rank endpoints are labelled explicitly and do not affect footprint ranking. Models missing either selected value have no point. Choosing the other axis’s value swaps the axes. Numeric axes fit the visible points with padding; percentage axes normally stay at 0–100%. Points outside fitted axes remain at their actual, true-scale positions.
Solved % = upstream resolved attempts / frozen attempt population × 100. Tests are not rerun by Parsimony. Footprint = sum of net coding units / N measured in-scope attempts: each attempt has equal weight 1/N, including failures. Missing measurements are not zero; solved-only footprint is separate. Benchmarks and languages are not pooled and have no benchmark weights. Default footprint uses no correctness weight.
Archived score only: passing results blend 80% net / 20% churn per-task percentiles; failed results use the same 80/20 weights in a bounded growth/churn penalty. It never affects footprint rank. Full formula.
Artificial Analysis indices are external, model-level capability scores from a pinned snapshot, not results on this benchmark. They are matched to a configuration only for the same model and reasoning effort, or where Artificial Analysis lists a single unlabelled entry or resolves an undated model ID; these weaker matches are disclosed on hover. Configurations without a match have no value and no point. They never affect footprint rank.
Self-run models were run by the site owner on the first 3–10 DeepSWE Python tasks (OpenAI models via pi, Anthropic models via Claude Code), without the task containers. They are unverified, not graded and never ranked, and cover only a task subset, so compare them mainly with each other. GPT-5.6 Sol is the one model run both ways: its self-run footprint was 0.88× its published one on the same tasks. Claude Code has no such check.
Combined-score CI endpoints are not net-unit uncertainty or inputs to footprint rank.
| Rank | Model |
|---|
Mean net units added, including failures. Smaller footprint ≠ better coding.
Separate boards, never pooled. Bars show task counts, not weights.
Each leaderboard uses only its own benchmark and language: 100% from that board, 0% from the others. These are source shares, not combined-score weights. The self-run pilot is the only exception: its points are shown faded on the DeepSWE Python board, outside the ranking.
Any benchmark, any language: if you know the amount of code used, you can explore it here. No patches or Parsimony analyzer required.
Download the example JSON, replace its benchmark name, unit and results, then open it below. Use total for code used, net for added minus removed, or churn for added plus removed. Units can be lines, tokens, bytes or your own clearly defined measure. Keep different units and measurement methods in separate files.
Your file stays in this browser tab: nothing is uploaded, published or saved. Up to 5 MB / 10,000 results. Imported counts are unverified and never mixed into the published leaderboard.
Want to share it publicly? See the format and contribution guide.
Task: does a finite list of numbers contain a negative value? Both solutions return the same answer—including False for an empty list.
def has_negative(values):
found = False
for value in values:
if value < 0:
found = True
return founddef has_negative(values):
return any(value < 0 for value in values)For [3, -1, 2] both return True; for [0, 2] and [] both return False. The second removes the mutable flag and explicit nested if, using Python’s built-in any.
Illustrative example, not a benchmark result. The shorter version stops at the first match; equivalence here assumes a plain finite list of numbers, not side-effecting iterators. Less bookkeeping can be easier to read, but line count and coding units are footprint measures—not proof of lower semantic or runtime complexity. Both have O(n) worst-case time.