Parsimony

Parsimony

How much code do coding agents write to fix real issues? Research preview.

Parsimony measures published SWE-bench Verified patches in coding units: each statement, call, comparison, name and literal counts as one unit. Formatting, comments and punctuation are not counted. Patches that pass the tests score higher when they are smaller. Failed patches score zero or below.

Leaderboard

Model Score 95% CI Solved Per solve Median churn

Score: mean over all tasks, from −25 to 100. Per solve: mean score over the tasks the model solved, where 50.5 is average. Median churn: units added plus removed per solved task.

Method

Tasks

Enter a task ID. The page opens on the task where solutions differ most in size.


Maintainers' pull request

ModelResultNet unitsChurnTask score

Limits