Functionality
Tasks passing the functional evaluator divided by all tasks in the selected set. No-patch and unresolved runs do not count as successes. Functional + all norms additionally requires every applicable norm to pass.
CODING AGENTS · REPOSITORY-LEVEL EVALUATION
Measure what agents implement—and whether their patches follow the repository’s norms. An evaluation of incremental development, beyond functional correctness alone.
RELEASED BASELINES
Loading verified release data…
| Model / knowledge | Functional success | Overall NCR | Contribution NCR | Prompt-omitted NCR | Functional + all norms | Decision coverage | Details |
|---|
NCR columns: task-macro averages. Coverage: decided / applicable norm instances. Missing and error results are retained. All released entries use mini-swe-agent v2; provider, reasoning and recovery configurations vary—open Details before comparing. No composite score or significance claim.
READ THE SCORES
Tasks passing the functional evaluator divided by all tasks in the selected set. No-patch and unresolved runs do not count as successes. Functional + all norms additionally requires every applicable norm to pass.
For each task, calculate pass / (pass + fail), then average across tasks with a decided norm. Contribution NCR uses the six published contribution categories; Prompt-omitted NCR uses only associations labeled omitted.
error, unavailable and review_required are not converted to fail. They reduce decision coverage. not_applicable is excluded. A coarse proxy pass does not certify documentation quality or test sufficiency.
Contribution categories: authored_tests, docs, changelog, inline_change, style, typing. Every entry covers the full 121-task benchmark. Shared snapshots mean tasks are not fully independent.
ADD YOUR METHOD
Use the corrected default tests and final coarse norm oracle. Submit patches, receipts and execution settings; maintainers validate the evidence before adding a row.
Submission guide ↗