CODING AGENTS · REPOSITORY-LEVEL EVALUATION

Correct code.
Repository-aware contributions.

Measure what agents implement—and whether their patches follow the repository’s norms. An evaluation of incremental development, beyond functional correctness alone.

Tasks
121
Python repositories / snapshots
6 / 14
Reference norms
160
Task–norm associations
1,024

RELEASED BASELINES

Leaderboard

1.0.0-rc1 · coarse_auto

Loading verified release data…

Functional and coarse norm results. NCR columns are task-macro averages; all metric values are percentages.
Model / knowledgeFunctional
success
Overall
NCR
Contribution
NCR
Prompt-omitted
NCR
Functional +
all norms
Decision
coverage
Details

NCR columns: task-macro averages. Coverage: decided / applicable norm instances. Missing and error results are retained. All released entries use mini-swe-agent v2; provider, reasoning and recovery configurations vary—open Details before comparing. No composite score or significance claim.

Reproducible, version-pinned resultsSummary CSV ↓Task-level CSV ↓Full JSON ↓

READ THE SCORES

Two dimensions. No hidden total score.

01

Functionality

Tasks passing the functional evaluator divided by all tasks in the selected set. No-patch and unresolved runs do not count as successes. Functional + all norms additionally requires every applicable norm to pass.

02

Norm compliance

For each task, calculate pass / (pass + fail), then average across tasks with a decided norm. Contribution NCR uses the six published contribution categories; Prompt-omitted NCR uses only associations labeled omitted.

03

Coverage & limits

error, unavailable and review_required are not converted to fail. They reduce decision coverage. not_applicable is excluded. A coarse proxy pass does not certify documentation quality or test sufficiency.

Contribution categories: authored_tests, docs, changelog, inline_change, style, typing. Every entry covers the full 121-task benchmark. Shared snapshots mean tasks are not fully independent.

Full metric definitions ↗ · Pinned dataset revision ↗

ADD YOUR METHOD

Evaluate locally.
Submit verifiable results.

Use the corrected default tests and final coarse norm oracle. Submit patches, receipts and execution settings; maintainers validate the evidence before adding a row.

Submission guide ↗
  1. Freeze the setup.Record the dataset revision, model, provider, agent, reasoning, budget and recovery policy.
  2. Run and preserve.Keep one declared selected result per task, including failures. Retain patch hashes and both evaluator receipts.
  3. Share for verification.Open a dataset discussion with the result bundle. This site does not execute code or automatically accept self-reported scores.

Run details