# Leaderboard metrics and comparison boundaries

Dataset: `SYSUSELab/RepoNormBench`, **1.0.0-rc1**, revision
`db1c44ca2e1c67f7bdcc10ae7ff1a9176b75c29f`. Norm profile: **coarse_auto**.

All rates in the JSON and CSV are fractions in [0, 1]; the page displays percentages.
An undefined rate is `null`, not zero. Every leaderboard entry covers all 121 tasks.

## Definitions

For a task t and a selected scope, let P(t), F(t) be the counts of norm `pass`, `fail`.
A task contributes to an NCR macro average only when P(t) + F(t) > 0.

| Metric | Calculation |
| --- | --- |
| Functional success | Functional passes / all selected tasks |
| Overall NCR (task macro) | Mean of P(t)/(P(t)+F(t)) across tasks with a decided instance |
| Contribution NCR (task macro) | Same formula, restricted to the six published contribution categories |
| Prompt-omitted NCR (task macro) | Same formula, restricted to associations explicitly labeled `omitted` |
| Instance NCR (micro, details) | Sum of P(t) / sum of P(t)+F(t) |
| Functional-subset NCR (micro, details) | Same micro formula, restricted to functionally passed tasks using current functional labels |
| Decision coverage | All `pass` + `fail` instances / all applicable instances |
| Functional + all norms | Tasks with functional success and every applicable instance explicitly `pass` / all selected tasks; at least one applicable instance required |

Contribution categories are the existing reproduction mapping:
`authored_tests`, `docs`, `changelog`, `inline_change`, `style`, `typing`.
The mapping uses `norm_kind` in the original decision CSVs, not a newly inferred
classification or the exported norm's `oracle_kind`. Prompt-omitted uses the
existing association labels. `implicit` and `partial` are **not** silently merged
into `omitted`. Scope-specific task counts, micro rates and coverage are in row details.

`not_applicable` is excluded from applicable denominators. `error`, `unavailable`
and `review_required` are applicable but undecided: they reduce coverage and prevent
a joint pass. They are not converted to `fail`. A no-patch result cannot count as a
functional pass. Infrastructure outcomes remain in source receipts; this dashboard
does not reinterpret them as successful runs or drop them from the task total.

## Corrected functional labels

The leaderboard consumes **data/results.jsonl**, not the historical functional labels
embedded in the old norm CSVs. `JIN1559-F07` DeepSeek + CodeWiki changes from fail
to pass after exact-patch rescoring on the corrected default tests. Functional success
is therefore **82/121 (67.77%)**, up from 81/121 (66.94%). Its norm decisions do not
change; functional-subset metrics are rejoined and recalculated. Joint success
remains 5/121 because this patch does not pass all applicable norms.

Evidence:
[exact-patch rescore](https://huggingface.co/datasets/SYSUSELab/RepoNormBench/blob/db1c44ca2e1c67f7bdcc10ae7ff1a9176b75c29f/evidence/deepseek/RESULTS.json).
The build verifies patch/test hashes and fails if the old label or count reappears.
All six historical message-literal bindings have corrected-test evidence in the
dataset; other results retain their recorded default-test evaluations. This is not
a claim that all agents or all functional tests were rerun for this website.

## Interpretation

- There is no weighted total score. Sorting a column is a viewing convenience,
  not a significance result or a claim of statistically established superiority.
- Compare methods within the same model and inspect configurations. Providers,
  reasoning settings, recovery policies and infrastructure retries are recorded,
  and are not identical across all historical executions.
- Report coverage alongside NCR. High NCR on fewer decided tasks is not necessarily
  better compliance. Included-task counts are shown instead of filling missing
  values with invented scores.
- Coarse norm checks are automated proxies. A changed documentation file does not
  establish good documentation; an authored-test pass does not prove sufficient tests.
- Public reference material must not be supplied as agent input in the standard
  track. Declare any benchmark-derived knowledge or alternative evaluation protocol.
- Shared snapshots and public tests/solutions limit independence and contamination
  control. Historical significance tests are not presented as newly recomputed ones.
- Recorded token usage is incomplete for some requests. This first leaderboard
  does not rank monetary costs or imply that recorded sums are complete bills.

