# Submit a RepoNormBench result

This leaderboard is maintained by SYSUSELab. It displays verified result bundles;
the Space does not run agents, accept executable uploads, or automatically publish
self-reported scores. No API key is needed to read the leaderboard.

## 1. Pin the evaluation

Use dataset `SYSUSELab/RepoNormBench` revision
`db1c44ca2e1c67f7bdcc10ae7ff1a9176b75c29f` (1.0.0-rc1), corrected default
functional tests, and the final `coarse_auto` norm oracle. Use the
[published loading/scoring guide](https://huggingface.co/datasets/SYSUSELab/RepoNormBench/blob/db1c44ca2e1c67f7bdcc10ae7ff1a9176b75c29f/QUICKSTART.md)
and its runtime kit. Do not select the historical `message-literal` test variant.

Every leaderboard entry must cover all 121 tasks. Partial-task entries are not accepted.
Retain failed/no-patch/error tasks. Never select a successful attempt after observing
the score without declaring the attempt-selection policy and all attempts.

## 2. Record a minimal configuration

Copy [submission.example.json](submission.example.json) and replace every placeholder.
Use [submission.schema.json](submission.schema.json) for basic metadata validation.
This schema is an intake contract, not a replacement for the functional/norm evaluators.

Record exact model identifier, provider, agent version/commit, reasoning configuration,
temperature, call-budget accounting, retries, workspace recovery, knowledge source and
how the agent receives it. Unavailable historical settings must be explicitly disclosed,
not guessed. Different budgets/providers may be displayed but must not be called a
matched controlled experiment without justification.

For the standard track, keep reference norms, task–norm associations, evaluation tests,
reference solutions, historical patches and labels unavailable to the agent and knowledge
acquisition. The public task statement and source/history at the frozen snapshot are
allowed. Report access controls and contamination limitations. Gold-guided or alternative
protocol runs require a separate declared track, not silent mixing with the standard one.

## 3. Include verifiable artifacts

```text
submission.json                 # completed metadata; no credentials
results.jsonl                   # exactly one declared selected outcome per task
norm_decisions.csv              # task_id,norm_id,status; one per annotated association
patches/                        # nonempty submitted/recovered patches, when available
receipts/functional/            # evaluator outputs + logs, bound to patch/test hashes
receipts/norm/                  # coarse_auto outputs, bound to patch/oracle hashes
SHA256SUMS.json                 # relative path -> SHA-256
```

Each task record should include `task_id`, `functional_success`, `outcome`,
`termination_reason`, `patch_path` (null if unavailable), `patch_sha256`,
`configuration_id`, `functional_receipt`, and `norm_receipt`. Preserve native
evaluator receipts; do not rewrite an evaluator error into a pass or fail. Include
`not_applicable`, `error`, `unavailable`, and `review_required` decisions where returned.
When there is no patch, declare it and retain unavailable norm instances rather than
omitting associations. Provide provenance for any derived unavailable labels.

Include selected/retried attempt IDs and how a recovered workspace patch was obtained.
Functional and norm scoring must refer to the same selected patch. A schema-valid
bundle alone does not establish test success or provenance. Remove credentials,
private paths and unrelated logs before sharing; retain the information needed to
audit outcome selection. Share token counts only with completeness metadata.

## 4. Request inclusion

Open a [dataset discussion](https://huggingface.co/datasets/SYSUSELab/RepoNormBench/discussions)
titled `Leaderboard submission: <model> / <method>`. Link the artifact bundle and
completed metadata; do not paste credentials. Large bundles can be shared through
a public, versioned HF/Git repository with an explicit license and file hashes.

Maintainers check task/association coverage, duplicate IDs, patch and evaluator
bindings, configuration disclosures and current-test compatibility, then recompute
metrics. Incoming scores remain unverified until this review is complete. Any
execution needed for verification occurs in a controlled offline evaluation
environment, never in this public Space. Do not run untrusted patches on a shared
host without appropriate isolation.

