# Answers and scoring

> In this document, `$API` is https://ehl-challenge.atira.ai (the CLI and API guide call it `$ATIRA_API`).

Each challenge defines its answer format and how quality is scored; the platform measures cost and time and
ranks everyone the same way. Test runs, scored submissions and the scouting logins use the public dataset (the
websites as you see them, with scenario `public-1`), and the leaderboard ranks those results. The correct answer
isn't published; your system has to work it out from the websites. When the hackathon ends, a [final run](#final-run)
on a private dataset decides the winners.

## Per run

- **Quality** (0 to 100): how right the answer is, by the challenge's scorer.
- **Eligible:** whether the answer meets the challenge's mandatory criteria. For RFQ: the right decision (bid
  or no_bid) about the right tender reference.
- **Cost:** the run's LLM usage, priced per model, in USD.
- **Time:** from the start of your run command to your accepted answer.

A malformed answer is rejected with a `400` and does not end the run. A well-formed answer is final, even if it
is wrong. A run without an answer scores quality 0 and is not eligible.

## Leaderboard

Each scored submission runs once in the sandbox on the public dataset and is scored; one run per submission
keeps your model bill down. A submission appears once its run has finished; each team is ranked by its best
complete submission. The API reports the values in its `median_*` fields, which organizers can also configure
to take the median of several runs.

1. Eligible before ineligible.
2. Higher quality.
3. Lower cost × time: the submission's cost in USD times its time, with cost counted as at least $0.01 and time
   as at least 1 second. Cost and time weigh the same: halving your model bill counts as much as halving your run
   time, and among free runs the faster one wins.
4. The submission ID, when everything above is equal.

On the leaderboard's plot, each team is a dot, quality against cost or time; the dotted line is the Pareto front,
the teams no other eligible team beats on both.

Test runs on the public scenario `public-1` show the points for each part of the answer, never the expected
values. Scored runs show quality, eligibility, cost and time without a breakdown.

## Final run

When the hackathon ends, the organizers run every team's latest submission that passed review (passed or
flagged; rejected ones don't count) once on a private dataset: the same websites and rules with different tenders,
specifications, materials, prices and approvals. That run decides the winners, so build a system that works the
answer out from the websites, not one tuned to the public scenario. It ranks in the same order as above.

It is the latest submission, not the best one: your best submission holds your place on the leaderboard; your
latest one that passed review goes into the final run. The final run's results stay with the organizers until they
publish them. Then the leaderboard shows the final results, tagged "Private dataset · final", with each team's
public-dataset result and its change in rank beside them, and the public-dataset board stays at `$API/challenges/rfq/leaderboard?dataset=public` (Markdown:
`leaderboard.md?dataset=public`; API: `GET /v1/challenges/rfq/leaderboard?dataset=public`).

## RFQ scoring

| Part | Points |
| --- | --- |
| Tender reference | 5 |
| Discount percent | 5 |
| Total net (to the cent, consistent with the positions) | 15 |
| Positions (split evenly): material 35 %, quantity 20 %, unit price 30 %, lead time 15 % | 75 |

Money is compared to the cent. Positions are matched by their position number; extra or duplicate positions cost
one position's share each. For a no_bid scenario, full marks need the right decision and reference, no positions
and a total of 0.
