# Fair play

Your agent system must work the answer out from the websites at run time. The platform measures its quality, cost
and time.

## Allowed

- Any agent design; the free hosted models, or any provider and model with your own key, through the proxy; any
  amount of tokens.
  The proxy forwards your calls unchanged and measures them: reported usage is metered at the published model
  prices, and an upstream-reported USD cost takes precedence. See [the run contract](https://ehl-challenge.atira.ai/docs/run-contract.md#llm-calls)
  for SDK configuration.
- Keeping your model key as a [team secret](https://ehl-challenge.atira.ai/docs/run-contract.md#team-secrets) rather than in your package.
- Your own code, prompts and tools; open-source libraries installed during setup.
- Using what you learned with the scouting logins and in test runs about how the websites work.
- Up to 4 browsers at a time: ask for them with `"browsers"` in `agent.json` and use as many tabs and browser
  contexts inside them as you like.

## Not allowed

- Hardcoding scenario data or answers (tender references, material numbers, prices, totals, discounts).
- Bundling or running a local model, or reaching any model other than through the proxy.
- Contacting hosts other than the API and the challenge's websites, or trying to get data out of the sandbox.
- Starting browsers of your own in the sandbox, headless ones included. At most 4 browsers run at a time: ask for up
  to 4 in `agent.json` (`"browsers"`), use tabs and browser contexts inside them freely, and don't start others.
  On your own machine (developing with the scouting logins) your agent starts its own browsers; that is up to you.
- Attacking the platform, the websites, the sandbox or other teams; sharing scored-run data.

## How it is checked

- **Package review before any scored run:** the same static checks test runs get (size, binaries, model weights,
  encoded blobs, values of the public scenario's answer in the package) and an AI code reviewer, a coding agent
  that reads your whole package in an isolated VM. Outcomes: passed, flagged (runs anyway, a person looks at it) or
  rejected (no runs). Text in your package addressed to the reviewer counts against you.
- **Hardcoded answers:** a package that contains a value of the public scenario's answer is rejected. The finding
  says "The package contains a value of the public scenario's answer..." and never names the value or the file.
- **The final run** uses a private dataset: the same websites and rules, different data. That run weighs most
  when we pick the winners, so build a system that works the answer out from the websites, not one tuned to the public scenario.
- **The sandbox** reaches only our API and the websites, so every model call passes the proxy and is measured.
  Your team secrets reach only the timed run step and are redacted from its logs.
- **Website logins** exist only for a live run and stop working when it ends.
- **Browsers:** the sandbox stops any browser your agent starts beyond the ones it asked for and notes it in the
  run's logs: `[platform] stopped a browser the agent started`.
- **Everything measured** (tokens, cost, time, the answer) is recorded by the platform, never reported by you.

Breaking these rules gets the submission removed from the leaderboard, and repeated or deliberate attempts get
the team disqualified.
