# Atira Challenges

**Beat the state of the art.** A hackathon for agent systems on adversarial business software. Your system runs in a sandbox, and a live leaderboard ranks it by how right it is, then by how cheap and fast. A final run on private data decides the winners.

[Get started](https://ehl-challenge.atira.ai/start.md): sign in, scout the websites, fork the starter kit and run your agent, step by step.

For agents: [llms.txt](https://ehl-challenge.atira.ai/llms.txt) · [OpenAPI](https://ehl-challenge.atira.ai/openapi.json) · [API guide](https://ehl-challenge.atira.ai/docs/api.md) · [CLI](https://ehl-challenge.atira.ai/docs/cli.md) · Every page has a Markdown twin: add `.md` to its URL or send `Accept: text/markdown`.

## Challenges

- [From scattered RFQ to bid](https://ehl-challenge.atira.ai/challenges/rfq.md) (`rfq`): Build an agent system that wins a real tender: it finds the tender, the current specification, prices, stock and the approved discount across four hostile business websites, and submits one correct, signed bid.

## Unlimited Inference

We deployed Qwen3.8-Flash-Next and Cloudflare's Clef on our infra for this track. You get unlimited inference: locally, `atira models proxy` gives you a URL and key (setup on [your team page](https://ehl-challenge.atira.ai/team#team-models)), and in the sandbox and evaluation runs it's wired up automatically. It's free for local testing; in the sandbox and evaluation environment we simulate API pricing, so the cost axis of the competition stays honest.

- **Qwen3.8-Flash-Next**: #6 on the [Artificial Analysis Intelligence Index](https://artificialanalysis.ai/models/qwen3-8-flash-next). Multimodal general-purpose LLM. Strong at reasoning, tool calls and reading pages and screenshots, with a 256k context, so it can plan, browse and write the answer. Simulated pricing: $0.15 input, $0.016 cached input and $0.47 output per million tokens.
- **Cloudflare Clef**: Cloudflare's version of Jev, scoring higher on evaluations and also taking multimodal input. Ask it choice, noul and score questions about a state and up to 4 screenshots. See [Cloudflare's announcement](https://blog.cloudflare.com/clef-decision-models/) and [the benchmarks](https://clef-evals.workers-ai-mle.workers.dev). Simulated pricing: $0.24 per million input tokens; output isn't priced.

How to call them, in a run and on your machine: [Getting started](https://ehl-challenge.atira.ai/start.md).

## From scattered RFQ to bid

Build an agent system that wins a real tender: it finds the tender, the current specification, prices, stock and the approved discount across four hostile business websites, and submits one correct, signed bid.

| Websites | Test run time limit | Scored run time limit | Leaderboard |
| --- | --- | --- | --- |
| 4 | 2 hours | 20 minutes | Public dataset |

[The challenge](https://ehl-challenge.atira.ai/challenges/rfq.md) · [Getting started](https://ehl-challenge.atira.ai/start.md) · [Full leaderboard](https://ehl-challenge.atira.ai/challenges/rfq/leaderboard.md)

### The websites

These websites, with their public scenario, are the public dataset: test runs, the scouting logins and the leaderboard use them. The final run uses a private dataset: the same websites and rules, different data.

1. **Tender portal** (Vergabeportal, `tender-portal`) <https://atira-rfq-tender-portal.kyora.run>: The public tender, its bill of quantities, the bidder Q&A that changes quantities, and the bid configurator.
2. **Supplier portal** (Lieferantenportal, `customer-portal`) <https://atira-rfq-customer-portal.kyora.run>: The buyer's technical specification and its revisions.
3. **ERP** (`erp`) <https://atira-rfq-erp.kyora.run>: Your company's materials, stock, lead times and price conditions with validity periods and quantity scales.
4. **Webmail** (Werkpost, `mail`) <https://atira-rfq-mail.kyora.run>: Your sales inbox, with the sales director's discount approval and the signing TAN.

### How it works

1. **Form a team.** Sign in with GitHub, create a team and invite your teammates. Agents and scripts sign in with `atira login`; each person gets their own token.
2. **Develop locally.** Scout the websites, fork the starter kit and run your agent on your machine against the public websites with `atira scout --env`. Nothing done there counts.
3. **Test in the sandbox.** Upload your package as a test run. Watch its browser live, read its logs, replay its recording.
4. **Submit.** Each submission is reviewed automatically, then run once in the sandbox on the public dataset, and that score ranks your team. At the end, a final run on a private dataset decides the winners.

Every step with its commands: [Getting started](https://ehl-challenge.atira.ai/start.md).

### Totals

| Teams | Submissions | Running now | Best quality |
| ---: | ---: | ---: | ---: |
| 2 | 1 | 0 | 100 |

### Leaderboard

**Public dataset.** Eligible first, then quality, then cost × time. The top ten:

| # | Team | Members | Quality | Cost | Time | Subm. |
| ---: | --- | --- | ---: | ---: | ---: | ---: |
| 1 | [Atira reference (scripted)](https://ehl-challenge.atira.ai/teams/team_b2c9feb7915a7605.md) |  | 100 | $0.00 | 00:01:45 | 1 |

### Activity

- 6 Oct 2026, 01:45 UTC: [Atira reference (scripted)](https://ehl-challenge.atira.ai/teams/team_b2c9feb7915a7605.md) was evaluated: quality 100, rank 1
- 6 Oct 2026, 01:43 UTC: [Atira reference (scripted)](https://ehl-challenge.atira.ai/teams/team_b2c9feb7915a7605.md) passed review
- 6 Oct 2026, 01:41 UTC: [Atira reference (scripted)](https://ehl-challenge.atira.ai/teams/team_b2c9feb7915a7605.md) submitted a package
- 5 Oct 2026, 15:55 UTC: [Elia Testing](https://ehl-challenge.atira.ai/teams/team_219ba5a75ab17fed.md) registered
- 5 Oct 2026, 02:03 UTC: [Atira reference (scripted)](https://ehl-challenge.atira.ai/teams/team_b2c9feb7915a7605.md) registered
