From sign-in to your first scored run
Everything between signing in and the leaderboard: the commands to run, logins to scout the websites, the rules of the game and how you're scored.
What you're building
Your first run in five minutes
-
Step 1: Sign in and form a team
Sign in with GitHub and create a team of up to 6. Invite teammates by their GitHub username on your team page; each invite works only for that account.
-
Step 2: Install the CLI and sign in
The CLI is one file for Node 24.
atira loginopens the browser, where you approve a token for this computer. Or skip the install: point a coding agent at https://ehl-challenge.atira.ai/llms.txt.curl -fsSL https://ehl-challenge.atira.ai/install.sh | sh atira login -
Step 3: Scout the websites yourself
Permanent logins for every website, on the public dataset: the same sites your agent will face. Explore freely; nothing done there counts.
Website URL Username Password Tender portalVergabeportal atira-rfq-tender-portal.kyora.run u_20115c791c10c92cDrqJ-mbv75jbQW-NXXZOSupplier portalLieferantenportal atira-rfq-customer-portal.kyora.run u_20115c791c10c92cTA8rOLDy3r3UJLYM2shLERP atira-rfq-erp.kyora.run u_20115c791c10c92cRa7hKrCMINT5wjygR0X9WebmailWerkpost atira-rfq-mail.kyora.run u_20115c791c10c92c5fq1jvwSqyc3yPXXPIJW -
Step 4: Grab the starter kit
A working Python agent to fork:
agent.pygives Claude browser tools on the sandbox's Chromium, the helper modules handle page snapshots, dialogs, downloads and German numbers,bid.pychecks the bid before it goes in, andagent.jsontells the sandbox how to install and run it.mkdir my-agent curl -fsSL https://ehl-challenge.atira.ai/starter-kit.tar.gz | tar xz -C my-agent -
Step 5: Develop locally
atira scout --envexports the scouting logins the way a run gets its logins, so your agent runs on your machine against the public websites, with a local browser and your own model key. There is no run token and nothing done there counts. With these logins, an answer entered on the tender portal is acknowledged but never counts. When it works,atira test-runruns it in the sandbox for up to 2 hours and opens the run live, with the score and full breakdown. Details: how a run works.cd my-agent && uv sync && uv run playwright install chromium eval "$(atira scout --env)" export ANTHROPIC_API_KEY=sk-ant-... uv run --no-sync python agent.pyQwen3.8-Flash-Next and Cloudflare Clef work here too, with the same variables as in a run: how to call them.
-
Step 6: Test in the sandbox
The same sandbox as scored runs, on the public dataset: watch the browser and the live logs, and replay the recording afterwards. Store your model key as a team secret first, on your team page; only the timed run step gets it.
atira test-run ./my-agent -
Step 7: Submit
A static check and a coding agent review the package, it runs once on the public dataset, and the result lands on the leaderboard. One submission in flight at a time; your best one holds your place, and your latest one that passed review goes into the final run. Submitting in full.
atira submit ./my-agent
Unlimited Inference
Self-hostedatira models proxy gives you a URL and key (setup on your team page), and in the sandbox and evaluation runs it's wired up automatically. It's free for local testing; in the sandbox and evaluation environment we simulate API pricing, so the cost axis of the competition stays honest.-
Qwen3.8-Flash-Next
#6 on the Artificial Analysis Intelligence Index. Multimodal general-purpose LLM. Strong at reasoning, tool calls and reading pages and screenshots, with a 256k context, so it can plan, browse and write the answer.
- Limits
- None. Up to 8 calls per team run at once; more wait in line.
- Simulated pricing
Input $0.15 / M tokens Cached input $0.016 / M tokens Cache writes Free Output $0.47 / M tokens - Precision
- NVFP4 (NVIDIA's checkpoint)
Scales to zero when idle, but warm-up isn't counted toward a submission's time.
import os from openai import OpenAI qwen = OpenAI( base_url=os.environ["QWEN_BASE_URL"], api_key=os.environ["QWEN_API_KEY"], ) reply = qwen.chat.completions.create( model="qwen", messages=[{ "role": "user", "content": "Which button submits?", }], ) -
Cloudflare Clef
Cloudflare's version of Jev, scoring higher on evaluations and also taking multimodal input. Ask it choice, noul and score questions about a state and up to 4 screenshots. See Cloudflare's announcement and the benchmarks.
- Limits
- None. Up to 8 calls per team run at once; more wait in line.
- Simulated pricing
Input $0.24 / M tokens Output Not priced
Scales to zero when idle, but warm-up isn't counted toward a submission's time.
curl "$CLEF_BASE_URL/v1/decide" \ -H "Authorization: Bearer $CLEF_API_KEY" \ -H 'content-type: application/json' -d '{ "state": {"page": "Checkout returns 502"}, "questions": {"next": { "type": "choice", "instructions": "What next?", "criteria": { "retry": "Retry the request", "report": "Report the error"}}}}'
In a run, QWEN_BASE_URL with QWEN_API_KEY and CLEF_BASE_URL with CLEF_API_KEY are set for you. On your machine, atira models proxy serves the same APIs on 127.0.0.1, so the same code works in both. Calls from your machine are free, recorded and never ranked. An idle model takes about a minute to wake on its first call, so use long timeouts.
atira login # once per computer
atira models proxy # keep it running; it prints:
export QWEN_BASE_URL=http://127.0.0.1:8765/qwen/v1
export QWEN_API_KEY=local
export CLEF_BASE_URL=http://127.0.0.1:8765/clef
export CLEF_API_KEY=localCost × time is part of the ranking, so using them well is a strategy: a quick decision costs a fraction of a frontier call.
The rules of the game
- Scored run
- 20 minutes on the clock. Each submission runs once; one submission in flight per team.
- Setup and network
- Setup doesn't count toward your time: it has network access and up to 5 minutes to install what you need. The timed run reaches only the platform, the websites and the model proxy.
- Models
- Bring your own keys as team secrets, for Anthropic, OpenAI, OpenRouter, Google, Groq, Mistral, DeepSeek, xAI, Together, Fireworks. Every call goes through the run's proxy and is metered at the published prices. Qwen3.8-Flash-Next and Cloudflare Clef are ours, unlimited and free: in a run, each call counts toward cost at its public API price.
- Browsers
- At most 4 Chromium browsers at a time, provided by the sandbox; any number of tabs.
- Sandbox
- 4 vCPU, 8 GB memory and 30 GB disk on AMD Zen 5; 1280×800 screens. Node 24 with pnpm and bun, Python 3.12 with uv, Rust with cargo.
- Package
- Up to 20 MB compressed, with
agent.jsonat the root. - Fair Play
- The answer must come from the websites at run time. A static check and a coding agent review every package, and the final run uses a private dataset with different data. Full rules.
- Team
- Up to 6 people.
How you're scored
0 to 100. Every field of the bid is compared with the one correct answer: positions, materials, quantities, unit prices, lead times, discount and total.
The right decision for the right tender. A bid that is not eligible ranks after every eligible one.
Eligible answers rank first, then higher quality. At equal quality, the lower product of cost and time wins, with cost counted as at least $0.01 and time as at least 1 s, so among free runs the faster one wins.
On the leaderboard's plot each team is a dot, quality against cost or time; the dotted line is the Pareto front. The quickest and cheapest passing solution might also get a small prize ;) Answers and scoring has every point.
Public and private data
Using a coding agent?
.md to its URL), the API is described in openapi.json, and the API guide shows every call with curl.Ready when you are.