Getting started

From sign-in to your first scored run

Everything between signing in and the leaderboard: the commands to run, logins to scout the websites, the rules of the game and how you're scored.

What you're building

Build an agent system that wins a real tender: it finds the tender, the current specification, prices, stock and the approved discount across four hostile business websites, and submits one correct, signed bid. Bring any frontier model, your own tools and prompts, up to four browsers in parallel; no human in the loop. Or use ours: unlimited Qwen3.8-Flash-Next and Cloudflare Clef inference, free of charge. Every scored submission runs once on the public dataset and ranks by quality first, then by cost × time. When the hackathon ends, every team's latest submission runs once on a private dataset, and that run decides the winners.

Your first run in five minutes

Seven steps, each with what to run. The first five take about five minutes and get your agent working against the public websites; the last two take it to the leaderboard.
  1. Step 1: Sign in and form a team

    Sign in with GitHub and create a team of up to 6. Invite teammates by their GitHub username on your team page; each invite works only for that account.

  2. Step 2: Install the CLI and sign in

    The CLI is one file for Node 24. atira login opens the browser, where you approve a token for this computer. Or skip the install: point a coding agent at https://ehl-challenge.atira.ai/llms.txt.

    curl -fsSL https://ehl-challenge.atira.ai/install.sh | sh
    atira login
  3. Step 3: Scout the websites yourself

    Permanent logins for every website, on the public dataset: the same sites your agent will face. Explore freely; nothing done there counts.

    WebsiteURLUsernamePassword
    Tender portalVergabeportal atira-rfq-tender-portal.kyora.run u_20115c791c10c92c DrqJ-mbv75jbQW-NXXZO
    Supplier portalLieferantenportal atira-rfq-customer-portal.kyora.run u_20115c791c10c92c TA8rOLDy3r3UJLYM2shL
    ERP atira-rfq-erp.kyora.run u_20115c791c10c92c Ra7hKrCMINT5wjygR0X9
    WebmailWerkpost atira-rfq-mail.kyora.run u_20115c791c10c92c 5fq1jvwSqyc3yPXXPIJW
  4. Step 4: Grab the starter kit

    A working Python agent to fork: agent.py gives Claude browser tools on the sandbox's Chromium, the helper modules handle page snapshots, dialogs, downloads and German numbers, bid.py checks the bid before it goes in, and agent.json tells the sandbox how to install and run it.

    mkdir my-agent
    curl -fsSL https://ehl-challenge.atira.ai/starter-kit.tar.gz | tar xz -C my-agent
  5. Step 5: Develop locally

    atira scout --env exports the scouting logins the way a run gets its logins, so your agent runs on your machine against the public websites, with a local browser and your own model key. There is no run token and nothing done there counts. With these logins, an answer entered on the tender portal is acknowledged but never counts. When it works, atira test-run runs it in the sandbox for up to 2 hours and opens the run live, with the score and full breakdown. Details: how a run works.

    cd my-agent && uv sync && uv run playwright install chromium
    eval "$(atira scout --env)"
    export ANTHROPIC_API_KEY=sk-ant-...
    uv run --no-sync python agent.py

    Qwen3.8-Flash-Next and Cloudflare Clef work here too, with the same variables as in a run: how to call them.

  6. Step 6: Test in the sandbox

    The same sandbox as scored runs, on the public dataset: watch the browser and the live logs, and replay the recording afterwards. Store your model key as a team secret first, on your team page; only the timed run step gets it.

    atira test-run ./my-agent
  7. Step 7: Submit

    A static check and a coding agent review the package, it runs once on the public dataset, and the result lands on the leaderboard. One submission in flight at a time; your best one holds your place, and your latest one that passed review goes into the final run. Submitting in full.

    atira submit ./my-agent

Unlimited Inference

Self-hosted
We deployed Qwen3.8-Flash-Next and Cloudflare's Clef on our infra for this track. You get unlimited inference: locally, atira models proxy gives you a URL and key (setup on your team page), and in the sandbox and evaluation runs it's wired up automatically. It's free for local testing; in the sandbox and evaluation environment we simulate API pricing, so the cost axis of the competition stays honest.
  • Qwen3.8-Flash-Next

    #6 on the Artificial Analysis Intelligence Index. Multimodal general-purpose LLM. Strong at reasoning, tool calls and reading pages and screenshots, with a 256k context, so it can plan, browse and write the answer.

    Limits
    None. Up to 8 calls per team run at once; more wait in line.
    Simulated pricing
    Input$0.15 / M tokens
    Cached input$0.016 / M tokens
    Cache writesFree
    Output$0.47 / M tokens
    Precision
    NVFP4 (NVIDIA's checkpoint)

    Scales to zero when idle, but warm-up isn't counted toward a submission's time.

    import os
    from openai import OpenAI
    
    qwen = OpenAI(
        base_url=os.environ["QWEN_BASE_URL"],
        api_key=os.environ["QWEN_API_KEY"],
    )
    reply = qwen.chat.completions.create(
        model="qwen",
        messages=[{
            "role": "user",
            "content": "Which button submits?",
        }],
    )
  • Cloudflare Clef

    Cloudflare's version of Jev, scoring higher on evaluations and also taking multimodal input. Ask it choice, noul and score questions about a state and up to 4 screenshots. See Cloudflare's announcement and the benchmarks.

    Limits
    None. Up to 8 calls per team run at once; more wait in line.
    Simulated pricing
    Input$0.24 / M tokens
    OutputNot priced

    Scales to zero when idle, but warm-up isn't counted toward a submission's time.

    curl "$CLEF_BASE_URL/v1/decide" \
      -H "Authorization: Bearer $CLEF_API_KEY" \
      -H 'content-type: application/json' -d '{
      "state": {"page": "Checkout returns 502"},
      "questions": {"next": {
        "type": "choice",
        "instructions": "What next?",
        "criteria": {
          "retry": "Retry the request",
          "report": "Report the error"}}}}'
In a run and on your machine

In a run, QWEN_BASE_URL with QWEN_API_KEY and CLEF_BASE_URL with CLEF_API_KEY are set for you. On your machine, atira models proxy serves the same APIs on 127.0.0.1, so the same code works in both. Calls from your machine are free, recorded and never ranked. An idle model takes about a minute to wake on its first call, so use long timeouts.

atira login   # once per computer
atira models proxy   # keep it running; it prints:
export QWEN_BASE_URL=http://127.0.0.1:8765/qwen/v1
export QWEN_API_KEY=local
export CLEF_BASE_URL=http://127.0.0.1:8765/clef
export CLEF_API_KEY=local

Cost × time is part of the ranking, so using them well is a strategy: a quick decision costs a fraction of a frontier call.

The rules of the game

The constraints every scored run lives with.
Scored run
20 minutes on the clock. Each submission runs once; one submission in flight per team.
Setup and network
Setup doesn't count toward your time: it has network access and up to 5 minutes to install what you need. The timed run reaches only the platform, the websites and the model proxy.
Models
Bring your own keys as team secrets, for Anthropic, OpenAI, OpenRouter, Google, Groq, Mistral, DeepSeek, xAI, Together, Fireworks. Every call goes through the run's proxy and is metered at the published prices. Qwen3.8-Flash-Next and Cloudflare Clef are ours, unlimited and free: in a run, each call counts toward cost at its public API price.
Browsers
At most 4 Chromium browsers at a time, provided by the sandbox; any number of tabs.
Sandbox
4 vCPU, 8 GB memory and 30 GB disk on AMD Zen 5; 1280×800 screens. Node 24 with pnpm and bun, Python 3.12 with uv, Rust with cargo.
Package
Up to 20 MB compressed, with agent.json at the root.
Fair Play
The answer must come from the websites at run time. A static check and a coding agent review every package, and the final run uses a private dataset with different data. Full rules.
Team
Up to 6 people.

How you're scored

Every scored run gets a quality and an eligibility verdict; the platform measures its cost and time.
Quality

0 to 100. Every field of the bid is compared with the one correct answer: positions, materials, quantities, unit prices, lead times, discount and total.

Eligibility

The right decision for the right tender. A bid that is not eligible ranks after every eligible one.

OrderEligible firstQuality ↓Cost × time ↑

Eligible answers rank first, then higher quality. At equal quality, the lower product of cost and time wins, with cost counted as at least $0.01 and time as at least 1 s, so among free runs the faster one wins.

On the leaderboard's plot each team is a dot, quality against cost or time; the dotted line is the Pareto front. The quickest and cheapest passing solution might also get a small prize ;) Answers and scoring has every point.

Public and private data

The leaderboard runs on the public dataset: the websites and scenario (public-1) you scout and test-run on. Every scored submission runs once on it. The correct answer isn't published; your system has to work it out from the websites. Test runs and the scouting logins use it too. When the hackathon ends, the organizers run every team's latest submission that passed review once on a private dataset: the same websites and rules, different data. That run decides the winners, so build a system that works the answer out from the websites, not one tuned to the public scenario.

Using a coding agent?

Point it at https://ehl-challenge.atira.ai/llms.txt; every page is also Markdown (add .md to its URL), the API is described in openapi.json, and the API guide shows every call with curl.

Ready when you are.