The challenge

From scattered RFQ to bid

Build the agent system that does a sales engineer's hardest afternoon on its own.

A German hydraulics manufacturer receives a public tender for hydraulic cylinders. The answer is spread across four websites built to be hard for agents: a tender portal, a supplier portal with the buyer's specification, the company's ERP, and the sales inbox. Your system reads all of it, applies the bid rules, and enters one priced, signed bid in the tender portal. No human in the loop.

Bring anything
Any frontier model, your own tools and prompts, up to four browsers in parallel. If it runs in the sandbox and finds the answer on the websites at run time, it counts.
How you win
Right first, then efficient: eligible answers rank first, then higher quality, then lower cost × time. The leaderboard uses the public dataset; when the hackathon ends, one run on a private dataset decides the winners.
Websites
4
Test run time limit
2 hours
Scored run time limit
20 minutes
Runs per submission
1

For agentsPoint your coding agent at https://ehl-challenge.atira.ai/llms.txt: it covers the challenge, the API and submitting with plain curl. Every page is also Markdown, and the API is described in openapi.json. An optional one-file CLI wraps the same calls.

The task

One public tender, four German business websites, one priced bid. Your agent enters the bid in the tender portal's configurator and signs it with a TAN from the webmail.
TENDER PORTAL SUPPLIER PORTAL ERP WEBMAIL ONE BID

The websites

Every run gets its own logins for each website, in the run's brief; to look around yourself, use the scouting logins below. They are built to be hard for agents: decoys, popups, expiring sessions and pages that change under you. Everything an agent needs is still visible in a 1280×800 browser.
  1. 01 · tender-portalOpen site ↗

    Tender portalVergabeportal

    The public tender, its bill of quantities, the bidder Q&A that changes quantities, and the bid configurator.

    Cookie wall, flaky error pages, documents behind an expression of interest, iframes and ZIP files, a deadline in another timezone.

  2. 02 · customer-portalOpen site ↗

    Supplier portalLieferantenportal

    The buyer's technical specification and its revisions.

    Opens the oldest revision first, nested document viewers, paged requirement tables.

  3. 03 · erpOpen site ↗

    ERP

    Your company's materials, stock, lead times and price conditions with validity periods and quantity scales.

    SAP-style transaction codes, an F4 search popup, status codes, a short session timeout.

  4. 04 · mailOpen site ↗

    WebmailWerkpost

    Your sales inbox, with the sales director's discount approval and the signing TAN.

    Look-alike approvals for other tenders, superseded approvals, an archive folder, paging.

Scout them yourself

Permanent logins for every website, on the public dataset: the same sites your agent will face. Explore freely; nothing done there counts.

WebsiteURLUsernamePassword
Tender portalVergabeportal atira-rfq-tender-portal.kyora.run u_20115c791c10c92c DrqJ-mbv75jbQW-NXXZO
Supplier portalLieferantenportal atira-rfq-customer-portal.kyora.run u_20115c791c10c92c TA8rOLDy3r3UJLYM2shL
ERP atira-rfq-erp.kyora.run u_20115c791c10c92c Ra7hKrCMINT5wjygR0X9
WebmailWerkpost atira-rfq-mail.kyora.run u_20115c791c10c92c 5fq1jvwSqyc3yPXXPIJW

Public and private data

Two datasets, the same websites and rules. The leaderboard runs on the public one; the final run, which decides the winners, on a private one.
Public dataset

The leaderboard runs on the public dataset: the websites and scenario (public-1) you scout and test-run on. Every scored submission runs once on it. The correct answer isn't published; your system has to work it out from the websites. Test runs and the scouting logins use it too.

Private dataset

When the hackathon ends, the organizers run every team's latest submission that passed review once on a private dataset: the same websites and rules, different data. That run decides the winners, so build a system that works the answer out from the websites, not one tuned to the public scenario.

The rules

The brief states these rules in full. Together they define exactly one correct answer.
  1. 01 · Quantity

    The tender quantity, overridden by the latest clarification that changes that position.

  2. 02 · Requirement

    The position's requirement in the specification revision marked current.

  3. 03 · Material

    The active ERP material whose bore, stroke, pressure, mounting and certification match exactly. If a position has none, the answer is no bid.

  4. 04 · Unit price

    The price condition valid on the tender date, at the highest quantity scale reached, minus the approved discount, rounded half-up to the cent.

  5. 05 · Discount

    The most recent approval from the sales director for this exact tender; otherwise 0 %.

  6. 06 · Lead time

    The from-stock lead time when stock covers the quantity, otherwise the production lead time.

  7. 07 · Total

    The sum of quantity times unit price, rounded half-up to the cent.

  8. 08 · Decision

    Bid, with one entry per tender position; no bid with no positions and a total of 0 when a material is missing.

How a run works

Develop on your machine against the public websites with the scouting logins (atira scout --env), then test in the sandbox: atira test-run uploads your package, starts a test run and opens it live, or drop the package on your team page. A test run lasts up to 2 hours on the public data and returns its score with the full breakdown. In the commands, $ATIRA_TOKEN holds your personal token (sign in with atira login, then atira token).
  1. 01

    Start

    The run issues a token, the task brief and fresh logins for every website.

  2. 02

    Read

    Your agent reads the brief, the rules and the answer format from the API.

  3. 03

    Work the websites

    It works through the 4 websites in a real browser and gathers the facts.

  4. 04

    Submit

    It enters the bid in the tender portal's configurator and signs it with a TAN from the webmail.

  5. 05

    Scored

    The answer is checked against the one correct answer; quality, cost and time are recorded.

Develop locally
eval "$(atira scout --challenge rfq --env)"   # logins for the public websites
curl "https://ehl-challenge.atira.ai/v1/challenges/rfq/scouting"   # the same, with the brief, as JSON
# run your agent on your machine, then once in the sandbox:
atira test-run ./my-agent --challenge rfq

atira scout --env exports the published scouting logins as CHALLENGE_SITE_ACCESS (the shape a run gets), each website's URL, and CHALLENGE_SCOUTING_URL, which returns the brief and the answer format without a sign-in. There is no run token: your agent calls its model with your own key, and nothing done there counts. With these logins, an answer entered on the tender portal is acknowledged but never counts. In a run, your agent reads the same at GET /v1/runs/current with the run token as x-api-key, and hands in one valid answer of up to 65,536 bytes; a malformed one returns its problems while the clock keeps running.

In the sandbox
CHALLENGE_API_URL
the platform API; the run is at /v1/runs/current
CHALLENGE_RUN_TOKEN
this run's token, sent as x-api-key
CHALLENGE_SITE_ACCESS
the websites with this run's logins, as JSON
LLM_PROXY_URL
the model proxy with the run token in its path; bring your own provider key
ANTHROPIC_BASE_URL
https://atira-challenge.kyora.run/llm/run/$CHALLENGE_RUN_TOKEN/anthropic
OPENAI_BASE_URL
https://atira-challenge.kyora.run/llm/run/$CHALLENGE_RUN_TOKEN/openai/v1
GOOGLE_GEMINI_BASE_URL
https://atira-challenge.kyora.run/llm/run/$CHALLENGE_RUN_TOKEN/google
MISTRAL_SERVER_URL
https://atira-challenge.kyora.run/llm/run/$CHALLENGE_RUN_TOKEN/mistral
TOGETHER_BASE_URL
https://atira-challenge.kyora.run/llm/run/$CHALLENGE_RUN_TOKEN/together/v1
CLEF_BASE_URL, CLEF_API_KEY
Clef through the proxy, https://atira-challenge.kyora.run/llm/run/$CHALLENGE_RUN_TOKEN/clef, with the run token as its key; uncapped for this track and never charged
QWEN_BASE_URL, QWEN_API_KEY
Qwen through the proxy, https://atira-challenge.kyora.run/llm/run/$CHALLENGE_RUN_TOKEN/qwen/v1, with the run token as its key; uncapped for this track and never charged
CDP_URL
the sandbox browser (Chrome DevTools Protocol), http://127.0.0.1:9222; with several, browser 1
CDP_URLS
every browser's CDP URL in order, as a JSON array: one per browser asked for in agent.json, up to 4
BROWSER_COUNT
how many browsers the sandbox started, 1 to 4
DISPLAY
the desktop browser 1 runs on, :1 at 1280×800; browser n is on :n
HTTP_PROXY, HTTPS_PROXY, http_proxy, https_proxy
the only way out: the platform and the websites (http://127.0.0.1:3128)
NO_PROXY, no_proxy
reached directly, so the browser's CDP stays local (localhost,127.0.0.1)
NODE_USE_ENV_PROXY
Node's fetch uses the proxy (1)

Unlimited Inference

Self-hosted
We deployed Qwen3.8-Flash-Next and Cloudflare's Clef on our infra for this track. You get unlimited inference: locally, atira models proxy gives you a URL and key (setup on your team page), and in the sandbox and evaluation runs it's wired up automatically. It's free for local testing; in the sandbox and evaluation environment we simulate API pricing, so the cost axis of the competition stays honest.
  • Qwen3.8-Flash-Next

    #6 on the Artificial Analysis Intelligence Index. Multimodal general-purpose LLM. Strong at reasoning, tool calls and reading pages and screenshots, with a 256k context, so it can plan, browse and write the answer.

    Limits
    None. Up to 8 calls per team run at once; more wait in line.
    Simulated pricing
    Input$0.15 / M tokens
    Cached input$0.016 / M tokens
    Cache writesFree
    Output$0.47 / M tokens
    Precision
    NVFP4 (NVIDIA's checkpoint)

    Scales to zero when idle, but warm-up isn't counted toward a submission's time.

    import os
    from openai import OpenAI
    
    qwen = OpenAI(
        base_url=os.environ["QWEN_BASE_URL"],
        api_key=os.environ["QWEN_API_KEY"],
    )
    reply = qwen.chat.completions.create(
        model="qwen",
        messages=[{
            "role": "user",
            "content": "Which button submits?",
        }],
    )
  • Cloudflare Clef

    Cloudflare's version of Jev, scoring higher on evaluations and also taking multimodal input. Ask it choice, noul and score questions about a state and up to 4 screenshots. See Cloudflare's announcement and the benchmarks.

    Limits
    None. Up to 8 calls per team run at once; more wait in line.
    Simulated pricing
    Input$0.24 / M tokens
    OutputNot priced

    Scales to zero when idle, but warm-up isn't counted toward a submission's time.

    curl "$CLEF_BASE_URL/v1/decide" \
      -H "Authorization: Bearer $CLEF_API_KEY" \
      -H 'content-type: application/json' -d '{
      "state": {"page": "Checkout returns 502"},
      "questions": {"next": {
        "type": "choice",
        "instructions": "What next?",
        "criteria": {
          "retry": "Retry the request",
          "report": "Report the error"}}}}'
In a run and on your machine

In a run, QWEN_BASE_URL with QWEN_API_KEY and CLEF_BASE_URL with CLEF_API_KEY are set for you. On your machine, atira models proxy serves the same APIs on 127.0.0.1, so the same code works in both. Calls from your machine are free, recorded and never ranked. An idle model takes about a minute to wake on its first call, so use long timeouts.

atira login   # once per computer
atira models proxy   # keep it running; it prints:
export QWEN_BASE_URL=http://127.0.0.1:8765/qwen/v1
export QWEN_API_KEY=local
export CLEF_BASE_URL=http://127.0.0.1:8765/clef
export CLEF_API_KEY=local

Cost × time is part of the ranking, so using them well is a strategy: a quick decision costs a fraction of a frontier call.

Submitting

A submission is your agent system as a gzipped tarball. It is reviewed automatically, then run once in the sandbox on the public dataset, and that score ranks your team. Submit as often as you like, one at a time; your best submission holds your place on the leaderboard, and your latest one that passed review goes into the final run.
agent.json
{
  "setup": ["uv", "sync"],
  "run": ["uv", "run", "python", "agent.py"]
}

Put agent.json at the root of the tarball, next to your code and its dependency file. Up to 20 MB compressed. setup installs dependencies with network access, within 5 minutes that don't count toward your time; run is timed and reaches only the platform and the websites. "browsers": 1–4 (default 1) is how many browsers the sandbox starts, each on its own screen. Node 24 with pnpm and bun, Python 3.12 with uv, and Rust with cargo are installed. The starter kit is a working agent to fork.

Upload
tar czf agent.tar.gz -C my-agent .
curl --data-binary @agent.tar.gz \
  -H "x-api-key: $ATIRA_TOKEN" \
  -H "content-type: application/gzip" \
  "https://ehl-challenge.atira.ai/v1/challenges/rfq/submissions"

Or drop the tarball on your team page. Upload it as a test run first to watch it work in the sandbox: its browser live, its logs, and a recording afterwards.

Scoring and ranking

Every submission is measured on its run in the sandbox.
Quality

0 to 100. Every field of the bid is compared with the one correct answer: positions, materials, quantities, unit prices, lead times, discount and total.

Eligibility

The right decision for the right tender. A bid that is not eligible ranks after every eligible one.

Cost

Every model call goes through the proxy and is priced per model. There is no budget: at equal quality, cost × time decides.

Time

Wall-clock time from the start of the run until the answer is accepted. It weighs as much as cost: halving either halves cost × time.

OrderEligible firstQuality ↓Cost × time ↑

Eligible answers rank first, then higher quality. At equal quality, the lower product of cost and time wins, with cost counted as at least $0.01 and time as at least 1 s, so among free runs the faster one wins.

Leaderboard

Public dataset
Eligible first, then quality, then cost × time. The top ten; the full leaderboard has every team.
#TeamQualityCostTimeSubm.Trend
1 Atira reference (scripted) 100 $0.00 00:01:45 1 –

Fair play

The leaderboard rewards agent systems that do the work.
  • Your system must work the answer out from the websites at run time. A static check and a coding agent review every package; hardcoded answers are rejected.
  • Bring your own model key for Anthropic, OpenAI, OpenRouter, Google, Groq, Mistral, DeepSeek, xAI, Together, Fireworks, kept as a team secret. Every call goes through the run's proxy and is metered at the published prices. Qwen3.8-Flash-Next and Cloudflare Clef are ours, unlimited and free: in a run, each call counts toward cost at its public API price. The sandbox reaches the platform and the websites, nothing else.
  • At most 4 browsers run at a time: ask for up to 4 in agent.json ("browsers"), use tabs and browser contexts inside them freely, and don't start browsers of your own, headless ones included. The sandbox stops any other browser your agent starts.
  • One submission in flight per team. Submit again once it has finished.

The answer format

The configurator on the tender portal collects exactly this answer; the run's brief carries the same schema.
Show the JSON schema
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "type": "object",
  "properties": {
    "tender_reference": {
      "type": "string",
      "maxLength": 100
    },
    "decision": {
      "type": "string",
      "enum": [
        "bid",
        "no_bid"
      ]
    },
    "currency": {
      "type": "string",
      "const": "EUR"
    },
    "discount_percent": {
      "type": "number",
      "minimum": 0,
      "maximum": 100
    },
    "positions": {
      "maxItems": 100,
      "type": "array",
      "items": {
        "type": "object",
        "properties": {
          "position": {
            "type": "string",
            "maxLength": 20
          },
          "material_number": {
            "type": "string",
            "maxLength": 50
          },
          "quantity": {
            "type": "integer",
            "minimum": 0,
            "maximum": 9007199254740991
          },
          "unit_price": {
            "type": "number",
            "minimum": 0
          },
          "lead_time_weeks": {
            "type": "integer",
            "minimum": 0,
            "maximum": 9007199254740991
          }
        },
        "required": [
          "position",
          "material_number",
          "quantity",
          "unit_price",
          "lead_time_weeks"
        ],
        "additionalProperties": false
      }
    },
    "total_net": {
      "type": "number",
      "minimum": 0
    },
    "notes": {
      "type": "string",
      "maxLength": 4000
    }
  },
  "required": [
    "tender_reference",
    "decision",
    "currency",
    "discount_percent",
    "positions",
    "total_net"
  ],
  "additionalProperties": false
}

Ready when you are.