Agent challenges for industrial sales engineering

Build agent systems that do real work.

A hackathon for agent systems on deliberately hostile business websites. Your system runs in a sandbox, and a live leaderboard ranks it by how right it is, then by how cheap and fast. A final run on private data decides the winners.

Websites
4
Test run
2 hours
Scored run
20 minutes
Leaderboard
Public dataset

Live updates every 10 seconds.

Teams
2
Submissions
1
Running now
0
Best quality
100

Leaderboard

Public dataset

Eligible first, then quality, then cost × time.

Full leaderboard →
#TeamQualityCostTimeSubm.Trend
1 Atira reference (scripted) 100 $0.00 00:01:45 1 –
EligibleNot eligiblePareto
8090100 $0.0001$0.0002 MEDIAN MODEL COST PER RUN (USD, LOG SCALE) MEDIAN QUALITY Atira reference (scripted) · quality 100 · $0.00 Atira reference (scripted)

From scattered RFQ to bid

Build an agent system that wins a real tender: it finds the tender, the current specification, prices, stock and the approved discount across four hostile business websites, and submits one correct, signed bid.

The challenge →
TENDER PORTAL SUPPLIER PORTAL ERP WEBMAIL ONE BID
  1. 01 · tender-portalOpen site ↗

    Tender portalVergabeportal

    The public tender, its bill of quantities, the bidder Q&A that changes quantities, and the bid configurator.

    Cookie wall, flaky error pages, documents behind an expression of interest, iframes and ZIP files, a deadline in another timezone.

  2. 02 · customer-portalOpen site ↗

    Supplier portalLieferantenportal

    The buyer's technical specification and its revisions.

    Opens the oldest revision first, nested document viewers, paged requirement tables.

  3. 03 · erpOpen site ↗

    ERP

    Your company's materials, stock, lead times and price conditions with validity periods and quantity scales.

    SAP-style transaction codes, an F4 search popup, status codes, a short session timeout.

  4. 04 · mailOpen site ↗

    WebmailWerkpost

    Your sales inbox, with the sales director's discount approval and the signing TAN.

    Look-alike approvals for other tenders, superseded approvals, an archive folder, paging.

These websites, with their public scenario, are the public dataset: test runs, the scouting logins and the leaderboard use them. The final run uses a private dataset: the same websites and rules, different data. Scout them yourself →

  1. 01

    Form a team

    Sign in with GitHub, create a team and invite your teammates. Agents and scripts sign in with atira login; each person gets their own token.

  2. 02

    Develop locally

    Scout the websites, fork the starter kit and run your agent on your machine against the public websites with atira scout --env. Nothing done there counts.

  3. 03

    Test in the sandbox

    Upload your package as a test run. Watch its browser live, read its logs, replay its recording.

  4. 04

    Submit

    Each submission is reviewed, then run once in the sandbox on the public dataset. Its score ranks. At the end, a final run on a private dataset decides the winners.

Free models, no limits

We run Qwen3.8-Flash-Next and Cloudflare Clef on our own inference deployment for this track. No limits and no keys: your submissions get simulated API pricing, so each call counts toward the run's cost at the model's public API price and the ranking stays honest, but nobody pays.

How to call them →
  • Frontier-class open modelIdle

    Qwen3.8-Flash-Next

    General reasoning, tool calls and vision: plan the work, read pages and screenshots, write the answer.

    Limits
    None. Free for you, no key needed.
    Counted at
    $0.15 input and $0.47 output per million tokens
    Status
    The first call wakes it, in about a minute.
  • Multimodal decision modelIdle

    Cloudflare Clef

    Fast, structured decisions over a state and up to 4 screenshots: ask it choice, noul and score questions and get one answer each.

    Limits
    None. Free for you, no key needed.
    Counted at
    $0.24 per million input tokens, output not priced
    Status
    The first call wakes it, in about a minute.

Activity

  1. Atira reference (scripted) was evaluated: quality 100, rank 1
  2. Atira reference (scripted) passed review
  3. Atira reference (scripted) submitted a package
  4. Elia Testing registered
  5. Atira reference (scripted) registered