How battles work

Two AI-built games. One blind battle.

You are the judge. Play both separately, choose the stronger result, then reveal the evidence behind each build.

  1. Play two builds

    Two builds from the same randomly selected benchmark appear one at a time. Play each until you have seen enough, then continue.

  2. Judge them blind

    Model, provider, reasoning level, price, and token usage stay hidden while you play. Judge the output—not the logo.

  3. Choose the stronger result

    Pick the build you preferred based on the experience itself. Play first, then sign in only to record one accountable vote.

  4. Reveal and rotate

    Reveal both identities and their public results after voting, with the next fresh matchup ready in the same flow.

Every signed-in contributor earns 1 Bench Point for each fresh, completed blind comparison. Lifetime status milestones are marked at 10, 25, 50, and 100 points and never influence benchmark verdicts.

6 benchmark groups · 17 playable buildsStart battling →