Versioned benchmark collection

JRPG Boss Battle v1

Harness v1 · 4 measured runs across 1 target. Each entry is a single recorded attempt; repeat-run reliability is not established.

JRPG boss battle benchmark output in the JRPG Boss Battle v1 collection
JRPG boss battle benchmark output in the JRPG Boss Battle v1 collection
4Measured runs
1Frozen targets
4Model families
4Configurations
Collection scope

Recorded evidence, without a synthetic score.

Runs share the same frozen target and prompt within each test, with disclosed provider-specific stacks. Costs are recorded estimates as labelled in each ledger, not subscription invoices.

Targets

  • JRPG Boss Battle

Model coverage

  • Claude Fable 5 · Max
  • GPT-5.6 Sol · Ultra
  • Kimi K3 · Max
  • Qwen3.8 Max Preview · Mandatory thinking enabled

Evidence available

  • Three measured run ledgers
  • Public recorded outputs
  • Cost, token, and workflow timing

Still missing

  • Repeat-run reliability evidence
  • Minimum blind-vote sample
Public receipts

4 measured run ledgers

Open any row for its recorded workflow cost, timing, token usage, stack disclosure, artifact integrity, and explicit evidence gaps.