Versioned benchmark collection
JRPG Boss Battle v1
Harness v1 · 4 measured runs across 1 target. Each entry is a single recorded attempt; repeat-run reliability is not established.

4Measured runs
1Frozen targets
4Model families
4Configurations
Collection scope
Recorded evidence, without a synthetic score.
Runs share the same frozen target and prompt within each test, with disclosed provider-specific stacks. Costs are recorded estimates as labelled in each ledger, not subscription invoices.
Targets
- JRPG Boss Battle
Model coverage
- Claude Fable 5 · Max
- GPT-5.6 Sol · Ultra
- Kimi K3 · Max
- Qwen3.8 Max Preview · Mandatory thinking enabled
Evidence available
- Three measured run ledgers
- Public recorded outputs
- Cost, token, and workflow timing
Still missing
- Repeat-run reliability evidence
- Minimum blind-vote sample
Public receipts
4 measured run ledgers
Open any row for its recorded workflow cost, timing, token usage, stack disclosure, artifact integrity, and explicit evidence gaps.
01Claude Fable 5Max reasoning · JRPG Boss Battle$123.42 estimated API-equivalent costrecorded estimate50:48.1 wall-clockEnd-to-end workflowOpen 02GPT-5.6 SolUltra reasoning · JRPG Boss Battle$9.98 estimated API-equivalent costrecorded estimate25:51.8 wall-clockEnd-to-end workflowOpen 03Kimi K3Max reasoning · JRPG Boss Battle$2.86 estimated API-equivalent costrecorded estimate41:20.2 wall-clockEnd-to-end workflowOpen 04Qwen3.8 Max PreviewMandatory thinking enabled reasoning · JRPG Boss Battle$15.15 estimated API-equivalent costrecorded estimate58m 40.6s wall-clockEnd-to-end workflowOpen
