Measured harness ledgerPublic result
Qwen3.8 Max

Red Sands v1 — Qwen3.8 Max xhigh

Author one complete voiced quest inside the fixed Red Sands browser game, including dialogue, travel, a gunfight, a cutscene, journal state, and an ending.

xhigh reasoningHeadline result
Workflow cost
¥235.23 estimated API-list-price equivalent (upper bound)
Wall-clock
1h 5m 52.3s wall-clock
Processed tokens
19.32M processed
Record state
ledger
Public summary

Qwen3.8 Max xhigh public partial ledger: 1h 5m 52.3s wall-clock, 19.32M processed, and ¥235.23 estimated API-list-price equivalent (upper bound).

Cost basis
  • Computed here for comparability: applying the pricing page's generic 10% cache-hit rule (¥1.2/M) to the 18,405,588 cache-read tokens gives ¥36.45. The receipt reports only the ¥235.23 upper bound.
  • This row headlines the full-input-rate UPPER BOUND, matching the Stage 1 STARFALL, Fighter Game and Cathedral rows. The Stage 2 Landfall and JRPG rows headline the 10% cache-hit rule instead. Convert before comparing Qwen rows.
  • Qwen Code records no per-request cost field; null by construction.
  • Pricing source recorded as: official Alibaba Model Studio rate card, Beijing region, resolved live.
  • Non-USD upper bound; excluded from USD averages and the cost × time plot.
Visible evidence gaps
  • a human playthrough of the quest, start to finish
  • blind-evaluation record — the primary instrument for this task
  • reconciliation of the root ledger-vs-transcript token gap
  • the handback note the prompt asks for
Builder test available

This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.

  • Red Sands v1 · Builder projects v1
RemakeBenchResearch console