Measured harness ledgerPublic result
Qwen3.8 MaxRed Sands v1 — Qwen3.8 Max xhigh
Author one complete voiced quest inside the fixed Red Sands browser game, including dialogue, travel, a gunfight, a cutscene, journal state, and an ending.
xhigh reasoningHeadline result
- Workflow cost
- ¥235.23 Independently recomputed first-party API-list-price equivalent for the measured token mix; NOT a billed amount. This run used an Alibaba Model Studio token-plan subscription, under which usage is not itemized as cash. (upper bound)
- Wall-clock
- 1h 5m 52.3s wall-clock
- Processed tokens
- 19.32M processed
- Record state
- artifact_model_playthrough_complete_ledger_partial_metrics
Public summary
Qwen3.8 Max xhigh artifact_model_playthrough_complete_ledger_partial_metrics ledger: 1h 5m 52.3s wall-clock, 19.32M processed, and ¥235.23 Independently recomputed first-party API-list-price equivalent for the measured token mix; NOT a billed amount. This run used an Alibaba Model Studio token-plan subscription, under which usage is not itemized as cash. (upper bound).
Run identity and stack
- Result ID: red-sands-quest-threejs-qwen3.8-max-xhigh
- Technical model: qwen3.8-max
- Provider: Alibaba Cloud Model Studio token plan
- Client: Qwen Code 0.21.6 installed / 0.21.6 transcript-recorded
- Stack: Alibaba Cloud Model Studio token plan
- Stack: Qwen Code 0.21.6 installed / 0.21.6 transcript-recorded
- Stack: Technical model/configuration: qwen3.8-max
- Stack: Three.js quest-authoring fixture
- Stack: Harness v1 frozen prompt
Cost basis
- This row headlines the full-input-rate UPPER BOUND, matching the Stage 1 STARFALL, Courtyard Arena and Cathedral rows. The Stage 2 Landfall and JRPG rows headline the 10% cache-hit rule instead. Convert before comparing Qwen rows.
- Qwen Code records no per-request cost field; null by construction.
- Computed here for comparability: applying the pricing page's generic 10% cache-hit rule (¥1.2/M) to the 18,405,588 cache-read tokens gives ¥36.45. The receipt reports only the ¥235.23 upper bound.
Primary artifact integrity
- Kind: manifest-verified-artifact
- Path: artifacts/red-sands-quest-threejs-qwen3.8-max-xhigh/source/src/quest/quests/the-deputys-price.js
- SHA-256: 6fd3aabf6f5536ba8b509ca72633478ed808112acf5d5a4649784470382852c1
Recorded caveats
- Metrics status is PARTIAL: the root ledger and root transcript token sums disagree by 20,256 fresh input, 390,764 cache read, 8,369 visible output and 2,392 reasoning tokens. The ledger figures are reported and the gap is disclosed, not reconciled.
- This task has NO automated completion gate, by design.
- No operator playthrough exists — the verification browser cannot boot the game, because a backgrounded pane throttles the rAF-driven ~10 s terrain generation.
- npm run smoke runs under ?capture=1, which never constructs the quest system.
- The completion evidence is the model's own scripted playthrough record, not an operator reproduction.
- The prompt's required handback note was not delivered; no stated-gaps section exists.
- Cache reads are normally discounted; pricing them at the fresh-input rate makes ¥235.23 an upper bound.
- No cache-write counter exists on this stack, so that line item is unavailable rather than zero.
- The client version is reported as-is: the benchmark prompt declared 0.21.7 and the observed binary is 0.21.6.
- Context occupancy is a per-request footprint, not cumulative usage; quota fields are null with no before/after evidence.
- The endpoint host is unresolved because the provider settings file is off-limits under the measurement safety rule.
- The voice audio was not listened to; delivery, count and casting are confirmed, performance is not.
- This is a quest-authorship task over a fixed game, so its figures are not comparable to the archive's from-scratch build rows.
- The archive is a delta over the fixture, not a standalone runnable tree.
- No quality score and no blind-evaluation record exist — which for this task is where the entire measurement lives.
- A genuine one-shot run. The two Codex peers on this task each recorded 2 user turns.
Visible evidence gaps
- a human playthrough of the quest, start to finish
- blind-evaluation record — the primary instrument for this task
- reconciliation of the root ledger-vs-transcript token gap
- the handback note the prompt asks for
Builder test available
This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.
- Red Sands v1 · Builder projects v1
