Measured harness ledgerPublic result
Qwen3.8 Max

Fighter Game v1 — Qwen3.8 Max xhigh

Improve the complete arena environment of a fixed Three.js fighting game through Blender MCP without changing the fighters, combat, HUD, effects, or audio.

xhigh reasoningHeadline result
Workflow cost
¥147.93 Independently recomputed first-party API-list-price equivalent for the measured token mix; NOT a billed amount. This run used an Alibaba Model Studio token-plan subscription, under which usage is not itemized as cash. (upper bound)
Wall-clock
1h 17m 10.2s wall-clock
Processed tokens
12.01M processed
Record state
artifact_independent_boot_ledger_partial_metrics
Public summary

Qwen3.8 Max xhigh artifact_independent_boot_ledger_partial_metrics ledger: 1h 17m 10.2s wall-clock, 12.01M processed, and ¥147.93 Independently recomputed first-party API-list-price equivalent for the measured token mix; NOT a billed amount. This run used an Alibaba Model Studio token-plan subscription, under which usage is not itemized as cash. (upper bound).

Run identity and stack
  • Result ID: courtyard-arena-blender-mcp-qwen3.8-max-xhigh
  • Technical model: qwen3.8-max
  • Provider: Alibaba Cloud Model Studio token plan
  • Client: Qwen Code 0.21.6 installed / 0.21.6 transcript-recorded
  • Stack: Alibaba Cloud Model Studio token plan
  • Stack: Qwen Code 0.21.6 installed / 0.21.6 transcript-recorded
  • Stack: Technical model/configuration: qwen3.8-max
  • Stack: Blender MCP
  • Stack: Three.js fighting-game fixture
  • Stack: Harness v1 frozen prompt
Cost basis
  • This row headlines the full-input-rate UPPER BOUND, matching the Stage 1 STARFALL row. The Stage 2 Landfall and JRPG rows headline the 10% cache-hit rule instead. Convert before comparing any two Qwen rows.
  • Qwen Code records no per-request cost field; null by construction.
  • Beijing / US-Virginia / Germany / Japan tier: input ¥12/M, output ¥36/M. The Singapore tier (¥14.988 / ¥44.965) is disclosed but not applied. The international page does not list qwen3.8-max.
  • Computed here for comparability: applying the pricing page's generic 10% cache-hit rule (¥1.2/M) to the 10,753,881 cache-read tokens gives ¥31.79. The receipt itself reports only the ¥147.93 upper bound.
Primary artifact integrity
  • Kind: manifest-verified-artifact
  • Path: artifacts/courtyard-arena-blender-mcp-qwen3.8-max-xhigh/source/web/models/arena.glb
  • SHA-256: 17ceb265b01b533bde93b0767272cc7cc8ca46ce6b46f98e5a892ec561e4e5e0
Validation evidence
  • Result: PASS
Recorded caveats
  • Metrics status is PARTIAL: the root ledger and root transcript token sums disagree by 43,646 fresh input, 272,059 cache read, 24,197 visible output, and 14,912 reasoning tokens. The ledger figures are reported and the gap is disclosed, not reconciled.
  • Wall-clock (4,630.2 s) is end-to-end workflow latency including Blender builds, browser checks, and idle gaps — not model-only compute. Summed API duration (4,771.8 s) is separate and exceeds it because subagents ran concurrently.
  • One Qwen Code agent turn fans out to many billable requests and subagent sessions under a single session id (101 root + 14 subagent here).
  • Visible output and reasoning are separate fields; both are output-priced on this stack.
  • inputTokens includes cached tokens and outputTokens includes reasoning tokens, so the ledger's totalTokens and total_processed are different quantities and must never be swapped.
  • Qwen Code records no cache-write counter, so that line item is unavailable rather than zero.
  • Cache reads are normally discounted; pricing them at the fresh-input rate makes ¥147.93 an upper bound.
  • The endpoint host is unresolved because reading the provider settings file is forbidden by the measurement safety rules.
  • The client version is reported as-is: installed and transcript-recorded both read 0.21.6 while the expected stack listed 0.21.7.
  • Eleven of 139 tool calls errored during the run; all 115 model requests returned HTTP 200.
  • This is an environment-improvement task over a fixed starter, so its figures are not comparable to the archive's from-scratch build rows.
  • The archive is a delta over the committed fixture, not a standalone runnable tree.
  • No frame-rate receipt exists — neither the model nor the operator recorded a valid FPS sample.
  • No blind evaluation and no scored side-by-side against the reference run or the other two candidates are recorded.
  • The sanitized session export named in receipt.sha256 is withheld from the public archive by its own privacy label.
Visible evidence gaps
  • blind-evaluation record
  • standardized FPS and hitch receipt under sustained play
  • scored side-by-side against the reference run's judging anchors
  • reconciliation of the root ledger-vs-transcript token gap
Builder test available

This result is part of a Builder test. Open it for the exact prompt and any released projects, RemakeBench Harness workflows and production skills. Public proof and known evidence gaps stay visible here.

  • Fighter Game v1 · Builder projects v1
RemakeBenchResearch console