Measured harness ledgerPublic result
Qwen3.8 MaxJRPG boss battle — Qwen3.8 Max xhigh
Build an interactive Three.js JRPG boss battle with supplied character assets, required combat beats, animation, sound controls, effects, and a validator-ready game loop.
xhigh reasoningHeadline result
- Workflow cost
- ¥43.72 Independently recomputed first-party API-list-price equivalent for the measured token mix; NOT a billed amount. This run used an Alibaba Model Studio token-plan subscription, under which usage is not itemized as cash.
- Wall-clock
- 1h 42m 47.5s wall-clock
- Processed tokens
- 22.50M processed
- Record state
- published_artifact_validator_runtime_ledger_partial_metrics
Public summary
Qwen3.8 Max xhigh published_artifact_validator_runtime_ledger_partial_metrics ledger: 1h 42m 47.5s wall-clock, 22.50M processed, and ¥43.72 Independently recomputed first-party API-list-price equivalent for the measured token mix; NOT a billed amount. This run used an Alibaba Model Studio token-plan subscription, under which usage is not itemized as cash..
Run identity and stack
- Result ID: jrpg-boss-battle-qwen3.8-max-xhigh
- Technical model: qwen3.8-max
- Provider: Alibaba Cloud Model Studio token plan
- Client: Qwen Code 0.21.6 installed / 0.21.6 transcript-recorded
- Stack: Alibaba Cloud Model Studio token plan
- Stack: Qwen Code 0.21.6 installed / 0.21.6 transcript-recorded
- Stack: Technical model/configuration: qwen3.8-max
- Stack: Three.js / Vite game
- Stack: Supplied GLB fixture bundle
- Stack: Harness v1 JRPG validator
Cost basis
- This row's headline uses the cache-hit rule, matching the Stage 2 Landfall row. The Stage 1 STARFALL row headlines its full-input-rate upper bound instead. Convert before comparing.
- Qwen Code records no per-request cost field; null by construction.
- ¥275.64 if the 21,474,232 cache-read tokens were priced at the full ¥12/M input rate instead of the 10% cache-hit rule. Computed here for comparability; the receipt itself reports only the cache-hit-rule total.
Primary artifact integrity
- Kind: vite-threejs-jrpg-source-entry-point
- Path: artifacts/jrpg-boss-battle-qwen3.8-max-xhigh/source/src/main.js
- SHA-256: ea7505d8891a5c5c47e4b328508ff854992c5ef7ed06656bc3934a44595c21f0
Validation evidence
- Result: PASS — 41/41
- Path: artifacts/jrpg-boss-battle-qwen3.8-max-xhigh/evidence/validator/scorecard.sanitized.json
- SHA-256: 5da85956549970f9606d956a6b7d7708d9e50ed72d70f4ce00402dc9ec957da8
Recorded caveats
- Metrics status is PARTIAL: the root ledger and root transcript token sums disagree by 24,235 fresh input, 455,188 cache read, 20,886 visible output, and 9,379 reasoning tokens. The ledger figures are reported and the gap is disclosed, not repaired.
- Wall-clock (6,167.5 s) is end-to-end workflow latency including tool execution, installs, and browser checks — not model-only compute. Summed API duration (7,220.2 s) is separate and exceeds it because subagents ran concurrently.
- One Qwen Code agent turn fans out to many billable requests and subagent sessions under a single session id (171 root + 39 subagent here).
- Visible output and reasoning are separate fields; both are output-priced on this stack.
- inputTokens includes cached tokens and outputTokens includes reasoning tokens, so the ledger's totalTokens and total_processed are different quantities and must never be swapped.
- Qwen Code records no cache-write counter, so that line item is unavailable rather than zero.
- The cached-input rate comes from the pricing page's generic 10% cache-hit rule; no model-specific cached rate is published.
- Context occupancy is a single-request footprint, not cumulative usage; quota fields are null with no explicit before/after evidence.
- The client version is reported as-is: installed and transcript-recorded both read 0.21.6 while the expected stack listed 0.21.7.
- Seven of 197 tool calls errored during the run; all 210 model requests returned HTTP 200.
- The 60.2 fps figure is a headless validator sample on one machine, not a standardized cross-machine comparison.
- The scorecard is the run's own validator invocation; no independent operator rerun is recorded.
- The sanitized session export named in receipt.sha256 is withheld from the public archive by its own privacy label.
Visible evidence gaps
- blind-evaluation record
- independent operator rerun of the contract validator
- reconciliation of the root ledger-vs-transcript token gap
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
