Measured harness ledgerPublic result
Qwen3.8 Max PreviewExplorable space-flight game — Qwen3.8 Max Preview Unreported · Attempt 3
Build a responsive browser space-flight game with flight controls, a coherent star-system environment, lighting, assets, and a playable game loop.
Unreported reasoningHeadline result
- Workflow cost
- $45.89
- Wall-clock
- 2h 4m 17.3s wall-clock
- Processed tokens
- 29.90M processed
- Record state
- partial_token_timing_runtime_artifact_visual_self_verification_ledger
Public summary
Qwen3.8 Max Preview Unreported partial_token_timing_runtime_artifact_visual_self_verification_ledger ledger: 2h 4m 17.3s wall-clock, 29.90M processed, and $45.89 Qualified third-party NanoGPT API-list-price scenario for the measured token mix; not an Alibaba Token Plan cash charge.
Run identity and stack
- Result ID: space-flight-game-qwen3.8-max-preview-attempt-3-isolated-vision-pass
- Technical model: qwen3.8-max-preview
- Provider: Alibaba Cloud Model Studio
- Client: Qwen Code 0.20.0
- Attempt: 3
- Stack: Alibaba Cloud Model Studio
- Stack: Qwen Code 0.20.0
- Stack: Technical model/configuration: qwen3.8-max-preview
- Stack: Attempt 3
- Stack: Three.js / Vite game
- Stack: Generated spaceship assets
- Stack: Harness v1 space-flight prompt
Cost basis
- No separate first-party Qwen3.8 cache rate was established, so cached input uses the published third-party input rate. This is not Alibaba pricing, provider cost, or an itemized subscription charge.
- Qualified API-list-price equivalent.
Primary artifact integrity
- Kind: threejs-space-flight-game-source-entry-point
- Path: artifacts/space-flight-game-qwen3.8-max-preview-attempt-3-isolated-vision-pass/source/src/main.js
- SHA-256: 2f921e8c2e4e82092579563b295a36f7dc341f3accdf46d2d99c6ccf07f0add8
Recorded caveats
- Wall-clock is end-to-end workflow latency, not model-only compute time.
- Output tokens include Qwen-accounted thoughts/reasoning, visible prose/code, and tool-related output; 185,832 thought tokens are a subset of 296,697 output tokens.
- The actual captures are 1920×993; 1920×1080 was the requested Chrome window size.
- The 120–123 FPS values are HUD-reported on Apple M1 Pro / ANGLE Metal, not standardized cross-machine performance benchmarks.
- The image-isolation architecture succeeded but exceeded its intended agent-call budget, making this successful retry unusually expensive.
- No task-isolated Token Plan Credit charge can be assigned because the meter combines attempts 2 and 3.
- Formal RemakeBench quality scoring, standardized RTX capture, and blind preference remain unassigned.
Visible evidence gaps
- bounded visual-inspector retry with six or fewer inspector calls and three or fewer capture/edit cycles
- standardized local FPS measurement receipt
- standardized RTX capture
- formal RemakeBench quality scoring
- blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
