Measured harness ledgerPublic result
GLM 5.3 Flash

Explorable space-flight game — GLM 5.3 Flash Max

Build a responsive browser space-flight game with flight controls, a coherent star-system environment, lighting, assets, and a playable game loop.

Max reasoningHeadline result
Workflow cost
$4.46
Wall-clock
48m 14.1s wall-clock
Processed tokens
14.15M processed
Record state
partial_multi_turn_token_timing_and_artifact_ledger
Public summary

GLM 5.3 Flash Max partial_multi_turn_token_timing_and_artifact_ledger ledger: 48m 14.1s wall-clock, 14.15M processed, and $4.46 UNVERIFIED reference scenario: GLM-5.2/GLM-5.1 official rates applied to GLM-5.3 token counts; not a verified GLM-5.3 charge or subscription marginal-cash invoice.

Run identity and stack
  • Result ID: space-flight-game-glm-5.3-max
  • Technical model: glm-5.3-flash
  • Provider: ZCode
  • Stack: ZCode
  • Stack: Technical model/configuration: glm-5.3-flash
  • Stack: Three.js / Vite game
  • Stack: Generated spaceship assets
  • Stack: Harness v1 space-flight prompt
Primary artifact integrity
  • Kind: interactive-threejs-space-flight-game
  • Path: artifacts/space-flight-game-glm-5.3-max/source/index.html
  • SHA-256: f9f0d336f956818cf50bdb7fe5887e0c8806e4fbc1ef9e18ecf27f5dbf4264c3
Recorded caveats
  • This is a two-turn, 13-user-message ZCode session, not a one-shot benchmark run.
  • Wall-clock is end-to-end workflow latency, not model-only compute.
  • GLM-5.3 had no official published API pricing at the access date. $4.4642 is a GLM-5.2/GLM-5.1-rate reference scenario and must not be compared as a verified GLM-5.3 cost.
  • The included images are model-supplied verification shots, not independent RemakeBench replay captures.
  • Static syntax and HTTP serving checks are not a browser/WebGL flight-control, visual-quality, or FPS measurement.
Visible evidence gaps
  • A clean one-shot generation window
  • Independent browser/WebGL runtime replay
  • local FPS with browser, hardware, and viewport receipt
  • RTX Pro 6000 final capture
  • blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console