Measured harness ledgerPublic result
GLM 5.3 FlashCampfire under the stars — GLM 5.3 Flash Max
Build a polished interactive Three.js campfire scene under a starry night with convincing fire, lighting, environmental detail, and camera motion.
Max reasoningHeadline result
- Workflow cost
- $0.11
- Wall-clock
- 2m 40.772s wall-clock
- Processed tokens
- 90K processed
- Record state
- partial_one_shot_runtime_visual_failure
Public summary
GLM 5.3 Flash Max partial_one_shot_runtime_visual_failure ledger: 2m 40.772s wall-clock, 90K processed, and $0.11 OpenCode-recorded API-equivalent usage accounting, not an itemized subscription marginal-cash invoice.
Run identity and stack
- Result ID: campfire-threejs-glm-5.3-max
- Technical model: glm-5.3-flash
- Provider: OpenCode Go
- Client: OpenCode
- Stack: OpenCode Go
- Stack: OpenCode
- Stack: Technical model/configuration: glm-5.3-flash
- Stack: Three.js interactive scene
- Stack: Harness v1 campfire prompt
Primary artifact integrity
- Kind: interactive-threejs-campfire-scene
- Path: artifacts/campfire-threejs-glm-5.3-max/source/index.html
- SHA-256: 408e21d9bb66dcc61eff7eabb49a5fc53966e1540e893badde4b4eb67b78d604
Validation evidence
- Result: FAIL_VISUAL_VALIDITY
Recorded caveats
- The benchmark creation window is one-shot, but a later exact continue request after the metrics boundary is disclosed and excluded because it made no source change.
- Wall-clock is end-to-end workflow latency, not model-only compute.
- The local FPS result is performance-only and must not be interpreted as visual acceptance.
- The source requires network access to pinned unpkg.com Three.js modules.
- Private OpenCode receipts, exports, and the model-specific operator prompt are intentionally not published because they include local operational identifiers or paths.
Visible evidence gaps
- A visually valid model-authored Campfire render, regenerated in a new bounded benchmark run
- Independent browser capture demonstrating a legible campfire after that regeneration
- Blind-evaluation record
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.
