Measured harness ledgerPublic result
GLM 5.3 Flash

Campfire under the stars — GLM 5.3 Flash Max

Build a polished interactive Three.js campfire scene under a starry night with convincing fire, lighting, environmental detail, and camera motion.

Max reasoningHeadline result
Workflow cost
$0.11
Wall-clock
2m 40.772s wall-clock
Processed tokens
90K processed
Record state
partial_one_shot_runtime_visual_failure
Public summary

GLM 5.3 Flash Max partial_one_shot_runtime_visual_failure ledger: 2m 40.772s wall-clock, 90K processed, and $0.11 OpenCode-recorded API-equivalent usage accounting, not an itemized subscription marginal-cash invoice.

Run identity and stack
  • Result ID: campfire-threejs-glm-5.3-max
  • Technical model: glm-5.3-flash
  • Provider: OpenCode Go
  • Client: OpenCode
  • Stack: OpenCode Go
  • Stack: OpenCode
  • Stack: Technical model/configuration: glm-5.3-flash
  • Stack: Three.js interactive scene
  • Stack: Harness v1 campfire prompt
Primary artifact integrity
  • Kind: interactive-threejs-campfire-scene
  • Path: artifacts/campfire-threejs-glm-5.3-max/source/index.html
  • SHA-256: 408e21d9bb66dcc61eff7eabb49a5fc53966e1540e893badde4b4eb67b78d604
Validation evidence
  • Result: FAIL_VISUAL_VALIDITY
Recorded caveats
  • The benchmark creation window is one-shot, but a later exact continue request after the metrics boundary is disclosed and excluded because it made no source change.
  • Wall-clock is end-to-end workflow latency, not model-only compute.
  • The local FPS result is performance-only and must not be interpreted as visual acceptance.
  • The source requires network access to pinned unpkg.com Three.js modules.
  • Private OpenCode receipts, exports, and the model-specific operator prompt are intentionally not published because they include local operational identifiers or paths.
Visible evidence gaps
  • A visually valid model-authored Campfire render, regenerated in a new bounded benchmark run
  • Independent browser capture demonstrating a legible campfire after that regeneration
  • Blind-evaluation record
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console