Measured harness ledgerPublic result
GLM 5.3 Flash

Eiffel Tower Drawing — GLM 5.3 Flash Highest

Draw the Eiffel Tower with a friction-held pencil on a Marvin arm and Wuji hand in MuJoCo.

Highest reasoningHeadline result
Workflow cost
$0.12
Wall-clock
10m 28.2s wall-clock
Processed tokens
2.29M processed
Record state
judge_pass_independently_replayed
Public summary

GLM 5.3 Flash Highest judge_pass_independently_replayed ledger: 10m 28.2s wall-clock, 2.29M processed, and $0.12 Provider-recorded usage estimate, not an itemized subscription cash charge.

Run identity and stack
  • Result ID: dex-draw-mujoco-glm-5.3-flash-max-opencode-go-2026-09-16
  • Technical model: GLM 5.3 Flash
  • Provider: OpenCode Go
  • Client: OpenCode 1.18.23
  • Variant: max
  • Stack: OpenCode Go
  • Stack: OpenCode 1.18.23
  • Stack: Technical model/configuration: GLM 5.3 Flash
  • Stack: Requested variant: max
  • Stack: MuJoCo pencil-drawing fixture
  • Stack: Harness drawing prompt
  • Stack: Requested tool profile: python-mujoco-headless
Primary artifact integrity
  • Kind: open-loop-1khz-actuator-log
  • Path: artifacts/dex-draw-mujoco-glm-5.3-flash-max-opencode-go-2026-09-16/evidence/model-supplied/ctrl.npy
  • SHA-256: ebe5660b8b6edeb7317ea95d46c9daa19487b7429e49cec3a26b26bfec55484e
Validation evidence
  • Result: pass
  • Path: artifacts/dex-draw-mujoco-glm-5.3-flash-max-opencode-go-2026-09-16/validation/judge-replay.public.json
  • SHA-256: 467dbd4fed3d926dfca9098828dfd6c339866c417d22695406c05b4f743ed9d5
  • Validator SHA-256: 5c7af1a85044163afefca6abcd1ea9a19b109bf045ed8be40b4d5796098b738d
Recorded caveats
  • The judge measures task completion only (contact-gated ink against the given polyline, runtime limits); there is no visual quality or blind evaluation for this task.
  • The replay and progression clips are rendered by the operator's viewer from the judge's trajectory export and ink stream with a fixed camera; they are presentation evidence, not a measurement.
  • usedThreeFingers / threeFingerFraction are reported, never gating.
  • Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute, and all four runs of this batch shared one Mac.
  • The given harness exposes the judge's own simulate/rasterise/score pipeline, so the solver could self-score in-process; wall clock and cost are NOT comparable with the dex-cube-turn-mujoco cohorts.
  • This run is on fixture revision v1.1 (harness simulation archiving only). Verdicts stay comparable with the three v1 runs; scene_sha256 differs because it hashes harness.py.
  • One OpenCode agent turn contains many billable model requests; visible output and reasoning are separate output-priced fields; cache reads are deeply discounted so total processed tokens overstate cost.
  • The OpenCode-recorded figure is API-equivalent usage accounting under OpenCode Go, not an itemized marginal cash charge.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console