Measured harness ledgerPublic result
MiMo V2.6 Pro

Claw Unlock — MiMo V2.6 Pro Highest

Find a collision-free rigid motion that unlocks two interlocked wire claws using pure geometry.

Highest reasoningHeadline result
Workflow cost
$0.32
Wall-clock
73m 25.6s wall-clock
Processed tokens
9.82M processed
Record state
judge_fail_independently_replayed
Public summary

MiMo V2.6 Pro Highest judge_fail_independently_replayed ledger: 73m 25.6s wall-clock, 9.82M processed, and $0.32 Provider-recorded usage estimate, not an itemized subscription cash charge.

Run identity and stack
  • Result ID: claw-unlock-geometry-mimo-v2.6-pro-max-opencode-go-2026-09-23
  • Technical model: Xiaomi MiMo v2.6 Pro
  • Provider: OpenCode Go
  • Client: OpenCode 1.18.23
  • Variant: max
  • Stack: OpenCode Go
  • Stack: OpenCode 1.18.23
  • Stack: Technical model/configuration: Xiaomi MiMo v2.6 Pro
  • Stack: Requested variant: max
  • Stack: python-fcl geometry fixture
  • Stack: Harness claw-unlock prompt
  • Stack: Requested tool profile: python-fcl-geometry-headless
Primary artifact integrity
  • Kind: se3-rigid-body-path-polyline
  • Path: artifacts/claw-unlock-geometry-mimo-v2.6-pro-max-opencode-go-2026-09-23/evidence/model-supplied/path.npy
  • SHA-256: 91c8d279ea81abc7987ecf97505d3f35dea986e9d828402834250d755c984461
Validation evidence
  • Result: pass
  • Path: artifacts/claw-unlock-geometry-mimo-v2.6-pro-max-opencode-go-2026-09-23/validation/judge-replay.public.json
  • SHA-256: a789ef2d68b1a75f4dc13b6011a263010a098c304ff1dbb1457e32ec67a42b89
  • Validator SHA-256: ac0ad531aaa6a4f93213a4445af54d4bcd582fbe105ca369b8253050e49c24cd
Recorded caveats
  • The judge measures task completion only; there is no visual quality or blind evaluation for this task.
  • The replay clips are rendered by the operator's viewer from the judge's own trajectory export with a fixed camera; they are presentation evidence, not a measurement.
  • Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
  • Cost is the OpenCode-recorded figure. OpenCode Go subscription ($0 charged). OpenRouter-list API-equivalent at $0.435 in / $0.87 out per MTok equals the recorded figure.
  • provenance/orchestrator-results.md is the task project's own result section copied with absolute local paths and session ids replaced.
  • The raw OpenCode receipts carry the session id, the reasoning transcript and local paths; they are not committed, only their sha256 commitments in provenance/source-receipt-hashes.json.
  • A Blender Cycles HERO RE-RENDER of this judged replay is pinned in provenance/clips.public.json as presentation: cycles-hero. Poses identical, visual layer only; pipeline ~/hero-render; two RTX 4090 Secure pods EUR-IS-1 + US, 62.7 + 77.4 min, $0.77 + $0.95, plus a third RTX 4090 Secure (US) for draw (~161 min, $1.99, operator-killed) = $3.71 total. FAIL episodes rendered as-is. It is not a new run, adds no measurement and cannot change a verdict. It is not committed; it is pinned by name, sha256, bytes, frames and duration.
  • Nudge disclosure: none.
  • The harness self-check archive under evidence/sim-attempts/ is committed exactly as the harness wrote it, so its attempt records still carry the operator's workspace path in their cwd field.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console