Measured harness ledgerPublic result
Grok 4.6

Claw Unlock — Grok 4.6 xhigh

Find a collision-free rigid motion that unlocks two interlocked wire claws using pure geometry.

xhigh reasoningHeadline result
Workflow cost
$6.15
Wall-clock
72m 45s wall-clock
Processed tokens
19.87M processed
Record state
judge_fail_independently_replayed
Public summary

Grok 4.6 xhigh judge_fail_independently_replayed ledger: 72m 45s wall-clock, 19.87M processed, and $6.15 Provider-recorded usage estimate, not an itemized subscription cash charge.

Run identity and stack
  • Result ID: claw-unlock-geometry-grok-4.6-xhigh
  • Technical model: Grok 4.6
  • Provider: xAI Grok
  • Client: grok-sub wrapper (detached, headless, -t 10800, sandbox off), grok-4.6 at xhigh; ran to its own end_turn
  • Stack: xAI Grok
  • Stack: grok-sub wrapper (detached, headless, -t 10800, sandbox off), grok-4.6 at xhigh; ran to its own end_turn
  • Stack: Technical model/configuration: Grok 4.6
  • Stack: python-fcl geometry fixture
  • Stack: Harness claw-unlock prompt
  • Stack: Requested tool profile: python-fcl-geometry-headless
Cost basis
  • usage cost reported by xAI for the session (grok-sub receipt line, confirmed by the wrapper run's meta.json)
Primary artifact integrity
  • Kind: se3-rigid-body-path-polyline
  • Path: artifacts/claw-unlock-geometry-grok-4.6-xhigh/evidence/model-supplied/path.npy
  • SHA-256: 37e95e90dac6bdadb6958e81cf6d6c1cf797e1592de9923c3a2b549f6f8fd30e
Validation evidence
  • Result: PASS
  • Path: artifacts/claw-unlock-geometry-grok-4.6-xhigh/validation/judge-replay.public.json
  • SHA-256: 62d855a16ef7b84038c3b04b3ffdf067e4dffc5c90775daa4d0b3b8bf71195c7
  • Validator SHA-256: ac0ad531aaa6a4f93213a4445af54d4bcd582fbe105ca369b8253050e49c24cd
Recorded caveats
  • The judge measures task completion only (exact mesh-mesh geometry along the interpolated path, convex-hull separation and the escape probe); there is no visual quality or blind evaluation for this task.
  • The replay clips are not committed to the repository: they are presentation evidence rendered by the operator's viewer from the judge's own trajectory export with a fixed camera, they add no measurement, and they are pinned by sha256 in provenance/clips.public.json. The two published clips (grok-rig.mp4, grok-rig-full.mp4) supersede the first delivery's undressed renders, which stay pinned there under their own names. Two of those undressed pins -- the colliding self-checks grok-sim-a003.mp4 and grok-sim-a006.mp4 -- were recomputed on 2026-09-20: a 2026-09-19 re-render had dropped their 1.5 s post-collision trim, so both were re-rendered with the trim restored from each attempt's own verdict.json (549 frames / 18.3 s and 216 frames / 7.2 s, the published framing, camera and encoder settings) and their commitments recomputed from the files on disk; the new h264 encodes are not byte-identical to the 2026-09-18 ones. The renderer's shared delivery index manifest.json is re-pinned with them. Nothing judged changes.
  • PRESENTATION ONLY -- the two published clips are rendered from a DRESSED replay: tools/dress_replay.py in the task project (operator-side, not committed here) adds two MuJoCo Menagerie ALOHA 2 arms (Trossen Robotics, BSD-3-Clause, upstream commit 8161bba264d7fa7c99ca301e91e7fb44737676ad), the table, the extrusion frame and the docks as a visual layer. Both pieces are carried by one rigid transform per frame, so the judged relative pose inv(pose(fixed)) @ pose(moving) is reproduced to 3.3e-16 per frame; the global carry is Qineng Wang's recorded fixed-piece trajectory for the same puzzle, mirrored, which is what keeps both grasps reachable; the arms are posed by damped least squares IK with a residual of at most 2.46e-07 m / 0.0331 deg and no frame at a joint limit; a colliding attempt is cut 1.5 s after the collision instant. The judged replay is unchanged and nothing in the dressing affects the verdict, the stage, any metric or the cohort.
  • provenance/gold-results.md is the task author's own reference-results file copied verbatim; provenance/orchestrator-results.md is the orchestrator's own result file with absolute local paths replaced.
  • minClearance is reported, never gating: every known solution to this puzzle rides the contact surface, so the figure says little about quality.
  • Wall clock is end-to-end workflow latency including the solver's own search and self-check runs, not model-only compute.
  • The given checker runs the judge's own judge() in-process, so the solver could self-score exactly; wall clock and cost are NOT comparable with the physics tasks in this suite.
  • The solver never invoked validate.py; the verdict of record is the operator's canonical judge run on the submitted path, and the solver's own ten self-checks agree with it on the submitted trajectory.
  • The two claw meshes carry no established redistribution licence (see given/ATTRIBUTION.md); no licence is claimed for them and they will be removed on request.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's judging environment is recorded)
  • a second cohort run: this is the only cohort ledger on the task so far, so there is no field to compare against
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console