Measured harness ledgerPublic result
Grok 4.6

Dex Cube Turn — Grok 4.6 xhigh

Turn one face of a Rubik's cube a quarter turn by finger contact with two dexterous hands in MuJoCo.

xhigh reasoningHeadline result
Workflow cost
$1.22
Wall-clock
25m 22s wall-clock
Processed tokens
Not recorded
Record state
judge_pass_independently_replayed
Public summary

Grok 4.6 xhigh judge_pass_independently_replayed ledger: 25m 22s wall-clock, Not recorded, and $1.22 Provider-recorded usage estimate, not an itemized subscription cash charge.

Run identity and stack
  • Result ID: dex-cube-turn-mujoco-grok-4.6-xhigh
  • Technical model: Grok 4.6
  • Provider: xAI Grok
  • Client: grok-sub wrapper (detached, headless), grok-4.6 at xhigh
  • Stack: xAI Grok
  • Stack: grok-sub wrapper (detached, headless), grok-4.6 at xhigh
  • Stack: Technical model/configuration: Grok 4.6
  • Stack: MuJoCo dexterous-hand fixture
  • Stack: Harness cube-turn prompt
  • Stack: Requested tool profile: python-mujoco-headless
Cost basis
  • usage cost reported by xAI for the session (grok-sub receipt line)
Primary artifact integrity
  • Kind: open-loop-1khz-actuator-log
  • Path: artifacts/dex-cube-turn-mujoco-grok-4.6-xhigh/evidence/model-supplied/ctrl.npy
  • SHA-256: 982fb1f9803d60fa13d62e87506c0ee73099d44eca8063b47031af0612487ae7
Validation evidence
  • Result: PASS
  • Path: artifacts/dex-cube-turn-mujoco-grok-4.6-xhigh/validation/judge-replay.public.json
  • SHA-256: f9ff741114589717b51dc887e494ff29e8d00bb4430d7f45f011c3df0795a055
  • Validator SHA-256: 8425328aed969dfd11cc1e5107b1e0c83b888c59675bdf2ce86c4bc2878dcee9
Recorded caveats
  • The judge measures task completion only (facelet state, alignment, contact pads, runtime limits); there is no visual quality or blind evaluation for this task.
  • The replay clip is rendered by the operator's viewer from the judge's trajectory export with a fixed camera; it is presentation evidence, not a measurement.
  • provenance/orchestrator-results.md is the orchestrator's own result table copied verbatim and retains the operator's workspace paths.
  • Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
  • Grok token counts are not available from the grok-sub receipt; only the xAI usage dollars, elapsed seconds and turn count were recorded.
  • Both tier-1 cohort solutions exploit the pre-positioned D grasp; tier 1 does not test finger placement, which is why the stage-2 task exists. Do not compare the tiers.
  • Reconstructed attempts are this run's own progression, recovered and judged after the fact by the archive operator. They are unscored evidence: they do not change this run's verdict or cohort standing and are excluded from every comparison, including any of them that the judge passes.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
  • token-count receipt (xAI usage dollars only)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console