Measured harness ledgerPublic result
Grok 4.6

Dex Cube — 3 Move — Grok 4.6 xhigh

Complete the three-move Rubik's cube sequence D, F', U2 with regrasp, using two dexterous hands in MuJoCo.

xhigh reasoningHeadline result
Workflow cost
$4.55
Wall-clock
46m 37s wall-clock
Processed tokens
Not recorded
Record state
judge_fail_independently_replayed
Public summary

Grok 4.6 xhigh judge_fail_independently_replayed ledger: 46m 37s wall-clock, Not recorded, and $4.55 Provider-recorded usage estimate, not an itemized subscription cash charge.

Run identity and stack
  • Result ID: dex-cube-turn-mujoco-stage-2-three-moves-grok-4.6-xhigh
  • Technical model: Grok 4.6
  • Provider: xAI Grok
  • Client: grok-sub wrapper (detached, headless), grok-4.6 at xhigh
  • Stack: xAI Grok
  • Stack: grok-sub wrapper (detached, headless), grok-4.6 at xhigh
  • Stack: Technical model/configuration: Grok 4.6
  • Stack: MuJoCo dexterous-hand fixture
  • Stack: Harness cube-turn stage-2 prompt
  • Stack: Requested tool profile: python-mujoco-headless
Cost basis
  • usage cost reported by xAI for the session (grok-sub receipt line)
Primary artifact integrity
  • Kind: open-loop-1khz-actuator-log
  • Path: artifacts/dex-cube-turn-mujoco-stage-2-three-moves-grok-4.6-xhigh/evidence/model-supplied/ctrl.npy
  • SHA-256: 4d4826e4017c9227402a63929d58030d7ab149c4135acf2855f1ad8bc5cec56c
Validation evidence
  • Result: PASS
  • Path: artifacts/dex-cube-turn-mujoco-stage-2-three-moves-grok-4.6-xhigh/validation/judge-replay.public.json
  • SHA-256: edf57c9ec8b2f3afb863f7149aede7cad82ebd75ab57501477d2db50a05c9c09
  • Validator SHA-256: 329a7f94ec9d9a37edbb94b6fcad5a2e48e32aca160e93e683f5de2763f4ab34
Recorded caveats
  • The judge measures task completion only (facelet state, alignment, contact pads, runtime limits); there is no visual quality or blind evaluation for this task.
  • The replay clip is rendered by the operator's viewer from the judge's trajectory export with a fixed camera; it is presentation evidence, not a measurement.
  • provenance/orchestrator-results.md is the orchestrator's own result table copied verbatim and retains the operator's workspace paths.
  • Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
  • Grok token counts are not available from the grok-sub receipt; only the xAI usage dollars, elapsed seconds and turn count were recorded.
  • The single judge run is byte-identical to the final submission, so it collapses into the final segment of the progression clip; the solver never reached a detented D turn.
  • Reconstructed attempts are this run's own progression, recovered and judged after the fact by the archive operator. They are unscored evidence: they do not change this run's verdict or cohort standing and are excluded from every comparison, including any of them that the judge passes.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
  • token-count receipt (xAI usage dollars only)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console