Measured harness ledgerPublic result
Grok 4.6Dex Cube — 3 Move — Grok 4.6 xhigh
Complete the three-move Rubik's cube sequence D, F', U2 with regrasp, using two dexterous hands in MuJoCo.
xhigh reasoningHeadline result
- Workflow cost
- $4.55
- Wall-clock
- 46m 37s wall-clock
- Processed tokens
- Not recorded
- Record state
- judge_fail_independently_replayed
Public summary
Grok 4.6 xhigh judge_fail_independently_replayed ledger: 46m 37s wall-clock, Not recorded, and $4.55 Provider-recorded usage estimate, not an itemized subscription cash charge.
Run identity and stack
- Result ID: dex-cube-turn-mujoco-stage-2-three-moves-grok-4.6-xhigh
- Technical model: Grok 4.6
- Provider: xAI Grok
- Client: grok-sub wrapper (detached, headless), grok-4.6 at xhigh
- Stack: xAI Grok
- Stack: grok-sub wrapper (detached, headless), grok-4.6 at xhigh
- Stack: Technical model/configuration: Grok 4.6
- Stack: MuJoCo dexterous-hand fixture
- Stack: Harness cube-turn stage-2 prompt
- Stack: Requested tool profile: python-mujoco-headless
Cost basis
- usage cost reported by xAI for the session (grok-sub receipt line)
Primary artifact integrity
- Kind: open-loop-1khz-actuator-log
- Path: artifacts/dex-cube-turn-mujoco-stage-2-three-moves-grok-4.6-xhigh/evidence/model-supplied/ctrl.npy
- SHA-256: 4d4826e4017c9227402a63929d58030d7ab149c4135acf2855f1ad8bc5cec56c
Validation evidence
- Result: PASS
- Path: artifacts/dex-cube-turn-mujoco-stage-2-three-moves-grok-4.6-xhigh/validation/judge-replay.public.json
- SHA-256: edf57c9ec8b2f3afb863f7149aede7cad82ebd75ab57501477d2db50a05c9c09
- Validator SHA-256: 329a7f94ec9d9a37edbb94b6fcad5a2e48e32aca160e93e683f5de2763f4ab34
Recorded caveats
- The judge measures task completion only (facelet state, alignment, contact pads, runtime limits); there is no visual quality or blind evaluation for this task.
- The replay clip is rendered by the operator's viewer from the judge's trajectory export with a fixed camera; it is presentation evidence, not a measurement.
- provenance/orchestrator-results.md is the orchestrator's own result table copied verbatim and retains the operator's workspace paths.
- Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
- Grok token counts are not available from the grok-sub receipt; only the xAI usage dollars, elapsed seconds and turn count were recorded.
- The single judge run is byte-identical to the final submission, so it collapses into the final segment of the progression clip; the solver never reached a detented D turn.
- Reconstructed attempts are this run's own progression, recovered and judged after the fact by the archive operator. They are unscored evidence: they do not change this run's verdict or cohort standing and are excluded from every comparison, including any of them that the judge passes.
Visible evidence gaps
- hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
- token-count receipt (xAI usage dollars only)
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.