Measured harness ledgerPublic result
Grok 4.6Eiffel Tower Drawing — Grok 4.6 xhigh
Draw the Eiffel Tower with a friction-held pencil on a Marvin arm and Wuji hand in MuJoCo.
xhigh reasoningHeadline result
- Workflow cost
- $0.31
- Wall-clock
- 10m 23s wall-clock
- Processed tokens
- Not recorded
- Record state
- judge_pass_independently_replayed
Public summary
Grok 4.6 xhigh judge_pass_independently_replayed ledger: 10m 23s wall-clock, Not recorded, and $0.31 Provider-recorded usage estimate, not an itemized subscription cash charge.
Run identity and stack
- Result ID: dex-draw-mujoco-grok-4.6-xhigh
- Technical model: Grok 4.6
- Provider: xAI Grok
- Client: grok-sub wrapper (detached, headless, -t 7200), grok-4.6 at xhigh; ran to its own end_turn
- Stack: xAI Grok
- Stack: grok-sub wrapper (detached, headless, -t 7200), grok-4.6 at xhigh; ran to its own end_turn
- Stack: Technical model/configuration: Grok 4.6
- Stack: MuJoCo pencil-drawing fixture
- Stack: Harness drawing prompt
- Stack: Requested tool profile: python-mujoco-headless
Cost basis
- usage cost reported by xAI for the session (grok-sub receipt line)
Primary artifact integrity
- Kind: open-loop-1khz-actuator-log
- Path: artifacts/dex-draw-mujoco-grok-4.6-xhigh/evidence/model-supplied/ctrl.npy
- SHA-256: 2b8bcaab829e4d8bf1bf16adacaf03f5203ddbf5202c682726abdba6e3e9e497
Validation evidence
- Result: PASS
- Path: artifacts/dex-draw-mujoco-grok-4.6-xhigh/validation/judge-replay.public.json
- SHA-256: a15a3050d1afe93de28fbec1226530bbee4d225aaebb3d559773b5aff3375530
- Validator SHA-256: 5c7af1a85044163afefca6abcd1ea9a19b109bf045ed8be40b4d5796098b738d
Recorded caveats
- The judge measures task completion only (contact-gated ink against the given polyline, runtime limits); there is no visual quality or blind evaluation for this task.
- The replay clip is rendered by the operator's viewer from the judge's trajectory export and ink stream with a fixed camera; it is presentation evidence, not a measurement.
- provenance/gold-results.md is the orchestrator's own result file copied verbatim.
- usedThreeFingers / threeFingerFraction are reported, never gating.
- Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
- The given harness exposes the judge's own simulate/rasterise/score pipeline, so the solver could self-score in-process; wall clock and cost are NOT comparable with the dex-cube-turn-mujoco cohorts.
- Grok token counts are not available from the grok-sub receipt; only the xAI usage dollars, elapsed seconds and turn count were recorded.
- The spire stroke was skipped after the 74 s cap, so the log ends 0.8 s short of MAX_STEPS with 13 of 14 strokes; tier 3 still passed on coverage 0.950.
Visible evidence gaps
- hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
- token-count receipt (xAI usage dollars only)
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.