Measured harness ledgerPublic result
Grok 4.7

Eiffel Tower Drawing — Grok 4.7 xhigh

Draw the Eiffel Tower with a friction-held pencil on a Marvin arm and Wuji hand in MuJoCo.

xhigh reasoningHeadline result
Workflow cost
$2.04
Wall-clock
19m 18s wall-clock
Processed tokens
1.98M processed
Record state
judge_pass_independently_replayed
Public summary

Grok 4.7 xhigh judge_pass_independently_replayed ledger: 19m 18s wall-clock, 1.98M processed, and $2.04 Provider-recorded usage estimate, not an itemized subscription cash charge.

Run identity and stack
  • Result ID: dex-draw-mujoco-grok-4.7-xhigh-2026-09-22
  • Technical model: Grok 4.7
  • Provider: xAI Grok
  • Client: grok-sub wrapper (detached, headless, sandbox off), grok-4.7 at xhigh; ran to its own end_turn
  • Stack: xAI Grok
  • Stack: grok-sub wrapper (detached, headless, sandbox off), grok-4.7 at xhigh; ran to its own end_turn
  • Stack: Technical model/configuration: Grok 4.7
  • Stack: MuJoCo pencil-drawing fixture
  • Stack: Harness drawing prompt
  • Stack: Requested tool profile: python-mujoco-headless
Cost basis
  • usage cost reported by xAI for the session (grok-sub receipt line, confirmed by the wrapper run's meta.json)
Primary artifact integrity
  • Kind: open-loop-1khz-actuator-log
  • Path: artifacts/dex-draw-mujoco-grok-4.7-xhigh-2026-09-22/evidence/model-supplied/ctrl.npy
  • SHA-256: ff4045c4e0e85e5604e8aba8a2a3c4958bae2fbbbdb38cb778d59600e7ac398c
Validation evidence
  • Result: PASS
  • Path: artifacts/dex-draw-mujoco-grok-4.7-xhigh-2026-09-22/validation/judge-replay.public.json
  • SHA-256: de55e0708a56c02b71e0db75afe66d07de057ee99ccee3d123e5c88a1d9977c5
  • Validator SHA-256: 5c7af1a85044163afefca6abcd1ea9a19b109bf045ed8be40b4d5796098b738d
Recorded caveats
  • The judge measures task completion only; there is no visual quality or blind evaluation for this task.
  • The replay clips are rendered by the operator's viewer from the judge's own trajectory export with a fixed camera; they are presentation evidence, not a measurement.
  • Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
  • Cost is the xAI-reported session total from the grok-sub receipt line.
  • provenance/orchestrator-results.md is the task project's own result section copied verbatim with absolute local paths replaced.
  • The raw grok-sub receipts carry the session id, the reasoning transcript and local paths; they are not committed, only their sha256 commitments in provenance/source-receipt-hashes.json.
  • A Blender Cycles HERO RE-RENDER of this judged replay is pinned in provenance/clips.public.json as presentation: cycles-hero. Poses identical, visual layer only; pipeline ~/hero-render; two RTX 4090 Secure pods EU-RO-1 + EUR-IS-2, 52 + 154 min, $0.64 + $1.90 + $0.25 aborted batch = $2.79 total. FAIL episodes rendered as-is. It is not a new run, adds no measurement and cannot change a verdict. It is not committed; it is pinned by name, sha256, bytes, frames and duration.
  • Nudge disclosure: none.
  • The progression clip's burned-in caption says Final: FAIL because the compositor did not receive the PASS flag; the judged verdict is PASS tier 3.
  • The harness self-check archive under evidence/sim-attempts/ is committed exactly as the harness wrote it, so its attempt records still carry the operator's workspace path in their cwd field.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console