Measured harness ledgerPublic result
GPT-6 Astra (low) planner + Jev 1.13 actorEiffel Tower Drawing — GPT-6 Astra (low) planner + Jev 1.13 actor Unreported
Draw the Eiffel Tower with a friction-held pencil on a Marvin arm and Wuji hand in MuJoCo.
Unreported reasoningHeadline result
- Workflow cost
- $0.18
- Wall-clock
- Not recorded
- Processed tokens
- Not recorded
- Record state
- judge_pass_independently_replayed
Public summary
GPT-6 Astra (low) planner + Jev 1.13 actor Unreported judge_pass_independently_replayed ledger: Not recorded wall-clock, Not recorded, and $0.18 Recorded provider cost.
Run identity and stack
- Result ID: dex-draw-mujoco-astra-planner-jev-2026-09-21
- Technical model: GPT-6 Astra (low) planner + Jev 1.13 actor
- Provider: OpenAI Codex (planner) + TypeSafe via OpenRouter (Jev)
- Stack: OpenAI Codex (planner) + TypeSafe via OpenRouter (Jev)
- Stack: Technical model/configuration: GPT-6 Astra (low) planner + Jev 1.13 actor
- Stack: MuJoCo pencil-drawing fixture
- Stack: Harness drawing prompt
Validation evidence
- Path: artifacts/dex-draw-mujoco-astra-planner-jev-2026-09-21/validation/judge-replay.public.json
- SHA-256: 130c53c2ca6d464cbf93fc58fb0b7a44417899285c83a519022b732c70cf6d2c
Recorded caveats
- NOT COMPARABLE WITH THE CODING-AGENT COHORTS OR WITH BARE DECISION-LOOP ROWS. This is a jev-astra-hybrid run: GPT-6 Astra and Jev 1.13 were coupled under a protocol that is not the one-shot locked-prompt coding-agent track and not the two-call-per-cycle decision-loop track. cohort_eligible is false. Harry's criterion (same verdict as Astra-alone AND cheaper on API-equivalent) is recorded in the track README; it is not a league-table standing.
- Astra-alone (low) coding-agent baseline is registered (2026-09-21). Arm C is cheaper than both Astra-max and Astra-low on this task.
- A Blender Cycles HERO RE-RENDER of this judged replay is pinned in provenance/clips.public.json as presentation: cycles-hero. Poses identical, visual layer only; pipeline ~/hero-render; one RTX 4090 session, 130 min, $1.61 + $0.04 discarded pod. It is not a new run, adds no measurement and cannot change a verdict. It is not committed; it is pinned by name, sha256, bytes, frames and duration.
Public result only
This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.