Measured harness ledgerPublic result
GPT-6 Astra (low) planner + Jev 1.13 actor

Apple Carry — GPT-6 Astra (low) planner + Jev 1.13 actor Unreported

Carry an apple onto a plate on two embodiments: an xArm7 in OpenRoboto's MuJoCo scene, and an FR3 with a Robotiq 2F-85 on SuperDex.

Unreported reasoningHeadline result
Workflow cost
$3.83
Wall-clock
Not recorded
Processed tokens
Not recorded
Record state
judge_pass_independently_replayed
Public summary

GPT-6 Astra (low) planner + Jev 1.13 actor Unreported judge_pass_independently_replayed ledger: Not recorded wall-clock, Not recorded, and $3.83 Recorded provider cost.

Run identity and stack
  • Result ID: apple-carry-fr3-gripper-astra-planner-jev-2026-09-21
  • Technical model: GPT-6 Astra (low) planner + Jev 1.13 actor
  • Provider: Provider not separately recorded
  • Stack: Technical model/configuration: GPT-6 Astra (low) planner + Jev 1.13 actor
  • Stack: Embodied manipulation fixture
  • Stack: Harness apple-carry prompt
Validation evidence
  • Path: artifacts/apple-carry-fr3-gripper-astra-planner-jev-2026-09-21/validation/judge-replay.public.json
  • SHA-256: 77d2ac101dca03a79e68806aa11ec2d097d03d0956d9c7ee6cb434900b2c259b
Recorded caveats
  • NOT COMPARABLE WITH THE CODING-AGENT COHORTS OR WITH BARE DECISION-LOOP ROWS. This is a jev-astra-hybrid run: GPT-6 Astra and Jev 1.13 were coupled under a protocol that is not the one-shot locked-prompt coding-agent track and not the two-call-per-cycle decision-loop track. cohort_eligible is false. Harry's criterion (same verdict as Astra-alone AND cheaper on API-equivalent) is recorded in the track README; it is not a league-table standing.
  • Criterion NOT met on API-equivalent ($0.77 vs $0.54/seed). Met on charged $ only because of the subscription.
  • Cycles hero not-rendered: the operator only renders Cycles heroes for results that beat Astra-alone on cost, and apple-carry Arm C did not ($0.77 vs $0.54/seed API-equivalent).
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console