Measured harness ledgerPublic result
GPT-6 Astra

Apple Stem — GPT-6 Astra Max

Pinch an apple by the stem and lift it with a Wuji hand on an FR3 arm in SuperDex, without contacting the apple body.

Max reasoningHeadline result
Workflow cost
$3.48
Wall-clock
11m 05s wall-clock
Processed tokens
2.03M processed
Record state
judge_pass_independently_replayed
Public summary

GPT-6 Astra Max judge_pass_independently_replayed ledger: 11m 05s wall-clock, 2.03M processed, and $3.48 Standard API-equivalent estimate from the Codex wrapper receipt, not a subscription invoice: this was a ChatGPT-subscription run, $0 was charged and Codex emits no cost field..

Run identity and stack
  • Result ID: dex-apple-stem-superdex-gpt-6-astra-max-codex-2026-09-19
  • Technical model: GPT-6 Astra
  • Provider: OpenAI Codex
  • Client: codex-sub wrapper (detached, headless, -l dexapple-astra -m gpt-6-astra -e max -s danger-full-access -t 7200), Codex CLI 0.154.0; ran to its own end_turn
  • Stack: OpenAI Codex
  • Stack: codex-sub wrapper (detached, headless, -l dexapple-astra -m gpt-6-astra -e max -s danger-full-access -t 7200), Codex CLI 0.154.0; ran to its own end_turn
  • Stack: Technical model/configuration: GPT-6 Astra
  • Stack: SuperDex fingertip fixture
  • Stack: Harness apple-stem prompt
  • Stack: Requested tool profile: python-superdex-headless
Cost basis
  • actual marginal subscription charge usd: $0.00 USD.
Primary artifact integrity
  • Kind: open-loop-200hz-actuator-log
  • Path: artifacts/dex-apple-stem-superdex-gpt-6-astra-max-codex-2026-09-19/evidence/model-supplied/ctrl.npy
  • SHA-256: 3b90aff1b832c0edcff8cc9e8b5d6f88bd9b1f5db9435e83dd4e408d420013f1
Validation evidence
  • Result: PASS
  • Path: artifacts/dex-apple-stem-superdex-gpt-6-astra-max-codex-2026-09-19/validation/judge-replay.public.json
  • SHA-256: ddc75e946dcd051e2823770eb26da578f5d92696b9c459bfb682d285d1361b93
  • Validator SHA-256: c735d44ca72534f8058a7e1061ccf775f6e9f180cef26200d97a402f3de7403e
Recorded caveats
  • The judge measures task completion only; there is no visual quality or blind evaluation for this task.
  • The replay clips are rendered by the operator's viewer from the judge's own trajectory export with a fixed camera; they are presentation evidence, not a measurement.
  • Wall clock is end-to-end workflow latency including the solver's own simulation and judge runs, not model-only compute.
  • Output tokens include hidden reasoning, visible prose and tool-call JSON; cache reads are discounted, so processed-token volume overstates effective cost.
  • This was a ChatGPT-subscription Codex run, so there is no dollar receipt: $0 was charged and Codex emits no cost field. The dollar figure is an API-equivalent estimate at the GPT-6 Astra list rates already on record in this harness, not an invoice.
  • provenance/orchestrator-results.md is the task project's own result section copied verbatim with absolute local paths replaced.
  • The harness self-check archive under evidence/sim-attempts/ is committed exactly as the harness wrote it, so its attempt records still carry the operator's workspace path in their cwd field.
  • The raw Codex receipts (events.jsonl, meta.json, prompt and final-message files) carry the session id, the full reasoning transcript and local paths; they are not committed, only their sha256 commitments in provenance/source-receipt-hashes.json.
  • A Blender Cycles HERO RE-RENDER of this judged replay was added on 2026-09-20 (apple-astra-hero.mp4, pinned in provenance/clips.public.json as presentation: cycles-hero). It restages the SAME judged trajectory frame for frame on the same fixed camera - max pose error 4.77e-07, one float32 ulp at this magnitude - in a photo-real studio look. It is a visual layer only: it is not a new run, it adds no measurement and nothing judged changes. It is not committed; it is pinned by name, sha256, bytes, frames and duration.
Visible evidence gaps
  • hardware and interpreter receipt from the solver session (only the operator's replay environment is recorded)
Public result only

This result keeps its public summary and evidence, but it does not currently have a matching Builder test with prompts, projects, Harness workflows or skills.

RemakeBenchResearch console