# Benchmark subscription-usage estimate

**As of 25 September 2026.** This report estimates the share of a weekly subscription allowance consumed by the requested RemakeBench cohort. The percentages are **calibrated estimates**, not provider-reported per-benchmark charges or exact quota debits.

| Subscription / model | Measured cohort | Estimated share of one weekly allowance | Four identical cohorts |
| --- | ---: | ---: | ---: |
| OpenAI Pro 20x / GPT-6 Sol Max | 17 of 18 runs | **~6.55%** | ~26.19 percentage points of one weekly allowance |
| SuperGrok Heavy / Grok 4.7 XHigh | 17 of 18 runs | **~31.17%** | ~124.70 percentage points of one weekly allowance |

The four-cohort column adds four weekly debits and expresses the sum in **one-week-allowance units**. Against the capacity of *four* equal weekly allowances, the fractions remain ~6.55% and ~31.17%, respectively. It does not mean Grok necessarily permits a 124.70% debit in any single week.

## Scope and evidence

The requested 18-test grid comprises Courtyard Blossom Stage 2, JRPG Boss Battle, Neon Gunner, Sekiro Main Menu, Space Flight Stage 2: Landfall, Carve WebGPU Snowboarding, Turbofan, Bluetooth Speaker PCB, Infinite Cathedral, Neo-Gothic Storm City, Oasis Outpost, Jungle Temple, Apple Carry, Apple Stem, Claw Unlock, Dex Cube Turn, Dex Cube Three Moves, and Eiffel Drawing. The cohort is drawn from the [published RemakeBench results](https://www.remakebench.com/results) and local CLI usage records. Space Flight Stage 1 is a prerequisite for Stage 2 but is not a separate row in this grid; its independent cost is not added here. Retries or preparatory work outside selected result records may also be omitted.

**Turbofan is not measured in either 17-run total.** There is no final GPT-6 Sol result JSON for it. Grok has an artifact-level estimate, but no final CLI usage receipt; the optional projection below keeps that estimate separate. Claude is being assessed separately at the user's direction, and MiMo-v2.6 Pro used OpenRouter rather than either subscription, so neither is included.

## OpenAI Pro 20x / GPT-6 Sol Max

The 17 selected GPT-6 Sol results record **5,804,476 fresh input tokens**, **379,685,632 cached input tokens**, and **1,282,837 output tokens including reasoning**. Applying the published GPT-6 Sol credit-equivalent rates of 50, 5, and 250 credits per million tokens, respectively, gives:

```text
(5,804,476 × 50 + 379,685,632 × 5 + 1,282,837 × 250) / 1,000,000
= 2,509.36121 credit equivalents (~$100.37 at the listed credit conversion)
```

The user-provided Codex Usage Analytics screenshot shows **46.13% of the weekly limit** for the long-running “Find remaining DeepSeek V4 tests” task, with zero purchased credits used. That task includes orchestration, subagents, and models other than GPT-6 Sol; **46.13% is not the benchmark cohort's measured consumption**. Summing locally available, model-weighted token records for that task and its subagents gives **17,676.36806 credit equivalents**. If those records cover exactly the usage attributed to the screenshot, and if included-plan quota is proportional to these credit weights, the cohort estimate is:

```text
46.13% × 2,509.36121 / 17,676.36806 = 6.5487% of one weekly allowance
four identical cohorts = 10,037.44484 credit equivalents (~$401.50 API-list equivalent)
```

This is particularly conditional: the screenshot is a **seven-day analytics view**, whereas a quota resets on its own weekly boundary; some selected results predate the current quota window. Local task logs may not cover every item the dashboard attributes to the task. Credit-equivalent rates need not match the internal weighting of included Pro usage. The dollar values above are **API-list equivalents, not subscription cash charges**. See [OpenAI's pricing and usage explanation](https://learn.chatgpt.com/docs/pricing).

## SuperGrok Heavy / Grok 4.7 XHigh

The 17 selected Grok CLI result records contain **283,374,205 processed tokens** and **2,497,265 generated tokens**, with **$124.85507688 of provider-recorded API-equivalent usage**. This is a pricing proxy, not money billed on the subscription.

Read-only first-party Grok CLI meter snapshots for the same weekly window showed **56.0% used at 2026-09-23 18:22:15 UTC** and **75.0% used at 2026-09-25 09:09:24 UTC**. Deduplicated CLI receipts between those snapshots total **$76.09728984 API-equivalent usage**, approximately **96.1% XHigh**. Assuming that interval's 19-percentage-point increase came from the logged use and that quota debit scales comparably across the cohort:

```text
19 percentage points × $124.85507688 / $76.09728984
= 31.1739 percentage points of one weekly allowance
four identical cohorts = $499.42030752 API-equivalent usage
```

The unverified Turbofan artifact estimates another **$27.850909** at its midpoint. Including it would raise the projection to **~38.13% per cohort** and **~152.51 percentage points in one-week-allowance units** for four cohorts; the four-cohort API-equivalent total would be **~$610.82**. That is a scenario, **not an eighteenth measured receipt**.

Grok's weekly pool is shared across products, so unlogged Grok.com, Imagine, Voice, or Build activity between the meter readings could have contributed to the 19-point change. The calibration interval is mostly, but not entirely, XHigh. Result receipts can also omit interrupted or abandoned attempts. These factors prevent treating 31.17% or 38.13% as exact subscription debits. See the [xAI Grok usage FAQ](https://docs.x.ai/grok/faq).

## Interpretation

Token counts and available CLI receipts are measured; **mapping them to a subscription percentage is inferred**. Use ~6.6% for the 17 recorded Sol runs and ~31.2% for the 17 recorded Grok runs as planning estimates, with the named exclusions and uncertainties above. Do not publish their decimal precision as if it were measured accuracy.
