Research index

Choose a model. Inspect every run.

Start with the model, then open every benchmark and reasoning level tested with it. Public records keep the headline result, cost, time, record state, and evidence gaps visible. Paid research access adds the complete available ledger and disclosed configuration. Builder adds exact prompts, released projects, the RemakeBench Harness and production skills for the tests in its library.

Published model index

Order by average cost or average time. Run count stays visible for context.

Download for Google SheetsCSV
21 models · 274 runs

15 models shown, ordered by average cost. 21 runs match the search.

01DeepSeek V4 FlashOpenCode Go · Provider not separately recorded · opencode-goMaxVision exp24runs$0.2623/24 USD measured2h 3m13/24 measured02Qwen3.8 FlashProvider not separately recordedUnreported11runs$0.4311/11 USD measured0/11 measured03GPT-5.6 LunaOpenAI Codex · Provider not separately recordedMaxxhigh25runs$0.4914/25 USD measured23m 26s14/25 measured04GLM 5.3 FlashOpenCode Go · Provider not separately recorded · Unavailable · ZCode · opencode-goMax28runs$1.5124/28 USD measured1h 1m14/28 measured05GPT-5.6 TerraOpenAI CodexUltra8runs$1.648/8 USD measured20m 59s8/8 measured06Gemini 3.8 FlashGoogle · OpenRouter · UnavailableHighHighest14runs$2.8014/14 USD measured54m 25s14/14 measured07Grok 4.6Grok CLI subscription · xAI Grok Build · xAI via grok-subxhigh21runs$5.4418/21 USD measured29m 31s20/21 measured08Kimi K3Kimi Code · Kimi Code CLI · Kimi Code CLI / Moonshot AI · OpenCode Go / Moonshot AI · openai-compatibleMaxUnreported18runs≥$5.9615/18 USD measured1h 11m18/18 measured09GPT-5.6 SolOpenAI Codex · Provider not separately recordedMaxUltraxhigh26runs$8.4526/26 USD measured39m 15s26/26 measured10Qwen3.8 Max PreviewAlibaba Cloud Model Studio · Alibaba Cloud Model Studio Token Plan International · Alibaba ModelStudio Token Plan InternationalMandatory thinking enabledUnreported11runs$12.0011/11 USD measured1h 16m11/11 measured11Claude Sonnet 5Anthropic Claude CodeMax2runs$18.842/2 USD measured1h 3m2/2 measured12Qwen3.8 MaxAlibaba Cloud Model Studio Token Plan via Qwen Code · Alibaba Cloud Model Studio token plan · Alibaba Cloud Model Studio token plan via Qwen Codexhigh8runs$24.991/8 USD measured · 7 non-USD1h 17m8/8 measured13Claude Fable 5Anthropic Claude Code · not suppliedLowMaxMedium18runs$56.7417/18 USD measured34m 20s17/18 measured14Claude Opus 5Anthropic · Anthropic Claude Code · Anthropic Claude Code with Blender MCP · Claude Code · Provider not separately recordedMax19runs≥$80.7619/19 USD measured2h 54m19/19 measured15Claude Fable 5.1Anthropic Claude CodeMax17runs$105.6717/17 USD measured1h 48m17/17 measured
Every public run

Every run, cost × time.

Compare individual model runs without flattening them into an average. The lower-left region is lower-cost and faster.

Run matrix16 measured runs on one field

Each dot is one public run. Cost runs left to right; wall-clock time runs bottom to top. Logarithmic spacing keeps the full range legible.

GLM 5.3 Flash2GPT-5.6 Sol2GPT-5.6 Luna3DeepSeek V4 Flash2Grok 4.61Claude Opus 51Claude Fable 52Kimi K31Claude Fable 5.11Gemini 3.8 Flash1Qwen3.8 Flash1Qwen3.8 Max Preview1GPT-5.6 Terra1Qwen3.8 Max1Claude Sonnet 51
5 unplotted

GLM 5.3 FlashCampfire under the stars lacks measured cost or time · GPT-5.6 LunaCampfire under the stars Measured or estimated cost is unavailable; excluded from cost averages and the cost × time plot. · DeepSeek V4 FlashCampfire under the stars lacks measured cost or time · Qwen3.8 FlashCampfire under the stars lacks measured cost or time · Qwen3.8 MaxCampfire under the stars Non-USD cost; excluded from USD averages and the cost × time plot.

RemakeBenchResearch console