Research index

Choose a model. Inspect every run.

Start with the model, then open every benchmark and reasoning level tested with it. Public records keep the headline result, cost, time, record state, and evidence gaps visible. Paid research access adds the complete available ledger and disclosed configuration. Builder adds exact prompts, released projects, the RemakeBench Harness and production skills for the tests in its library.

Published model index

Order by average cost or average time. Run count stays visible for context.

Download for Google SheetsCSV
21 models · 274 runs

18 models shown, ordered by average cost. 23 runs match the search.

01DeepSeek V4 FlashOpenCode Go · Provider not separately recorded · opencode-goMaxVision exp24runs$0.2623/24 USD measured2h 3m13/24 measured02GLM 5.2OpenCode GoMax1run$0.271/1 USD measured6m 28s1/1 measured03DeepSeek V4 ProOpenCode GoMax12runs$0.3112/12 USD measured1h 16m12/12 measured04Qwen3.8 FlashProvider not separately recordedUnreported11runs$0.4311/11 USD measured0/11 measured05GPT-5.6 LunaOpenAI Codex · Provider not separately recordedMaxxhigh25runs$0.4914/25 USD measured23m 26s14/25 measured06GLM 5.3 FlashOpenCode Go · Provider not separately recorded · Unavailable · ZCode · opencode-goMax28runs$1.5124/28 USD measured1h 1m14/28 measured07GPT-5.6 TerraOpenAI CodexUltra8runs$1.648/8 USD measured20m 59s8/8 measured08Gemini 3.6 FlashGoogle Gemini via AntigravityHigh1run$1.921/1 USD measured7m 9s1/1 measured09Gemini 3.8 FlashGoogle · OpenRouter · UnavailableHighHighest14runs$2.8014/14 USD measured54m 25s14/14 measured10Grok 4.6Grok CLI subscription · xAI Grok Build · xAI via grok-subxhigh21runs$5.4418/21 USD measured29m 31s20/21 measured11Kimi K3Kimi Code · Kimi Code CLI · Kimi Code CLI / Moonshot AI · OpenCode Go / Moonshot AI · openai-compatibleMaxUnreported18runs≥$5.9615/18 USD measured1h 11m18/18 measured12GPT-5.6 SolOpenAI Codex · Provider not separately recordedMaxUltraxhigh26runs$8.4526/26 USD measured39m 15s26/26 measured13Qwen3.8 Max PreviewAlibaba Cloud Model Studio · Alibaba Cloud Model Studio Token Plan International · Alibaba ModelStudio Token Plan InternationalMandatory thinking enabledUnreported11runs$12.0011/11 USD measured1h 16m11/11 measured14Claude Sonnet 5Anthropic Claude CodeMax2runs$18.842/2 USD measured1h 3m2/2 measured15Qwen3.8 MaxAlibaba Cloud Model Studio Token Plan via Qwen Code · Alibaba Cloud Model Studio token plan · Alibaba Cloud Model Studio token plan via Qwen Codexhigh8runs$24.991/8 USD measured · 7 non-USD1h 17m8/8 measured16Claude Fable 5Anthropic Claude Code · not suppliedLowMaxMedium18runs$56.7417/18 USD measured34m 20s17/18 measured17Claude Opus 5Anthropic · Anthropic Claude Code · Anthropic Claude Code with Blender MCP · Claude Code · Provider not separately recordedMax19runs≥$80.7619/19 USD measured2h 54m19/19 measured18Claude Fable 5.1Anthropic Claude CodeMax17runs$105.6717/17 USD measured1h 48m17/17 measured
Every public run

Every run, cost × time.

Compare individual model runs without flattening them into an average. The lower-left region is lower-cost and faster.

Run matrix18 measured runs on one field

Each dot is one public run. Cost runs left to right; wall-clock time runs bottom to top. Logarithmic spacing keeps the full range legible.

GLM 5.3 Flash2GPT-5.6 Sol2GPT-5.6 Luna2DeepSeek V4 Flash2Grok 4.61Claude Opus 51Claude Fable 52Kimi K31Claude Fable 5.11Gemini 3.8 Flash1DeepSeek V4 Pro1Qwen3.8 Flash1Qwen3.8 Max Preview1GPT-5.6 Terra1Qwen3.8 Max1Claude Sonnet 51Gemini 3.6 Flash1GLM 5.21
5 unplotted

GLM 5.3 FlashInfinite cathedral corridor shader lacks measured cost or time · GPT-5.6 LunaInfinite cathedral corridor shader Measured or estimated cost is unavailable; excluded from cost averages and the cost × time plot. · DeepSeek V4 FlashInfinite cathedral corridor shader lacks measured cost or time · Qwen3.8 FlashInfinite cathedral corridor shader lacks measured cost or time · Qwen3.8 MaxInfinite cathedral corridor shader Non-USD cost; excluded from USD averages and the cost × time plot.

RemakeBenchResearch console