Research index

Choose a model. Inspect every run.

Start with the model, then open every benchmark and reasoning level tested with it. Public records keep the headline result, cost, time, record state, and evidence gaps visible. Paid research access adds the complete available ledger and disclosed configuration. Builder also includes the separate, immutable Tech Review 001 prompt and data package for its 28 named runs.

6 models found.

01Claude Fable 5Anthropic Claude CodeLowMaxMedium10benchmarks17runs≥$684.4802GPT-5.6 SolOpenAI CodexUltraxhigh10benchmarks17runs$83.7803Qwen3.8 Max PreviewAlibaba Cloud Model StudioMandatory thinking enabledUnreported10benchmarks11runs$131.9904Kimi K3Kimi Code CLIMax10benchmarks10runs≥$30.1105GPT-5.6 LunaOpenAI Codexxhigh8benchmarks8runs$4.3906GPT-5.6 TerraOpenAI CodexUltra8benchmarks8runs$13.14
RemakeBenchResearch console