Where STAR stands today, relative to published or in-house data.
FrontierRL is constantly running more benchmarks — results land on this page as they finish.
| Eval | FrontierRL Star | Claude Fable 5 | Claude Fable 5 (Xhigh) | Claude Opus 4.8 | Claude Opus 5 | Claude Sonnet 5 | DeepSeek v4 Flash | DeepSeek v4 Pro | Gemini 3 Pro | Gemini 3.1 Pro | Gemini 3.1 Pro Preview | Gemini 3.6 Flash | GPT‑5 | GPT‑5.5 | GPT‑5.6 Luna | GPT‑5.6 Sol | GPT‑5.6 Sol (High) | GPT‑5.6 Sol (Max) | GPT‑5.6 Terra | Kimi K3 | Qwen3.7 Max | Source for other model's metrics |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| BenchCAD (python tool) | 81.75% | — | — | 51.8% | — | — | — | — | — | — | — | — | — | 55.8% | 73.9% | 83.4% | — | — | 78.2% | — | — | openai.com/index/gpt-5-6 |
| IFEval | 97.97% | — | — | — | — | — | 91.6% | 93.4% | — | — | — | — | — | — | — | — | — | 96.64% | — | — | 94.2% | In-house tested by FrontierRL |
| OSWorld Verified | 85.36% | 85% | — | 83.4% | — | — | — | — | — | 76.2% | — | — | — | 78.7% | — | — | — | — | — | — | — | anthropic.com/claude/mythos |
| Terminal-Bench 2.1 | 75.28% | 80.52% | — | — | 84.64% | 74.53% | — | — | — | — | 70.79% | 73.78% | — | 76.40% | — | 85.77% | — | — | — | 80.90% | — | openai.com/index/gpt-5-6 |
| Video-MME v2 | 83% | — | 77% | — | — | — | — | — | 66.1% | — | — | — | 44.7% | — | — | — | 67% | 69% | — | — | — | In-house tested by FrontierRL |
FrontierRL Star results for these land here as they finish.
| Eval | FrontierRL Star | Claude Fable 5 | Claude Opus 4.8 | GPT‑5.5 | GPT‑5.6 Luna | GPT‑5.6 Sol | GPT‑5.6 Terra |
|---|---|---|---|---|---|---|---|
| DeepSWE v1.1 | — | 69.7% | 59% | 67% | 67.2% | 72.7% | 69.6% |
| GPQA Diamond | — | 92.6% | 92% | 93.6% | 92.3% | 94.6% | 92.9% |
| GraphWalks BFS 256k f1 | — | — | 85.9% | 73.7% | 81.3% | 90.7% | 76.9% |
| MMMU Pro (with tools) | — | — | — | 83.2% | 79.5% | 84.6% | 82% |
| Toolathlon | — | 61.7% | 59.9% | 55.6% | 53.4% | 58% | 53.1% |