FrontierRL
Book an intro

Benchmarks.

Where STAR stands today, relative to published or in-house data.
FrontierRL is constantly running more benchmarks — results land on this page as they finish.

Data source for other models: OpenAI, “GPT‑5.6” announcement — openai.com/index/gpt-5-6 (July 2026).
Figures recharted by FrontierRL; values reproduced as published.

EvalFrontierRL STARGPT‑5.6 SolGPT‑5.6 Sol UltraGPT‑5.6 TerraGPT‑5.6 LunaGPT‑5.5Claude Mythos 5Claude Mythos PreviewClaude Opus 4.8Gemini 3.1 Pro PreviewDeepSeek v4 FlashDeepSeek v4 ProQwen3.7 MaxGPT‑5.5 Medium
BenchCAD (python tool)81.75%83.4%78.2%73.9%55.8%65%61%51.8%

Prompt-level strict accuracy: IFEval (541 verifiable-instruction prompts), measured in-house by FrontierRL

IFEval97.97%91.6%93.4%94.2%96.2%

Dropping Soon!

FrontierRL STAR results for these land here as they finish.

EvalFrontierRL STARGPT‑5.6 SolGPT‑5.6 Sol UltraGPT‑5.6 TerraGPT‑5.6 LunaGPT‑5.5Claude Mythos 5Claude Mythos PreviewClaude Fable 5Claude Opus 4.8Gemini 3.1 Pro Preview
Terminal-Bench 2.188.8%91.9%87.4%84.7%85.6%88%83.1%78.9%70.7%
MMMU Pro (with tools)84.6%82%79.5%83.2%
GPQA Diamond94.6%92.9%92.3%93.6%94.1%94.6%92.6%92%94.3%
Toolathlon58%53.1%53.4%55.6%61.7%61.1%61.7%59.9%48.8%
GraphWalks BFS 256k f190.7%76.9%81.3%73.7%91.1%85.7%85.9%
HealthBench Professional60.5%57.7%55.7%49.5%60.9%53%
DeepSWE v1.172.7%69.6%67.2%67%69.7%59%11.8%

Next Gen Frontier Inference