Quality shortlist
GPT-6 Astra (max)
88±1.9
Expected performance %
Publisher’s expected performance, including estimated missing results.
A prompt for your current assistant, checks for the result and evidence when you need an alternative.
Only this choice is remembered on this device, when storage is available.
Your next step · Math problems
Use the prompt below, then check the result. State the domain and constraints. Ask for a checkable derivation and verify the result by substitution or a separate method.
Your text stays on this page. It is not saved, sent to AI or included in shared links.
Solve this problem: [problem]. State the domain and assumptions, show a checkable derivation, and verify the answer by substitution, boundary cases or a separate method. If a step cannot be justified, identify the gap. Do not substitute numerical examples for a general proof.
Fill any remaining brackets before sending.
Use these checks before you trust or publish it.
If it misses the mark, select “I want a better result” above. Consider another model when a clearer brief still leaves you stuck.
AgentBoards editorial guidance
There is no evidence here that you need to switch. Try your familiar tool on a clearly defined task first; an app name does not tell us its exact model or capabilities.
MathArena · Overall model analysis · Higher expected performance % is better within this benchmark.
Quality shortlist
88±1.9
Expected performance %
Publisher’s expected performance, including estimated missing results.
Alternative
83.1±5.0
Expected performance %
High effort; keep this distinct from the max result.
Effort comparison
73.9±3.7
Expected performance %
More reasoning effort does not guarantee a higher score.
What this cannot tell you: MathArena combines observed scores with estimates for missing problems across non-deprecated competitions. This is not a raw pass rate on one shared test, or proof that a solution is correct.
Sources checked 2026-09-17 · Publisher snapshot date not supplied. This snapshot is due for review. Check the publisher before choosing.
Start with one small example. Before committing, compare at least three representative examples using the same inputs, tools and settings; record failures too. AgentBoards has not run these models on your work.
Scores keep their publisher’s units and exact configurations. We do not average unrelated benchmarks. This is a curated shortlist, not a full leaderboard. Snapshots need review after 14 days; sponsorship cannot buy a recommendation.
Six everyday tasks. Learn a useful trick with every answer.