Quality shortlist
GPT-6 Astra (max)
1800±16
Preference score
$10 in / $50 outUSD per million tokens · listed by publisher
2,281 votes; publisher rank range 1–1.
A prompt for your current assistant, checks for the result and evidence when you need an alternative.
Only this choice is remembered on this device, when storage is available.
Your next step · Coding & web apps
Use the prompt below, then check the result. Give it a small, reproducible problem and a clear passing condition. Ask it to inspect the project, change one thing, and run the relevant checks.
Your text stays on this page. It is not saved, sent to AI or included in shared links.
Implement [small feature] in this project. First inspect its conventions and state the acceptance criteria. Handle empty, loading and error states. Include keyboard access and a mobile layout. Run the relevant checks and report exactly what passed, what failed and what was not tested. Do not claim success from a screenshot alone.
Fill any remaining brackets before sending.
Use these checks before you trust or publish it.
If it misses the mark, select “I want a better result” above. Consider another model when a clearer brief still leaves you stuck.
AgentBoards editorial guidance
There is no evidence here that you need to switch. Try your familiar tool on a clearly defined task first; an app name does not tell us its exact model or capabilities.
Arena · Code / WebDev · Higher preference score is better within this benchmark.
Quality shortlist
1800±16
Preference score
$10 in / $50 outUSD per million tokens · listed by publisher
2,281 votes; publisher rank range 1–1.
Alternative
1758±14
Preference score
$10 in / $50 outUSD per million tokens · listed by publisher
3,036 votes; publisher rank range 2–2.
Lower API price
1607±12
Preference score
$0.07 in / $0.25 outUSD per million tokens · listed by publisher
2,925 votes; publisher rank range 17–17. Lower score and lower listed API price.
What this cannot tell you: This measures web-development preferences in Arena’s setup. It is not a repository bug-fix success rate. Agent tools, instructions and test coverage affect the result.
Listed API prices are not chat subscription prices or cost per finished task. Reasoning tokens, tools, retries and provider terms can change the bill; confirm pricing before buying.
Sources checked 2026-09-17 · Publisher snapshot 2026-09-11. This snapshot is due for review. Check the publisher before choosing.
Fixing an existing repository? Also inspect SWE-bench’s comparable agent setups and the coding agent board. A model and the agent running it are different choices.
Start with one small example. Before committing, compare at least three representative examples using the same inputs, tools and settings; record failures too. AgentBoards has not run these models on your work.
Scores keep their publisher’s units and exact configurations. We do not average unrelated benchmarks. This is a curated shortlist, not a full leaderboard. Snapshots need review after 14 days; sponsorship cannot buy a recommendation.
Six everyday tasks. Learn a useful trick with every answer.