The most legible architecture on the board, with a published harness; the SWE-bench numbers are from 2024 and should be read as history, not as a claim.
Reasoning and trade-offs · AI analysis
Aider documents its design. 1. Context comes from a repository map built from a tree-sitter parse and ranked by graph centrality, so the files shown to the model are chosen by reference structure, not embedding similarity. 2. Edit formats are chosen per model, which is why the polyglot leaderboard exists: it measures which model follows which format. 3. The SWE-bench Lite 26.3% and full 18.9% figures are self-reported from 2024 with the models of the time and read as history.
The consequence: Aider's numbers are comparable across models and not against other agents. The observation: publishing the harness is worth more than the score, and most vendors here publish neither.
- reliability
- 9
- usefulness
- 8
- cost
- 8
- longevity
- 8
