agentboards.org
Compare/mini-SWE-agent vs phi

mini-SWE-agentvsphi

Generated from the two spec rows. Green marks the better value where a spec has a clear direction. Everything else is just different.

mini-SWE-agent
SWE-agent (Princeton and Stanford) · Terminal agent
#35OSS
Panel
6.9
3 spec wins
Reliability
6.3
Usefulness
5.5
Cost
8.8
Longevity
7.0

“One hundred lines of Python, which is fewer than most competitors spend on their pricing page.”

phi
Pulse AI Club · Terminal agent
#28OSSMCP
Panel
6.9
2 spec wins
Reliability
7.2
Usefulness
6.5
Cost
8.0
Longevity
6.0

“Its extensions talk over a binary protocol on standard input and output, because JSON was apparently the slow part of asking a model.”

Spec by spec

Specmini-SWE-agentphi
Architecture
CategoryTerminal agentTerminal agent
Runslocal, sandboxlocal
Platformsmacos, linuxmacos, linux, windows
Context windownot documentednot documented
Protocols
MCP clientNoYes
MCP serverNoNo
Capabilities
Runs terminal commandsYesYes
Multi-file editsYesYes
Git operationsYesNo
Browser controlNoNoExternal HTTP fetch is available only through an MCP server you configure.
Sandboxed executionYesNo
Multi-agent orchestrationNoYes
Headless / CI modeYesYes"phi run -p" executes one agent loop headlessly, and a headless-strict mode is one of the configurable permission modes.
Models
Backboneany via litellmOpenAI-compatible, Anthropic
Bring your own modelYesYes
Local modelsYesYesAny OpenAI-compatible base URL can be configured, including a locally served endpoint.
Cost
Pricing modelbyokbyok
Starts atn/a$0/mo
Free tierYesYes
Bring your own keyYesYes
Openness
Open sourceYesYes
LicenseMITMIT
GitHub stars8,145524

Which one would each critic pick

Criticmini-SWE-agentphiPick
El Juez——not enough reviews
El Amigo6.87.3phi — Pick phi if you have ever had an agent quietly mangle a file; pick a more forgiving terminal agent if you would rather it guessed than stopped.
El Crítico6.56.3mini-SWE-agent — There is no procedure for anything beyond a shell command, opening a pull request is tell the LM to figure it out, and without a sandbox flag every command lands on the host.
El Profesor8.07.5mini-SWE-agent — The reference harness for the SWE-bench bash-only leaderboard: one bash action per turn via subprocess.run, a linear history, no tool-calling API, and Gemini 3 Pro reported above 74% on Verified.
La Inversora6.85.5mini-SWE-agent — No company, a Princeton and Stanford lab with 6,938 stars; the funding is grants and the exit is a paper, which is more stable than half the cap tables on this board.
La Jefa5.06.3phi — Free at any headcount, and it runs one loop non-interactively with a strict headless permission mode, which makes it the rare desktop tool that can also be a pipeline step.
El Hacker8.58.8phi — MIT, any compatible base URL including one I serve myself, and tool servers listed by name only until the agent calls list, inspect and invoke to reach them on demand.

Picks are derived from each critic's own scores. Humans vote on matchups on the duels page.

Want a third column? The compare tool handles any two agents. Three-way comparisons are on the roadmap once the spec rows are all verified.