agentboards.org
Compare/mini-SWE-agent vs Stakpak

mini-SWE-agentvsStakpak

Generated from the two spec rows. Green marks the better value where a spec has a clear direction. Everything else is just different.

mini-SWE-agent
SWE-agent (Princeton and Stanford) · Terminal agent
#35OSS
Panel
6.9
2 spec wins
Reliability
6.3
Usefulness
5.5
Cost
8.8
Longevity
7.0

“One hundred lines of Python, which is fewer than most competitors spend on their pricing page.”

Stakpak
Stakpak · Terminal agent
#133OSSMCP
Panel
7.0
3 spec wins
Reliability
6.7
Usefulness
7.3
Cost
7.8
Longevity
6.3

“It debugs Kubernetes on its own initiative, which is either the future or the most thankless job ever handed to a machine.”

Spec by spec

Specmini-SWE-agentStakpak
Architecture
CategoryTerminal agentTerminal agent
Runslocal, sandboxlocal, sandbox
Platformsmacos, linuxmacos, linux
Context windownot documentednot documented
Protocols
MCP clientNoYes
MCP serverNoYes
Capabilities
Runs terminal commandsYesYes
Multi-file editsYesYes
Git operationsYesNo
Browser controlNoNo
Sandboxed executionYesYesSubagents run sandboxed analysis with restricted tool access, and the README recommends 2GB+ RAM for autopilot and sandbox runs.
Multi-agent orchestrationNoYes
Headless / CI modeYesYes
Models
Backboneany via litellmany OpenAI-compatible endpoint, Ollama, LM Studio
Bring your own modelYesYes
Local modelsYesYes
Cost
Pricing modelbyokbyok
Starts atn/a$0/mo
Free tierYesYes
Bring your own keyYesYes
Openness
Open sourceYesYes
LicenseMITApache-2.0
GitHub stars8,1451,806

Which one would each critic pick

Criticmini-SWE-agentStakpakPick
El Juez——not enough reviews
El Amigo6.87.3Stakpak — Pick it if you want infrastructure work that continues while you sleep and wakes you only when it must; pick a normal terminal agent if you want to watch every step.
El Crítico6.56.5no preference
El Profesor8.07.3mini-SWE-agent — The reference harness for the SWE-bench bash-only leaderboard: one bash action per turn via subprocess.run, a linear history, no tool-calling API, and Gemini 3 Pro reported above 74% on Verified.
La Inversora6.86.8no preference
La Jefa5.06.5Stakpak — It runs headless, so it becomes a pipeline step I can gate on, and there is still no directory integration or retention policy to put in front of my security team.
El Hacker8.58.0mini-SWE-agent — MIT, any model through litellm, OpenRouter or Portkey including my local server, a YAML config, and source short enough to read before breakfast; the missing MCP client is the only gap.

Picks are derived from each critic's own scores. Humans vote on matchups on the duels page.

Want a third column? The compare tool handles any two agents. Three-way comparisons are on the roadmap once the spec rows are all verified.