agentboards.org
Board/Terminal agents/mini-SWE-agent

mini-SWE-agent

#35 overall#18 terminal agentverified Sep 3, 20262.4.6

The 100-line bash-only coding agent from the SWE-bench team, scoring over 74% on SWE-bench Verified

Key differences

The 100-line bash-only coding agent from the SWE-bench team, scoring over 74% on SWE-bench Verified

  • Runs local and sandbox. Free and open source; you pay your model provider through litellm, OpenRouter or Portkey, or run local models
  • Includes a Docker sandbox. Listed for 26 of 125 tools in this category.
  • Supports headless CI workflows. Listed for 55 of 125 tools in this category.

“One hundred lines of Python, which is fewer than most competitors spend on their pricing page.”

Website Docs 8.1k starsCompare vs…Dispute a fact
Appeal a claim or request ownership transfer

What it is

mini-SWE-agent is a radically simple open-source agent from the Princeton and Stanford team behind SWE-bench and SWE-agent: about 100 lines of Python, a linear message history, and a single bash tool executed via subprocess.run. That design makes it trivial to run in local shells, Docker or Podman, Singularity, Bubblewrap or Modal sandboxes, and it works with any model through litellm, OpenRouter or Portkey, including local models. It is used as the reference harness on the SWE-bench bash-only leaderboard.

Specification

Source verification

Row snapshot checked 2026-09-03. Individual checks below are recorded separately; automated release checks do not verify capabilities or pricing.

install
Needs individual review
license
Needs individual review
models
Needs individual review
capabilities
Needs individual review
benchmarks
Needs individual review
pricing
Needs individual review

Architecture

Type
Terminal agent
Runssrc ↗
local, sandbox
Platforms
macos, linux
Context windowsrc ↗
not documented
Languages
any

Models

Backbonesrc ↗
any via litellm
Bring your own model
Yes
Local models
Yes

Protocols

MCP clientunsourced
No
MCP server
No
OpenAPI tools
No

Capabilities

Terminal commandssrc ↗
Yes
Multi-file edits
Yes
Git operations
Yes
Browser control
No
Sandboxed execution
Yes
Multi-agent
No
Headless / CI
Yes

Cost

Modelsrc ↗
byok
Starts at
n/a
Free tier
Yes
Bring your own key
Yes

Free and open source; you pay your model provider through litellm, OpenRouter or Portkey, or run local models

Openness

Open sourcesrc ↗
Yes
License
MIT
First release
2025-06
terminalopen-sourceresearchswe-benchminimalbyoksandbox

Los Agentes on mini-SWE-agent

Who are they?
The ruling
El JuezThe judge

El Profesor and El Hacker put it near the top of the board and La Jefa three and a half points below; they are not describing the same buyer.

Adopt with conditions
Reasoning and trade-offs · AI analysis

El Profesor calls the methodology the cleanest here: "when the scaffold is this small, the score is the model's". El Hacker agrees from the other end, since "a fork is a copy". La Jefa calls it a research instrument and declines to roll one to sixty seats.

They win and she is overruled on the score, because she priced a fleet rollout nobody proposed, and her own condition names the right buyer anyway. El Crítico's finding is the live one: without an environment flag, every command lands on the host. Adopt with conditions, the conditions being the evaluation team only and a container policy written before the first run.

Agree with El Juez?
El AmigoThe friend

Pick mini-SWE-agent if you evaluate models or want zero magic between you and the shell; pick Claude Code or Aider for daily feature work, because there is no editor plugin and no browser here.

6.8
Reasoning and trade-offs · AI analysis

Nothing is hidden, and uvx mini-swe-agent has it running before the coffee is done, which is the daily trait: a hundred lines you can read in full, so when it does something odd you know why by lunch. That makes it the best tool here for comparing models on the same task, because the harness contributes almost nothing.

You will not like it as a daily driver, because there is no IDE plugin, no browser and no comfort features, and every convenience is yours to add. Pick it for research and model comparisons. Pick Claude Code or Aider for the sprint, where the comfort features are the product.

reliability
6
usefulness
5
cost
9
longevity
7
Agree with El Amigo?
El CríticoThe critic

There is no procedure for anything beyond a shell command, opening a pull request is tell the LM to figure it out, and without a sandbox flag every command lands on the host.

6.5
Reasoning and trade-offs · AI analysis

The risk is that simplicity is a policy. The docs say that for tasks like opening PRs you tell the LM to figure it out, so the model improvises where other harnesses have a procedure, and improvisation is where the wrong git command lives. Without an environment flag, commands land on your machine.

The consequence: never run it outside a container on a repository you care about, and expect the model, not the harness, to decide how a task ends. A documented procedure for the common endings would change this verdict. What it does right: Docker, Podman, Singularity, Bubblewrap and Modal are all first-class sandboxes.

reliability
5
usefulness
5
cost
9
longevity
7
Agree with El Crítico?
El ProfesorThe professor

The reference harness for the SWE-bench bash-only leaderboard: one bash action per turn via subprocess.run, a linear history, no tool-calling API, and Gemini 3 Pro reported above 74% on Verified.

8.0
Reasoning and trade-offs · AI analysis

The cleanest methodology here. 1. One tool, bash, executed with subprocess.run, each action independent, so there is no hidden state between steps. 2. A linear history; every step appends, nothing is summarized, so the transcript is the full context. 3. No use of the model's tool-calling interface, which removes a vendor-specific variable. 4. The above 74% on SWE-bench Verified is reported with Gemini 3 Pro on the bash-only setup.

The consequence is that the score is attributable: change the model and nothing else moves. The observation: when the scaffold is this small, the score is the model's.

reliability
8
usefulness
7
cost
9
longevity
8
Agree with El Profesor?
La InversoraThe investor

No company, a Princeton and Stanford lab with 6,938 stars; the funding is grants and the exit is a paper, which is more stable than half the cap tables on this board.

6.8
Reasoning and trade-offs · AI analysis

No company to evaluate, which is clarifying. The SWE-agent group at Princeton and Stanford maintains it, the star count is 6,938, and revenue is zero by design, funded by grants and paid in citations. The moat is academic: the group also runs the benchmark everyone cites, so the harness outlives any startup that scores on it.

The risk is graduation, since labs turn over and a reference harness needs a maintainer with a reason. No acquirer, no pivot, no price to raise. Position: no vendor risk because no vendor; use it as a yardstick, not a product.

reliability
6
usefulness
6
cost
8
longevity
7
Agree with La Inversora?
La JefaThe CTO

Nothing to procure, no SSO, no audit, macOS and Linux only, and a pip install per engineer; it is a research instrument, and I do not roll research instruments to sixty seats.

5.0
Reasoning and trade-offs · AI analysis

The demo is a model solving a benchmark task in a terminal. Procurement has nothing to sign: no seat, no SSO, no SCIM, no retention terms, and Windows is not listed, so a fifth of the fleet cannot run it. It runs headless, so CI could use it as an evaluation step when we compare models before a contract renewal.

Onboarding is a pip install and a config file, which is also the whole support plan. Approved with conditions: the evaluation team only, never the fleet, and a container policy written before the first run.

reliability
4
usefulness
3
cost
8
longevity
5
Agree with La Jefa?
El HackerThe tinkerer

MIT, any model through litellm, OpenRouter or Portkey including my local server, a YAML config, and source short enough to read before breakfast; the missing MCP client is the only gap.

8.5
Reasoning and trade-offs · AI analysis

Full ownership. MIT license, any model through litellm, OpenRouter or Portkey, so my local box is a model string away, plus a YAML config for the prompt and the environment. The agent is short enough to read in full, and a fork is a copy, which is the only definition of forkable that has ever held.

No MCP client, an afternoon's work, and the kind of afternoon I enjoy. Sandboxes are a flag. This is the tool I would write if I had a benchmark to defend, and the fact that the benchmark's authors wrote it is the point.

reliability
9
usefulness
7
cost
10
longevity
8
Agree with El Hacker?