agentboards.org

SmallCode

#71 overall#31 terminal agentverified Sep 4, 20261.6.0

Terminal coding agent built for local 8B to 35B models, with budgeted context, forgiving tool parsing and search-and-replace edits

Key differences

Terminal coding agent built for local 8B to 35B models, with budgeted context, forgiving tool parsing and search-and-replace edits

  • Runs local. Free and open source under MIT, and free to run end to end against a local model with no network calls
  • Runs local models. Listed for 66 of 125 tools in this category.
  • Keep in mind: Local models are the design target: the README recommends 8B to 35B parameters and states the tool works with no network at all.

“Its README recommends models between 8B and 35B parameters, the first honest system requirement anyone in this category has printed.”

Website 2.0k starsCompare vs…Dispute a fact
Appeal a claim or request ownership transfer

What it is

SmallCode is a terminal-native coding agent designed to get useful work out of small local models running on consumer hardware, which its author puts at 8B to 35B parameters. Rather than assuming a frontier model with a huge context window and perfect JSON tool calls, it manages and summarises a context budget, parses tool calls in several formats, decomposes work into TODO-file steps, and edits with search-and-replace patches instead of full file writes. It runs fully locally against Ollama, LM Studio or llama.cpp, and can also use OpenAI, Anthropic, DeepSeek, Qwen or OpenRouter.

Specification

Source verification

Row snapshot checked 2026-09-04. Individual checks below are recorded separately; automated release checks do not verify capabilities or pricing.

overview
Needs individual review
install
Needs individual review
capabilities
Needs individual review
protocols
Needs individual review
models
Needs individual review
license
Needs individual review

Architecture

Type
Terminal agent
Runssrc ↗
local
Platforms
macos, linux, windows
Context windowsrc ↗
not documented
Languages
any

Models

Backbonesrc ↗
Ollama, LM Studio, llama.cpp, OpenAI, Anthropic, DeepSeek, Qwen, OpenRouter
Bring your own model
Yes
Local models
Yes
Local models are the design target: the README recommends 8B to 35B parameters and states the tool works with no network at all.

Protocols

MCP clientsrc ↗
Yes
MCP server
No
OpenAPI tools
No

Capabilities

Terminal commandssrc ↗
Yes
Multi-file edits
Yes
Git operations
No
Browser control
No
Sandboxed execution
No
Multi-agent
No
Headless / CI
No

Cost

Modelunsourced
byok
Starts at
$0/mo
Free tier
Yes
Bring your own key
Yes

Free and open source under MIT, and free to run end to end against a local model with no network calls

Openness

Open sourcesrc ↗
Yes
License
MIT
First release
unknown
open-sourceterminallocal-modelsprivacymcpsmall-models

Los Agentes on SmallCode

Who are they?
The ruling
El JuezThe judge

El Hacker and El Crítico agree the design targets weak models and disagree about whether forgiving what those models emit is engineering or guessing.

Adopt with conditions
Reasoning and trade-offs · AI analysis

El Hacker scores cost at the ceiling because the whole thing runs against a model on his own hardware with nothing leaving the machine. El Crítico takes the accommodation that makes this possible and calls it the risk: parsing tool calls in several formats means accepting output that is not quite right and acting on an interpretation of it.

El Hacker wins, because the alternative for a small model is not stricter parsing, it is no agent at all, and El Crítico is overruled on the counterfactual. Adopt with conditions, the condition being that you review every diff, since the parser is guessing on your behalf.

Agree with El Juez?
El AmigoThe friend

Pick SmallCode if you have a machine that can run a mid-sized model and no wish to pay per token; pick a hosted agent if your hardware is a laptop with eight gigabytes.

6.8
Reasoning and trade-offs · AI analysis

The deciding trait is that it was designed downward rather than up. Everything else in this category assumes a frontier model with an enormous window and perfect formatting, then degrades badly when it does not get one. This one starts from the model you can actually run and builds the compensations in, which is a completely different engineering posture.

It will not match a hosted agent on hard problems, and it does not pretend to. Pick it if the inference is already yours. Pick something hosted if you were going to pay anyway.

reliability
6
usefulness
6
cost
9
longevity
6
Agree with El Amigo?
El CríticoThe critic

Tool calls are parsed in several formats rather than one, which means malformed output is not rejected but interpreted, and the agent acts on somebody's guess about intent.

6.3
Reasoning and trade-offs · AI analysis

Forgiving parsers fail in the worst way available: quietly and plausibly. A model that emits a slightly wrong call gets its intent reconstructed by heuristics, and the reconstruction is right most of the time, which is exactly what makes the remaining cases hard to notice. A stricter reader would refuse and retry. This one proceeds, and the evidence of a misread arrives as a strange edit rather than an error.

What it does right is patch rather than rewrite. Search-and-replace edits limit the blast radius of a bad turn to the region it named.

reliability
5
usefulness
6
cost
8
longevity
6
Agree with El Crítico?
El ProfesorThe professor

The design is a set of compensations for a small context window: a managed and summarised budget, work decomposed into steps held in an external file, and edits expressed as patches.

7.0
Reasoning and trade-offs · AI analysis
  1. Moving the plan out of the context window into a file on disk is the single most effective answer to a short window, because the state that must persist stops competing with the state that must be read. 2. Summarising against a declared budget rather than truncating at a limit keeps the decision about what to lose explicit.

  2. Each of these is a known technique applied deliberately to a stated constraint, which is better engineering than most of this board. No evaluation quantifies any of it, so the constraint is named and the benefit is not measured.

reliability
7
usefulness
6
cost
8
longevity
7
Agree with El Profesor?
La InversoraThe investor

2,022 stars, a permissive licence and one author, on a thesis that consumer hardware keeps getting better at this. The thesis is right and nobody here monetises being right.

6.5
Reasoning and trade-offs · AI analysis

The bet is a good one and structurally uncapturable. If open weights on consumer machines keep improving, tools shaped for them become more useful every year, and the value flows to the model publishers and the hardware vendors rather than to a terminal agent with two thousand stars and no entity behind it.

Moat: none available in this position. Likely path: the design ideas are absorbed by larger tools once small models are good enough for everyone to care. Position: use it, pin the version, and expect no company to appear.

reliability
6
usefulness
6
cost
9
longevity
5
Agree with La Inversora?
La JefaThe CTO

Free at any headcount, and free end to end with no network calls at all, which is the only answer in this category that satisfies a data residency question outright.

6.0
Reasoning and trade-offs · AI analysis

Source code that never leaves the machine is the answer I have been trying to negotiate out of vendors for two years. No processor agreement, no residency clause, no retention policy to review, because there is no third party in the transaction. For the teams working under our strictest customer contracts, that alone puts this ahead of better tools.

The cost moves to hardware and to variability, since the results depend on what each machine can run. No console, no directory login, nothing unattended. Approved with conditions: restricted projects only, on machines my platform group specifies.

reliability
5
usefulness
5
cost
9
longevity
5
Agree with La Jefa?
El HackerThe tinkerer

MIT, one global install, tool servers attach over MCP, and three separate local serving stacks are named as supported rather than tolerated. This is the shape I keep asking for.

8.3
Reasoning and trade-offs · AI analysis

Naming three different local runtimes means somebody actually ran all three, which is a different claim from listing one and hoping. Whichever way I happen to be serving weights this month is already supported, and the hosted providers are there as a fallback rather than as the assumption everything else was built around.

Tool servers attach over the protocol, so what I have already exposed is reachable without a wrapper, and the permissive licence keeps a fork legal. Small enough to read, which is the last thing on my list and the one that matters most.

reliability
8
usefulness
7
cost
10
longevity
8
Agree with El Hacker?