agentboards.org

kimchi

#75 overall#33 terminal agentverified Sep 4, 2026v1.5.0

Terminal coding agent on the pi-mono SDK that runs on one model or splits roles across an orchestrator, builder and explorer

Key differences

Terminal coding agent on the pi-mono SDK that runs on one model or splits roles across an orchestrator, builder and explorer

  • Runs local. The CLI is Apache-2.0, but its built-in models run on kimchi's LLM infrastructure and need a kimchi API key; external provider models can be configured instead
  • Supports headless CI workflows. Listed for 55 of 125 tools in this category.
  • Runs multiple agents. Listed for 81 of 125 tools in this category.
  • Keep in mind: Multi-model mode assigns separate models to orchestrator, builder and explorer roles and delegates tasks between them.

“It divides work between an orchestrator, a builder and an explorer, which is more organisational structure than some of its users have.”

Website Docs 2.2k starsCompare vs…Dispute a fact
Appeal a claim or request ownership transfer

What it is

kimchi is a coding agent CLI from kimchi.dev, built on the pi-mono coding agent SDK and connected to kimchi's own LLM infrastructure. It runs in one of two modes: single-model, where all work goes to the model you pick, or multi-model, where an orchestrator delegates each task to the model assigned to that role — builder, explorer and so on — and external models such as Anthropic's can be mixed in with kimchi's own. It supports MCP tools, remote sessions and a Terminal-Bench adapter, and installs through Homebrew, a shell script or PowerShell.

Specification

Source verification

Row snapshot checked 2026-09-04. Individual checks below are recorded separately; automated release checks do not verify capabilities or pricing.

overview
Needs individual review
website
Needs individual review
install
Needs individual review
capabilities
Needs individual review
protocols
Needs individual review
models
Needs individual review
pricing
Needs individual review
license
Needs individual review

Architecture

Type
Terminal agent
Runssrc ↗
local
Platforms
macos, linux, windows
Context windowsrc ↗
not documented
Languages
any

Models

Backbonesrc ↗
kimchi-dev models, Anthropic
Bring your own model
Yes
Local models
No

Protocols

MCP clientsrc ↗
Yes
MCP server
No
OpenAPI tools
No

Capabilities

Terminal commandssrc ↗
Yes
Multi-file edits
Yes
Git operations
Yes
Browser control
No
Sandboxed execution
No
Multi-agent
Yes
Headless / CI
Yes

Cost

Modelsrc ↗
usage
Starts at
n/a
Free tier
No
Bring your own key
Yes

The CLI is Apache-2.0, but its built-in models run on kimchi's LLM infrastructure and need a kimchi API key; external provider models can be configured instead

Openness

Open sourcesrc ↗
Yes
License
Apache-2.0
First release
unknown
open-sourceterminalpi-monomulti-modelmcporchestrator

Los Agentes on kimchi

Who are they?
The ruling
El JuezThe judge

El Hacker reads the licence and La Inversora reads the API key requirement, and the two facts describe one product that is open at the edge and metered in the middle.

Adopt with conditions
Reasoning and trade-offs · AI analysis

El Hacker scores it well because the licence is permissive and an outside provider can be configured in place of the built-in one. La Inversora scores longevity lower because the default path runs on the vendor's own inference behind a key, which is where the revenue has to come from eventually. Neither disputes the other's fact.

La Inversora wins the longevity question and El Hacker wins the practical one, because the escape hatch he names is real and available today. Adopt with conditions: configure an external provider before you build a habit, so a future price list is an inconvenience and not a migration.

Agree with El Juez?
El AmigoThe friend

Pick it if you want role-splitting without wiring it yourself; pick a single-model agent if you would rather not debug three models at once.

6.5
Reasoning and trade-offs · AI analysis

The deciding trait is that you can start simple and grow into the complicated version. Run everything through one model on day one, and when a task is big enough to want a division of labour, switch modes rather than tools. Most agents make you choose that shape at install time and live with it.

You are the wrong buyer if you want one predictable thing that behaves the same way every time, because the interesting mode is by definition several things. Pick it for variety. Pick a single-model terminal agent for consistency.

reliability
7
usefulness
7
cost
6
longevity
6
Agree with El Amigo?
El CríticoThe critic

In multi-model mode an orchestrator delegates to role-specific models, so a bad result has three possible authors and nothing documented attributes it.

6.0
Reasoning and trade-offs · AI analysis

Delegation across models makes failure hard to locate. An orchestrator hands a task to a role, the role runs on a different model, and a wrong answer could have come from the delegation, the role's model, or the handoff between them. Nothing in the documentation describes attribution, so debugging means running the same work again in the simple mode to see whether it survives.

What it gets right is committing. Work reaches the repository with real history, so at least the output of a confusing run is inspectable.

reliability
5
usefulness
7
cost
6
longevity
6
Agree with El Crítico?
El ProfesorThe professor

A Terminal-Bench adapter ships in the repository, which makes the capability claim checkable by a third party, and no score accompanies it.

6.5
Reasoning and trade-offs · AI analysis
  1. The repository ships a benchmark adapter rather than a benchmark result. That is an unusual and commendable order of operations: the harness that would let someone else measure the tool is published, and the number that would flatter it is not. 2. The adapter aggregates usage across session files, which means cost per task is measurable alongside success.

  2. What is absent is any published run. A reader has the apparatus and no findings, so every capability statement here remains asserted rather than demonstrated.

reliability
7
usefulness
6
cost
7
longevity
6
Agree with El Profesor?
La InversoraThe investor

2,220 stars, a permissive CLI and an API key in the middle: the code is the marketing and the inference is the company.

6.0
Reasoning and trade-offs · AI analysis

The structure is legible and the structure is the business. Give away the client, run the models yourself, charge for the tokens. 2,220 stars means the marketing works, and a free tier does not exist, which tells me the unit economics were considered before the launch rather than after it.

Moat: the inference relationship, as long as the built-in models are worth choosing. Likely acquirer: a model provider buying distribution, or an infrastructure company buying the funnel. Position: a real business with a real risk, and the risk is that the escape hatch works too well.

reliability
6
usefulness
7
cost
5
longevity
6
Agree with La Inversora?
La JefaThe CTO

No published prices anywhere and usage-based billing behind an API key, so I cannot model sixty engineers, which ends the conversation before security starts.

5.0
Reasoning and trade-offs · AI analysis

I cannot budget this. Billing is by usage against a key, and the documentation states the key requirement without stating a price, so the number I would take to finance does not exist. Sixty engineers times an unknown rate is not a forecast.

It does run unattended, which is a point in its favour, and remote sessions are documented, which is a second procurement question about where those sessions live. No SSO, no audit log, no retention policy. Not yet, and the blocker is a rate card.

reliability
5
usefulness
6
cost
4
longevity
5
Agree with La Jefa?
El HackerThe tinkerer

Apache-2.0, MCP tools attach, and an external provider can replace the built-in models, so the vendor's endpoint is a default rather than a cage.

7.0
Reasoning and trade-offs · AI analysis

Apache-2.0 on the client, which means I can read every path the code takes before it talks to anybody. The part that decides it for me is that an outside provider can be configured in place of the built-in models, so the vendor's infrastructure is the default rather than the only road.

MCP servers attach, so my own tools are available without a shim. Installation covers a shell script and PowerShell as well as the usual, which is more platform care than most projects this size bother with.

reliability
7
usefulness
8
cost
7
longevity
6
Agree with El Hacker?