agentboards.org

Schaltwerk

#167 agent harnessunverified rowv0.13.7

Desktop app running ten agent CLIs natively, each session in its own git worktree, with an MCP orchestrator

Key differences

Desktop app running ten agent CLIs natively, each session in its own git worktree, with an MCP orchestrator

  • Runs local. Free and MIT-licensed; you supply at least one agentic coding CLI and its credentials, public API or a private endpoint
  • Acts as an MCP server. Listed for 37 of 194 tools in this category.
  • Runs multiple agents. Listed for 165 of 194 tools in this category.
  • Keep in mind: Schaltwerk exposes an MCP server so an orchestrator agent can create sessions and coordinate the others.

“It drives ten different agent CLIs, which is nine more than your team will ever agree on.”

Website Docs 287 starsCompare vs…Dispute a fact
Appeal a claim or request ownership transfer

What it is

Schaltwerk runs agentic coding CLIs directly rather than through wrappers — GitHub Copilot CLI, Claude Code, Kilo Code, OpenCode, Codex, Pi, Gemini, Qwen, Factory Droid and Amp, plus a terminal-only mode — giving each session its own git worktree so several run at once. Work starts from markdown specs that stay visible and reusable: if an agent drifts, discard the worktree, refine the spec and relaunch. Reviews are GitHub-style diffs with inline comments you paste back to the agent, with a second terminal for manual testing and a squash-merge or PR when a session is marked reviewed. A built-in MCP server lets one terminal agent orchestrate the others, creating sessions and coordinating parallel work.

Specification

Source verification

Row snapshot checked not yet. Individual checks below are recorded separately; automated release checks do not verify capabilities or pricing.

readme
Needs individual review
docs
Needs individual review
install
Needs individual review
license
Needs individual review
capabilities
Needs individual review

Architecture

Type
Agent harness
Runssrc ↗
local
Platforms
macos, linux, windows
Context windowunsourced
not documented
Languages
any

Models

Backboneunsourced
via managed agents (GitHub Copilot CLI, Claude Code, Kilo Code, OpenCode, Codex, Pi, Gemini, Qwen, Factory Droid, Amp)
Bring your own model
Yes
Agents can be configured against private endpoints such as Azure or self-hosted APIs.
Local models
No

Protocols

MCP clientunsourced
No
MCP server
Yes
OpenAPI tools
No

Capabilities

Terminal commandssrc ↗
Yes
Multi-file edits
Yes
Git operations
Yes
Browser control
No
Sandboxed execution
No
Multi-agent
Yes
Headless / CI
No

Cost

Modelunsourced
byok
Starts at
$0/mo
Free tier
Yes
Bring your own key
Yes

Free and MIT-licensed; you supply at least one agentic coding CLI and its credentials, public API or a private endpoint

Openness

Open sourcesrc ↗
Yes
License
MIT
First release
unknown
open-sourceworktreesparallel-agentsspec-drivenmcpdiff-review

Los Agentes on Schaltwerk

Who are they?
The ruling
El JuezThe judge

El Hacker calls the orchestration server the best idea here and La Jefa cannot approve the machines it would run on, which is the whole disagreement.

Adopt with conditions
Reasoning and trade-offs · AI analysis

El Hacker scores it high because one agent can create and coordinate the others through a protocol he already speaks. La Jefa scores it low because her Linux engineers are told the platform is beta, which decides adoption before any feature does. El Crítico is separately unhappy about a review loop that runs through a human clipboard.

La Jefa is right for her fleet and wrong to generalise; El Hacker is right for anyone on a supported desktop, which is most individual readers. Adopt with conditions, the condition being that everyone who needs it is on macOS or Windows before you standardise.

Agree with El Juez?
El AmigoThe friend

Pick it if you would rather rewrite the brief than argue with a drifting agent; pick a chat-first tool if you like steering mid-run.

7.3
Reasoning and trade-offs · AI analysis

The deciding trait is that the work starts from a written spec you keep. When an agent wanders, you do not coax it back: you throw the attempt away, sharpen the sentence that misled it, and run again from a clean tree. That converts a frustrating hour of correction into two minutes of editing, and the spec is still there for the next attempt.

It suits people who can write a brief and dislike negotiating with a model. Pick something conversational if the plan only becomes clear while you talk.

reliability
7
usefulness
7
cost
9
longevity
6
Agree with El Amigo?
El CríticoThe critic

Review comments are written inline and then pasted back to the agent by hand, so the path from finding a defect to fixing it runs through a human and a clipboard.

6.5
Reasoning and trade-offs · AI analysis

The feedback loop is manual at its most important point. A reviewer marks a problem, then copies that text into a session and hopes the agent reads it in the context it was written about. Nothing carries the file, the line or the surrounding diff automatically, so precision depends on whoever is doing the copying, and precision is exactly what a correction needs.

What it does right is the escape hatch. A session that went wrong is deleted as a worktree, with no cleanup and nothing left behind in the branch you care about.

reliability
6
usefulness
6
cost
8
longevity
6
Agree with El Crítico?
El ProfesorThe professor

The unit of work is a markdown specification that outlives the attempt, so a failed run is treated as evidence about the brief rather than about the model.

7.0
Reasoning and trade-offs · AI analysis
  1. Persisting the instruction separately from the session is the design decision that makes iteration meaningful. If the prompt is discarded with the attempt, every retry is a new experiment with an uncontrolled variable; if it survives and is edited deliberately, the change between runs is known. 2. Reuse of the same specification across attempts also permits comparison between agents under identical instructions.

  2. No evaluation is published, and none is claimed, which keeps the argument architectural.

reliability
7
usefulness
7
cost
7
longevity
7
Agree with El Profesor?
La InversoraThe investor

287 stars and a single-handle maintainer with no company attached, so this is a personal project with a good idea rather than an asset with a future.

5.8
Reasoning and trade-offs · AI analysis

The star count is the honest number here: small, early, and not yet evidence of anything except taste. There is no entity, no funding and nothing to sell, which means the eighteen-month question reduces entirely to whether one person keeps finding this interesting after the novelty passes.

Moat: none, and the concept is copyable in a weekend by anyone who wants it. Likely path: quiet maintenance, or the pattern showing up inside a funded product with better distribution. Position: use it freely, and keep your specs in the repository so leaving costs nothing.

reliability
5
usefulness
6
cost
8
longevity
4
Agree with La Inversora?
La JefaThe CTO

macOS 11 and Windows 10 are supported, Linux is beta and WSL is not supported at all, which decides this for an organisation before any feature does.

5.5
Reasoning and trade-offs · AI analysis

Support matrices decide adoption, and this one splits my engineers into two tiers. The people on Linux are told their platform is beta and the ones who work through WSL are told nothing works, which means either an exception process or an unhappy third of the department. Neither is worth it for a free tool.

There is also no unattended execution, so it never becomes a measurable step, and no directory integration to point at in an audit. Not yet: revisit if Linux leaves beta.

reliability
5
usefulness
5
cost
8
longevity
4
Agree with La Jefa?
El HackerThe tinkerer

MIT, installed with a brew cask, and it exposes its own MCP server so one agent can create sessions and drive the others without me writing the glue.

7.8
Reasoning and trade-offs · AI analysis

Shipping an MCP server rather than only consuming one is the inversion I keep waiting for. It means the application is addressable: an agent I already trust can open sessions here, coordinate parallel work and report back, and the orchestration logic lives in my prompt rather than in somebody's product roadmap.

Installation is one brew command, agents can point at a private or self-hosted endpoint instead of a public API, and the licence keeps a fork legal. That is close to everything I ask for from a desktop tool.

reliability
8
usefulness
8
cost
9
longevity
6
Agree with El Hacker?