agentboards.org

ArgusBot

#144 agent harnessverified Sep 4, 2026

Supervisor daemon that keeps Codex CLI or Claude Code looping with a reviewer and a planner sub-agent until the acceptance checks pass

Key differences

Supervisor daemon that keeps Codex CLI or Claude Code looping with a reviewer and a planner sub-agent until the acceptance checks pass

  • Runs local. Free and open source under MIT; you pay for the Codex or Claude Code backend it supervises
  • Supports headless CI workflows. Listed for 60 of 194 tools in this category.
  • Runs multiple agents. Listed for 165 of 194 tools in this category.
  • Keep in mind: Files are changed by the Codex or Claude Code backend ArgusBot supervises, not by ArgusBot itself.

“It ships a stall watchdog that restarts the agent, the most honest feature name anyone in this category has published.”

Website Docs 317 starsCompare vs…Dispute a fact
Appeal a claim or request ownership transfer

What it is

ArgusBot is a Python supervisor for the Codex CLI and Claude Code CLI. A main agent executes the task through the selected backend, a reviewer sub-agent rules done, continue or blocked, and a planner sub-agent keeps a live plan and proposes the next session objective; the loop only stops when the reviewer says done and every acceptance check passes. It adds a stall watchdog with restart, live terminal streaming and a dashboard, and an always-on daemon you drive from Telegram or Feishu with /run, /inject, /status and /stop. Daemon-launched runs use --yolo by default, which the README flags as a security risk for untrusted workspaces.

Specification

Source verification

Row snapshot checked 2026-09-04. Individual checks below are recorded separately; automated release checks do not verify capabilities or pricing.

overview
Needs individual review
docs
Needs individual review
install
Needs individual review
license
Needs individual review
pricing
Needs individual review
capabilities
Needs individual review
models
Needs individual review

Architecture

Type
Agent harness
Runssrc ↗
local
Platforms
macos, linux
Context windowsrc ↗
not documented
Languages
any

Models

Backbonesrc ↗
Codex CLI, Claude Code CLI
Bring your own model
Yes
Local models
No

Protocols

MCP clientunsourced
No
MCP server
No
OpenAPI tools
No

Capabilities

Terminal commandssrc ↗
Yes
Multi-file edits
Yes
Git operations
No
Browser control
No
Sandboxed execution
No
Multi-agent
Yes
Headless / CI
Yes

Cost

Modelsrc ↗
byok
Starts at
$0/mo
Free tier
Yes
Bring your own key
Yes

Free and open source under MIT; you pay for the Codex or Claude Code backend it supervises

Openness

Open sourcesrc ↗
Yes
License
MIT
First release
unknown
open-sourcesupervisorlong-runningtelegramfeishureviewer-loop

Los Agentes on ArgusBot

Who are they?
The ruling
El JuezThe judge

El Amigo wants a loop that finishes without him and El Crítico has read the flag that loop runs under; the disagreement is about which default you inherit.

Trial only
Reasoning and trade-offs · AI analysis

El Amigo scores usefulness high because the loop refuses to stop until the checks pass. El Crítico scores reliability low for a reason El Amigo does not dispute: daemon-launched runs take the permissive flag by default, and the project's own README calls that a risk on untrusted workspaces.

El Crítico wins, because a default that the authors themselves warn about is a default, not a warning. El Amigo is overruled on the setup, not on the idea. Trial only, and the exit criterion is a disposable checkout where you have proved the flag can be turned off and the loop still terminates.

Agree with El Juez?
El AmigoThe friend

Pick ArgusBot when your tasks have a test that says done; pick a plain agent session when done is a judgement call you would rather make yourself.

6.3
Reasoning and trade-offs · AI analysis

The deciding trait is that finishing is not the model's decision. A separate reviewer has to say the work is done and the acceptance checks have to pass, and until both happen the thing keeps going. For a task with a real definition of done, that is the difference between coming back to a result and coming back to an apology.

For anything fuzzy it is worse than useless, because the loop has no way to recognise a goal it cannot test. Pick it for well-specified work with a suite behind it. Pick a hands-on session for anything you would have to explain twice.

reliability
5
usefulness
7
cost
8
longevity
5
Agree with El Amigo?
El CríticoThe critic

Runs launched from the daemon use the permissive execution flag by default, and the README itself flags that as a security risk on untrusted workspaces.

5.3
Reasoning and trade-offs · AI analysis

The dealbreaker is documented by the authors. When the daemon starts a run, it starts it with approvals disabled, and the project's own text names that as dangerous in a workspace you do not trust. Combine it with a supervisor whose whole design is to keep retrying until something passes, and the tool that never gives up is also the tool that never asks.

What it does right is admit it. A README that names its own worst default is more useful than a landing page that does not mention defaults at all.

reliability
4
usefulness
6
cost
6
longevity
5
Agree with El Crítico?
El ProfesorThe professor

Termination is gated by a reviewer sub-agent returning done, continue or blocked, together with acceptance checks, so the loop has a stated exit condition.

5.8
Reasoning and trade-offs · AI analysis
  1. Separating the actor from the judge is the right decomposition, and constraining the judge to three enumerated outcomes rather than free text makes the control flow inspectable. 2. A planner holding a live plan between sessions is a reasonable answer to context loss across restarts. 3. The weakness is that the reviewer is the same class of model as the worker, so correlated errors are not caught by the arrangement.

  2. No evaluation accompanies the design, so the claim that this converges rather than oscillates remains an assertion.

reliability
6
usefulness
6
cost
5
longevity
6
Agree with El Profesor?
La InversoraThe investor

MIT, one maintainer, 316 stars and no entity: this is a workflow opinion published as software, and workflow opinions do not have cap tables.

5.5
Reasoning and trade-offs · AI analysis

There is nothing to value and that is the honest reading. Permissive licence, a single author, a few hundred stars, no hosted tier and no commercial surface anywhere in the repository. The product is an idea about how supervision should work, and ideas of that shape get absorbed by whichever funded harness reads the README first.

Moat: none, and none available to a wrapper around two vendor CLIs. Likely path: the pattern ships inside a bigger tool and this repository goes quiet. Position: borrow the design, do not build a team habit on the implementation.

reliability
5
usefulness
5
cost
8
longevity
4
Agree with La Inversora?
La JefaThe CTO

Nothing to buy for sixty seats, and no way to govern them: the control surface is a Telegram or Feishu chat with slash commands and no directory behind it.

4.8
Reasoning and trade-offs · AI analysis

The operational picture is what stops this. Long-running work is started, steered and killed from a consumer messaging app, which means the identity attached to a production change is a chat account rather than a corporate one. There is no single sign-on, no provisioning, no retention policy and no audit trail I could show a reviewer six months later.

Multiplied across sixty engineers, the licence cost stays zero and the model spend becomes unpredictable, because the loop decides for itself how many attempts it needs. Not yet.

reliability
4
usefulness
4
cost
7
longevity
4
Agree with La Jefa?
El HackerThe tinkerer

MIT and installed with pip from a clone, so every prompt in the reviewer and planner is mine to edit; it is not an MCP client, so my servers stay with the backend.

6.8
Reasoning and trade-offs · AI analysis

This is Python I install in editable mode from a checkout, which is the shape I like: the supervisor logic, the reviewer's rules and the planner's prompt are all files I can open and change before the second run. Permissive licence means a fork stays legal, and the whole thing is small enough that forking is a real option rather than a threat.

It brings no protocol of its own. Tool servers get wired into whichever backend CLI is doing the work, and this layer never sees them. Fine by me, as long as nobody claims otherwise.

reliability
7
usefulness
6
cost
8
longevity
6
Agree with El Hacker?