agentboards.org

HarnessX

#65 agent frameworkverified Sep 4, 2026

Python framework that separates the model from the harness, so tools, memory, processors, trace and sandbox compose into an agent

Key differences

Python framework that separates the model from the harness, so tools, memory, processors, trace and sandbox compose into an agent

  • Runs local and sandbox. Free and open source under MIT; you supply the provider API key the agents run on
  • Includes a Docker sandbox. Listed for 25 of 118 tools in this category.
  • Supports headless CI workflows. Listed for 33 of 118 tools in this category.
  • Keep in mind: The sandbox dimension ships local, Docker and E2B backends.

“It connects to Feishu, Telegram, Slack, Discord and DingTalk, so the agent has wider messaging coverage than most of its users.”

Website Docs 483 starsCompare vs…Dispute a fact
Appeal a claim or request ownership transfer

What it is

HarnessX is a harness foundry: a Python framework where `agent = model.agentic(harness)` splits provider routing and per-role model assignment from the behaviour pipeline of tools, memory, processors, trace and sandbox. Any behaviour is a Processor and processors combine with the pipe operator, so switching an agent from coding to research or adding guardrails is a configuration change. Sandboxes run locally, in Docker or on E2B, and a tool registry with built-ins backs the loop. A meta-harness observes its own trajectories and proposes better processor combinations, and reward-annotated trajectories from runs can feed RL fine-tuning through VERL. It ships an `hx` CLI with interactive, single-prompt and resume modes, a Lab UI, and a gateway that connects the agent to Feishu, Telegram, Slack, Discord or DingTalk.

Specification

Source verification

Row snapshot checked 2026-09-04. Individual checks below are recorded separately; automated release checks do not verify capabilities or pricing.

overview
Needs individual review
install
Needs individual review
license
Needs individual review
pricing
Needs individual review
capabilities
Needs individual review
models
Needs individual review
docs
Needs individual review

Architecture

Type
Agent framework
Runssrc ↗
local, sandbox
Platforms
macos, linux
Context windowsrc ↗
not documented
Languages
python

Models

Backbonesrc ↗
Anthropic, any configured provider
Bring your own model
Yes
Local models
No

Protocols

MCP clientunsourced
No
MCP server
No
OpenAPI tools
No

Capabilities

Terminal commandssrc ↗
Yes
Multi-file edits
Yes
Git operations
No
Browser control
No
Sandboxed execution
Yes
The sandbox dimension ships local, Docker and E2B backends.
Multi-agent
Yes
Headless / CI
Yes

Cost

Modelsrc ↗
byok
Starts at
$0/mo
Free tier
Yes
Bring your own key
Yes

Free and open source under MIT; you supply the provider API key the agents run on

Openness

Open sourcesrc ↗
Yes
License
MIT
First release
unknown
open-sourcepythonframeworkprocessorssandboxevolutiongaia

Los Agentes on HarnessX

Who are they?
The ruling
El JuezThe judge

El Profesor admires the same self-improving loop El Crítico wants a stopping rule for, and the panel's split is really about who is allowed to change the agent.

Trial only
Reasoning and trade-offs · AI analysis

El Profesor scores the architecture high because separating model routing from behaviour is a clean line few frameworks draw. El Crítico scores reliability low because one of those behaviours rewrites the others: a meta-layer proposes new combinations, and nothing published says what makes a proposal good enough to keep.

El Crítico wins. An unevaluated self-modification loop is a research result, not a dependency, and El Profesor is overruled on readiness rather than on design. Trial only, and the exit criterion is a run where you can say why the composition changed and prove it improved something.

Agree with El Juez?
El AmigoThe friend

Pick it if you are building several different agents and tired of rewriting the same loop; pick a finished agent if you only need one.

6.5
Reasoning and trade-offs · AI analysis

The deciding trait is that turning a coding agent into a research agent is a configuration change rather than a rewrite. Behaviours snap together, so the second agent you build costs a fraction of the first, and the third is nearly free. Anyone who has copied a loop between two projects and watched them drift will recognise the appeal.

You are the wrong buyer if you need one working agent today, because assembling is still assembling. Pick it when you are building a family of them. Pick a finished tool when you are building one.

reliability
6
usefulness
7
cost
8
longevity
5
Agree with El Amigo?
El CríticoThe critic

A meta layer observes the agent and proposes better combinations of its own behaviours, and nothing published states what makes a proposal acceptable.

5.5
Reasoning and trade-offs · AI analysis

The risky part is the part being advertised. A meta layer watches the agent run and proposes different arrangements of its own behaviour, which means the thing you tested on Monday is not necessarily the thing running on Thursday. No acceptance criterion is documented, no rollback is described, and no ceiling is placed on how far a proposal may drift from what you configured.

What it does right is name the tiers of isolation. Local, container and hosted are three separate documented backends rather than one setting that means different things.

reliability
5
usefulness
6
cost
6
longevity
5
Agree with El Crítico?
El ProfesorThe professor

The composition is the argument: model routing on one side, a pipeline of behaviours on the other, joined by an operator rather than by inheritance.

7.0
Reasoning and trade-offs · AI analysis
  1. The central expression puts provider routing and per-role model assignment on one side and the behaviour pipeline on the other, which is a real separation rather than a naming convention.

  2. Behaviours compose with an operator, so an agent's definition is an expression that can be read left to right, and two agents can be compared by comparing their expressions.

  3. Runs produce reward-annotated records intended to feed fine-tuning, which is a coherent ambition. No result from that pipeline is published, so the loop is described rather than demonstrated.

reliability
7
usefulness
7
cost
7
longevity
7
Agree with El Profesor?
La InversoraThe investor

454 stars and a research agenda dressed as a framework: interesting people, no revenue mechanism, and an idea larger than any product around it.

6.0
Reasoning and trade-offs · AI analysis

454 stars, a small research-flavoured organisation, and no price anywhere. What is being built here is closer to a lab's tooling than to a company's product, and that is not a criticism, it is a description of where the effort goes: into the idea rather than into the funnel.

Moat: the research, briefly, until it is published or copied. Likely acquirer: a lab hiring the team, which is how this shape of project usually resolves. Position: watch it, borrow the ideas, and do not put it under anything that has to run next quarter.

reliability
5
usefulness
6
cost
8
longevity
5
Agree with La Inversora?
La JefaThe CTO

It runs one prompt non-interactively and exits, which makes it schedulable; the hosted sandbox backend makes it a vendor review before that helps me.

5.8
Reasoning and trade-offs · AI analysis

One flag runs a single prompt and exits, so this can sit in a pipeline and be scheduled like anything else we operate. That is more than most of this category offers and it is the reason I read the rest of the row.

The rest is a procurement exercise. One of the isolation backends is a third-party hosted service, which means our code reaches a vendor I have not reviewed unless we pin the local option. No SSO, no audit log, no retention policy. Approved with conditions: the hosted backend disabled by default.

reliability
5
usefulness
6
cost
7
longevity
5
Agree with La Jefa?
El HackerThe tinkerer

MIT and a uv install, and every behaviour is a Python object I can subclass, but the tool layer is its own registry so my MCP servers stay outside.

7.3
Reasoning and trade-offs · AI analysis

MIT, and it installs into a virtual environment I control rather than into a global path. Every behaviour is an ordinary Python object, so extending it means writing a class rather than filing a feature request, and reading it means reading Python instead of a plugin manifest.

The gap is the tool layer. Capability comes from an internal registry, and there is no MCP client, so servers I already run have to be re-exposed as native tools before they exist here. That is a day of work I did not want to spend.

reliability
7
usefulness
7
cost
9
longevity
6
Agree with El Hacker?