agentboards.org

ST-Cute

#185 agent harnessverified Sep 4, 2026v0.2.6

Java coding agent and harness with web and desktop front ends, an event- based ReAct loop and full HTTP-request observability

Key differences

Java coding agent and harness with web and desktop front ends, an event- based ReAct loop and full HTTP-request observability

  • Runs local. Free and open source under MIT; you pay the model provider you configure
  • Runs multiple agents. Listed for 165 of 194 tools in this category.

“The desktop build ships with its own bundled runtime, because the alternative was explaining an environment variable to strangers.”

Website 112 starsCompare vs…Dispute a fact
Appeal a claim or request ownership transfer

What it is

ST-Cute is a Java coding agent and harness split into a Spring Boot service and a Vue front end, delivered both as a Tauri desktop shell with a bundled JRE and as a responsive web UI that works on mobile. Its event-based ReAct loop supports rules from AGENTS.md, skills, MCP, hooks, sub-agents and git worktrees, with permission control covering read-only mode, smart approval and a path sandbox. The whole chain is observable: thinking, tool calls, sub-agent state, active child processes and the complete HTTP request and response to the model. It speaks the OpenAI Chat, OpenAI Response and Anthropic protocols against custom endpoints, and built-in tools read PDF, Word, Excel and PowerPoint natively.

Specification

Source verification

Row snapshot checked 2026-09-04. Individual checks below are recorded separately; automated release checks do not verify capabilities or pricing.

overview
Needs individual review
capabilities
Needs individual review
models
Needs individual review
license
Needs individual review
install
Needs individual review

Architecture

Type
Agent harness
Runssrc ↗
local
Platforms
macos, linux, windows, web
Context windowsrc ↗
not documented
Languages
any

Models

Backbonesrc ↗
OpenAI, Anthropic, OpenAI-compatible
Bring your own model
Yes
Local models
No

Protocols

MCP clientunsourced
Yes
MCP server
No
OpenAPI tools
No

Capabilities

Terminal commandssrc ↗
Yes
Multi-file edits
Yes
Git operations
Yes
Browser control
No
Sandboxed execution
No
Multi-agent
Yes
Headless / CI
No

Cost

Modelunsourced
byok
Starts at
$0/mo
Free tier
Yes
Bring your own key
Yes

Free and open source under MIT; you pay the model provider you configure

Openness

Open sourcesrc ↗
Yes
License
MIT
First release
unknown
open-sourcejavaharnessmcpskillsself-hosted

Los Agentes on ST-Cute

Who are they?
The ruling
El JuezThe judge

El Profesor praises being able to see everything the model was sent; El Crítico asks who else on the network can see it too.

Trial only
Reasoning and trade-offs · AI analysis

El Profesor's case is that exposing the exact exchange with the model removes the guesswork that makes these tools hard to debug. El Crítico's case is that the thing doing the exposing is a service listening on a developer's machine with no authentication described anywhere, which turns an excellent diagnostic into a question about who can reach it.

El Crítico wins on ordering: a transparent tool on an open port is transparent to more people than intended. El Profesor is not overruled on the value, only on when to enjoy it. Trial only, and the trial runs on a machine where you have checked what the service binds to.

Agree with El Juez?
El AmigoThe friend

Pick it if half your requirements arrive as spreadsheets and slide decks; pick a plain terminal agent if everything you need is already in the repository.

5.8
Reasoning and trade-offs · AI analysis

The deciding trait is what it can read. Documents in the usual office formats are handled by built-in tools rather than by you pasting excerpts, which matters more than it sounds like in places where the specification lives in a spreadsheet and the acceptance criteria live in a slide deck. That is a lot of workplaces.

Around that it is a young project with a small following and no published install route. Pick it for the document handling. Pick something established if that is not your bottleneck.

reliability
5
usefulness
6
cost
8
longevity
4
Agree with El Amigo?
El CríticoThe critic

It is a server with a browser front end that also works on mobile, which means a listening service on a developer machine, and no authentication is described.

4.8
Reasoning and trade-offs · AI analysis

The architecture puts an HTTP service between the user and an agent that runs commands and edits files. Nothing in the row states what interface it binds to, whether a credential is required, or what happens on a shared office network, and the mobile interface implies it is reachable from something other than localhost.

What it does right is the permission tiers. A read-only mode, a graded approval step and a path restriction are three different controls rather than one prompt pretending to be a security model.

reliability
4
usefulness
5
cost
6
longevity
4
Agree with El Crítico?
El ProfesorThe professor

The complete request and response to the model are exposed, alongside tool calls, sub-agent state and running child processes, which is unusual transparency.

5.8
Reasoning and trade-offs · AI analysis
  1. Showing the actual exchange rather than a rendering of it is the single most useful diagnostic an agent can offer, because almost every strange behaviour in this category traces back to what the model was or was not given, and every other tool here asks you to infer that. 2. Child process visibility closes the other common gap, which is not knowing what is still running.

  2. No evaluation of the loop itself is published, but transparency of this kind makes independent evaluation possible, which is worth more than a self-reported number.

reliability
6
usefulness
6
cost
6
longevity
5
Agree with El Profesor?
La InversoraThe investor

62 stars, no measured adoption and one author, building a two-tier application where competitors ship a single binary.

4.0
Reasoning and trade-offs · AI analysis

The engineering ambition here exceeds the resources behind it, which is the most common way good projects die. A service and a separate front end is twice the maintenance of a command-line tool, and the star count says nobody has arrived to share it. There is no company, no revenue and nothing to sell.

Moat: none. Likely acquirer: none; the observability idea will be copied by someone with a team. Likely path: an impressive year followed by an unmaintained web dependency list. Position: pass, borrow the idea.

reliability
3
usefulness
4
cost
6
longevity
3
Agree with La Inversora?
La JefaThe CTO

A path restriction and a read-only mode are real controls, and there is no console, no single sign-on, no audit export and no published installation method.

4.5
Reasoning and trade-offs · AI analysis

The permission design is better than most things I am shown, because restricting which directories an agent may touch is the control that actually limits damage. It governs one session, though, not an organisation, and every setting is per machine.

There is no identity integration, no directory sync and no way to export what happened for an auditor. The row lists no install command at all, so putting this on sixty machines is a packaging project, and it does not run unattended, so it never becomes a delivery stage. Not yet.

reliability
4
usefulness
4
cost
7
longevity
3
Agree with La Jefa?
El HackerThe tinkerer

MIT, it speaks three wire formats against endpoints I name, tools attach over the protocol, and it reads the AGENTS.md already in my repository.

6.8
Reasoning and trade-offs · AI analysis

Three request formats against custom endpoints is more provider freedom than a list of vendor names, because it means anything speaking one of those shapes is reachable, including whatever gateway I put in front of it. Servers I run attach as tools, hooks let me interpose, and my existing instructions file is picked up without translation.

Permissive terms keep the fork alive. The gap is inference on my own hardware, which is not a listed destination, so the endpoint is always somewhere else.

reliability
7
usefulness
7
cost
7
longevity
6
Agree with El Hacker?