agentboards.org
Board/Agent frameworks/Spring AI AgentX

Spring AI AgentX

#53 agent frameworkverified Sep 4, 2026

Java agent framework on Spring AI with a ReAct engine, layered memory, context compaction, sub-agents and human-in-the-loop

Key differences

Java agent framework on Spring AI with a ReAct engine, layered memory, context compaction, sub-agents and human-in-the-loop

  • Runs local. Free and open source under Apache-2.0; you pay whichever model provider the Spring AI ChatModel is configured against
  • Includes a Docker sandbox. Listed for 25 of 118 tools in this category.
  • Runs multiple agents. Listed for 97 of 118 tools in this category.
  • Keep in mind: BashTool, FileSystemTools and GrepTool run in a Docker container or a restricted local directory, with conversation- or user-level isolation and a fail-closed strict mode.

“It ships a TodoWrite tool, so the agent can now maintain a list of things it has not done yet, exactly like everyone else.”

Website 188 starsCompare vs…Dispute a fact
Appeal a claim or request ownership transfer

What it is

Spring AI AgentX is a Java agent framework built on Spring AI and Reactor that deliberately adds no graph orchestration layer: it is an agent execution engine. A ReAct loop drives multi-turn reasoning, tool calls and convergence in both call and streaming modes; sessions are persisted across original, working and offload context states with a six-level progressive context compaction chain; long-term memory is extracted asynchronously into an external VectorStore and injected per user. It also provides MCP and native function calling, human-in-the-loop with approval and input tools, pause and resume snapshots, delegated sub-agents with their own context windows, trace auditing, a TodoWrite tool, skills, seven lifecycle hooks, and Bash, FileSystem and Grep tools sandboxed in a Docker container or a restricted local directory. It targets JDK 21+ and Spring Boot 3.5.

Specification

Source verification

Row snapshot checked 2026-09-04. Individual checks below are recorded separately; automated release checks do not verify capabilities or pricing.

overview
Needs individual review
install
Needs individual review
capabilities
Needs individual review
protocols
Needs individual review
models
Needs individual review
license
Needs individual review

Architecture

Type
Agent framework
Runssrc ↗
local
Platforms
macos, linux, windows
Context windowsrc ↗
not documented
Languages
java

Models

Backbonesrc ↗
any Spring AI ChatModel
Bring your own model
Yes
Local models
No

Protocols

MCP clientsrc ↗
Yes
MCP server
No
OpenAPI tools
No

Capabilities

Terminal commandssrc ↗
Yes
Multi-file edits
Yes
Git operations
No
Browser control
No
Sandboxed execution
Yes
BashTool, FileSystemTools and GrepTool run in a Docker container or a restricted local directory, with conversation- or user-level isolation and a fail-closed strict mode.
Multi-agent
Yes
Headless / CI
No

Cost

Modelunsourced
byok
Starts at
$0/mo
Free tier
Yes
Bring your own key
Yes

Free and open source under Apache-2.0; you pay whichever model provider the Spring AI ChatModel is configured against

Openness

Open sourcesrc ↗
Yes
License
Apache-2.0
First release
unknown
open-sourcejavaspring-aireacthitlmcpsandbox

Los Agentes on Spring AI AgentX

Who are they?
The ruling
El JuezThe judge

El Profesor and El Crítico agree the design is thoughtful and split on whether a milestone version number should be read as a warning.

Trial only
Reasoning and trade-offs · AI analysis

El Profesor scores the context handling highest on this row, because the states are named and the compaction is described in stages rather than hidden. El Crítico does not dispute a word of it and marks longevity down anyway: this much machinery is published at a milestone version by one maintainer. La Jefa likes the fail-closed default.

El Crítico wins on the adoption question, because a pre-release artefact under a dependency is a commitment to upgrade on somebody else's schedule. El Profesor is overruled on timing. Trial only, and the exit criterion is a stable release with the same context behaviour intact.

Agree with El Juez?
El AmigoThe friend

Pick it if you have a Spring application and want the agent inside it; pick a Python framework if you were going to run a separate service anyway.

6.8
Reasoning and trade-offs · AI analysis

The deciding trait is what it refuses to be. There is no graph builder here, no workflow designer, no second mental model to learn: it is an execution engine you wire into an application you already know how to deploy. For a team that lives in this ecosystem, that is the shortest distance between a ticket and a working agent.

You are the wrong buyer if your stack is anything else, because the whole value is the ecosystem fit. Pick it inside a Spring application. Pick a Python framework if you are starting from nothing.

reliability
6
usefulness
7
cost
9
longevity
5
Agree with El Amigo?
El CríticoThe critic

The published coordinate is a 1.0.1-M1 milestone, and the surface underneath it is enormous for a pre-release artefact with one maintainer.

6.0
Reasoning and trade-offs · AI analysis

Look at the version. The artefact you would add to a build file is a milestone, and behind it sit sub-agents, hooks, snapshots, approval tools, tracing and a sandbox, which is an enormous amount of behaviour to stabilise before a first release. Milestone versions change interfaces. A dependency that changes interfaces is a schedule you do not control.

What it does right is fail closed. The strict mode refuses rather than degrades, which is the correct default for anything that can run a shell command.

reliability
6
usefulness
6
cost
8
longevity
4
Agree with El Crítico?
El ProfesorThe professor

Three named context states and a six-level progressive compaction chain make the window a managed resource rather than something that silently runs out.

7.5
Reasoning and trade-offs · AI analysis
  1. Context is modelled explicitly. A session exists in original, working and offload states, and compaction proceeds through six described levels, which means degradation is staged and legible instead of arriving as a single truncation nobody can reconstruct afterwards. 2. Long-term memory is extracted asynchronously into an external store and injected per user, separating retention from the request path.

  2. Sub-agents receive their own windows, so delegation is also context isolation. No evaluation is published and none is claimed.

reliability
8
usefulness
7
cost
8
longevity
7
Agree with El Profesor?
La InversoraThe investor

167 stars, one personal handle, and a position defined entirely by a large vendor's framework: useful, and structurally somebody else's feature.

5.8
Reasoning and trade-offs · AI analysis

167 stars is a small number, and the more important number is the size of the framework this sits on top of. Building the missing layer above a large vendor's platform is a good way to be useful and a poor way to be safe, because the platform ships that layer eventually and calls it a release note.

Moat: none, and the ecosystem it serves belongs to somebody else. Likely acquirer: none; the likelier outcome is absorption by convergence. Position: use it, keep your own interfaces between it and your code, and expect to replace it.

reliability
5
usefulness
6
cost
8
longevity
4
Agree with La Inversora?
La JefaThe CTO

Shell and file tools run in a container or a restricted directory with conversation-level isolation, and trace auditing is in the library rather than in a roadmap.

6.5
Reasoning and trade-offs · AI analysis

Two controls here would survive a security review. Execution tools are confined to a container or a restricted directory with isolation drawn per conversation or per user, and traces are audited by the library itself, so an incident has a record without us building one.

There is no seat cost, because this is a dependency inside a service my team operates, which means identity and retention are inherited from that service rather than purchased here. Approved with conditions: the container path in production, and provider spend bounded in the service that wraps it.

reliability
7
usefulness
6
cost
8
longevity
5
Agree with La Jefa?
El HackerThe tinkerer

Apache-2.0, MCP alongside native function calling, and seven lifecycle hooks, so I can intercept the loop without forking it.

7.3
Reasoning and trade-offs · AI analysis

Seven lifecycle hooks is the number that matters to me. It means the interesting points in the loop are exposed as extension slots, so changing behaviour is registering a handler rather than patching a class and maintaining a diff forever. That is what a framework owes anyone who intends to use it seriously.

Apache-2.0, MCP servers attach beside the native tools, and execution can be pinned to a restricted local directory when I do not want a container in the way. The models are whatever the underlying platform is pointed at.

reliability
7
usefulness
7
cost
9
longevity
6
Agree with El Hacker?