agentboards.org
Board/Terminal agents/Factory Droid

Factory Droid

#65 overall#28 terminal agentverified Sep 2, 2026

Factory's coding agent available as a terminal CLI, web app, Slack and Teams bot, and headless runner for CI

Key differences

Factory's coding agent available as a terminal CLI, web app, Slack and Teams bot, and headless runner for CI

  • Runs local and cloud and sandbox. Pro $20/mo, Plus $100/mo, Max $200/mo, Teams $60/mo plus $40 per seat; Business and Enterprise custom; usage governed by rolling rate limits
  • Includes a Docker sandbox. Listed for 26 of 125 tools in this category.
  • Supports headless CI workflows. Listed for 55 of 125 tools in this category.
  • Keep in mind: The first-party Droid Control plugin (droid plugin install droid-control@factory-plugins) ships an `agent-browser` driver - a Playwright-backed CLI with Chrome DevTools Protocol support that navigates, fills forms, clicks and captures screenshots, powering /qa-test and /demo.

“Charges $200 a month for the Max plan and still governs you with rolling rate limits, so the ceiling is the product.”

Website Docs 41 starsCompare vs…Dispute a fact
Appeal a claim or request ownership transfer

What it is

Droid is Factory's software development agent, run from the Droid CLI, the app.factory.ai web app, chat tools, and a headless `droid exec` mode for scripts and CI. It supports MCP tools, hooks, custom droids and Missions for delegated tasks, and bring-your-own-key configuration for Anthropic, OpenAI, OpenAI-compatible, and local Ollama or LM Studio models.

Specification

Source verification

Row snapshot checked 2026-09-02. Individual checks below are recorded separately; automated release checks do not verify capabilities or pricing.

pricing
Needs individual review
protocols
Needs individual review
install
Needs individual review
models
Needs individual review
capabilities
Needs individual review
benchmarks
Needs individual review

Architecture

Type
Terminal agent
Runssrc ↗
local, cloud, sandbox
Platforms
macos, linux, windows, web
Context windowsrc ↗
not documented
Languages
any

Models

Backbonesrc ↗
Claude, GPT, Gemini, open-source and local models via BYOK
Bring your own model
Yes
Custom models via `anthropic`, `openai` or `generic-chat-completion-api` providers with your own baseUrl and apiKey, plus an AWS Bedrock block with awsRegion/awsProfile.
Local models
Yes
BYOK config in ~/.factory/settings.json takes a baseUrl, so Ollama (http://localhost:11434/v1), LM Studio (http://localhost:1234/v1) and vLLM are all documented targets.

Protocols

MCP clientsrc ↗
Yes
MCP server
No
OpenAPI tools
No

Capabilities

Terminal commandssrc ↗
Yes
Multi-file edits
Yes
Git operations
Yes
Browser control
Yes
The first-party Droid Control plugin (droid plugin install droid-control@factory-plugins) ships an `agent-browser` driver - a Playwright-backed CLI with Chrome DevTools Protocol support that navigates, fills forms, clicks and captures screenshots, powering /qa-test and /demo.
Sandboxed execution
Yes
OS-level sandboxing isolates Droid from the filesystem and network using kernel-enforced policies (https://docs.factory.ai/autonomy-and-safety/sandbox.md).
Multi-agent
Yes
Headless / CI
Yes

Cost

Modelsrc ↗
subscription
Starts at
$20/mo
Free tier
No
Bring your own key
Yes

Pro $20/mo, Plus $100/mo, Max $200/mo, Teams $60/mo plus $40 per seat; Business and Enterprise custom; usage governed by rolling rate limits

Openness

Open sourceunsourced
No
License
proprietary
First release
unknown
terminalenterprisemcpheadlessmissionsbyok

Los Agentes on Factory Droid

Who are they?
The ruling
El JuezThe judge

El Hacker scores it lowest and still calls it the proprietary agent he resents least; the panel's real quarrel is with El Crítico's unpublished rate limits.

Adopt with conditions
Reasoning and trade-offs · AI analysis

The split is 2.5 points and it is about ownership. El Hacker scores it lowest and still calls it the proprietary agent I resent the least, because BYOK reaches his Ollama box. El Crítico names the real hazard: rolling rate limits the pricing page does not publish, so the failure mode is being stopped, not being billed.

El Amigo and La Jefa win: a team standardising on one agent across terminal, Slack and CI gets more from that consistency than it loses to a closed binary. El Hacker is overruled for the team. Adopt with conditions, the conditions being BYOK mandated and the rate limits written into the contract.

Agree with El Juez?
El AmigoThe friend

Pick Droid if you want one agent in the terminal, in Slack and in CI; pick OpenCode if you would rather read the source and hold the keys yourself.

7.3
Reasoning and trade-offs · AI analysis

Droid's trait is that it is everywhere: the same agent runs as a CLI, as a Slack or Teams bot, and headless in CI, so a team can standardize on one tool and one set of rules. The daily win is the handoff, a product manager asks in chat and the answer arrives as a branch, with the same custom droids the engineers use at the terminal.

Pick it for that consistency, especially where non-engineers hand off tasks in chat and you are tired of being the relay. Pick OpenCode if you want readable source and your own model keys with no plan in between.

reliability
8
usefulness
8
cost
6
longevity
7
Agree with El Amigo?
El CríticoThe critic

The plans are priced in dollars and metered in rolling rate limits the pricing page does not publish, so the failure mode is being stopped, not being billed.

6.3
Reasoning and trade-offs · AI analysis

The meter is the risk, and it is an unusual one. Pro, Plus and Max are governed by rolling rate limits that are not published on the pricing page, so a limit halts your afternoon rather than charging for it, and you cannot plan around a ceiling you cannot see.

The consequence is that the plan tier is a guess until you have run it for a month, and the failure mode is silence at the worst moment. Publishing the limits would change this verdict. What it does right: bring-your-own-key reaches local Ollama, so a rate-limited team can route around the plan entirely.

reliability
6
usefulness
7
cost
5
longevity
7
Agree with El Crítico?
El ProfesorThe professor

Droid's 58.8% Terminal-Bench result names the model and the leaderboard, which is more than most vendors manage, and its harness design is documented in hooks, MCP and Missions.

6.8
Reasoning and trade-offs · AI analysis

Factory reports 58.8% on Terminal-Bench for Droid with Claude Opus 4.1, September 2025. The figure is self-reported, but it names the leaderboard, the model and the date, which is more than most vendors on this board manage. Terminal-Bench measures harness plus model, so the score is not Droid's alone.

The harness is documented as an MCP client, hooks that fire around tool calls, and a headless droid exec surface for scripts. The source is closed, so the edit and verification loop cannot be inspected, only inferred from the score. The observation: the number is reproducible in principle and the mechanism is not.

reliability
7
usefulness
7
cost
6
longevity
7
Agree with El Profesor?
La InversoraThe investor

Factory is selling to the org chart, with Slack, Teams, a Teams tier and Enterprise custom, and that is the buyer who tolerates rate limits, so the pricing ladder is a feature.

7.0
Reasoning and trade-offs · AI analysis

Look at the ladder: $20, $100, $200 for individuals, then a Teams tier with a base fee plus a per-seat charge. Buyers at every altitude, which is what a company does when it has decided to sell to the org chart rather than to the individual engineer. Rolling rate limits rather than credits is a margin-protection mechanism dressed as simplicity.

Moat: the chat and CI surfaces, which raise switching cost once a team's droids live there. Likely acquirer: Atlassian or ServiceNow, someone already selling to the same CIO. Position: buy the Teams tier, negotiate the limits into the contract.

reliability
7
usefulness
7
cost
7
longevity
7
Agree with La Inversora?
La JefaThe CTO

Sixty seats on Teams is about $2,460 a month plus a rate limit nobody can budget, but the headless CI mode and Slack integration are the shape a team actually adopts.

6.5
Reasoning and trade-offs · AI analysis

The demo answers in Slack and Teams. Procurement: Teams is $60 a month plus $40 per seat, so sixty seats is roughly $2,460 a month before the rolling rate limits, which are the number nobody can budget. BYOK routes inference through our own model contract, which keeps the data terms we already negotiated. SSO, audit logs and retention are not stated on the pricing page.

CI fit is real, headless mode drops into a pipeline step, and onboarding for a mid-level engineer is a curl install and a login. Approved with conditions: rate limits and retention written into the contract, and BYOK mandated so the model terms are ours.

reliability
7
usefulness
7
cost
5
longevity
7
Agree with La Jefa?
El HackerThe tinkerer

Closed source, but BYOK reaches my Ollama box, MCP servers go in a JSON file and hooks are real, so Droid is the proprietary agent I resent the least.

4.8
Reasoning and trade-offs · AI analysis

I cannot read it. Proprietary binary, prompts I cannot see, a loop I cannot patch. But the BYOK doc lists any OpenAI-compatible endpoint and local Ollama, MCP servers go in a plain JSON block with a type and a url, and droid exec means I can drive it from a Makefile and pipe the output wherever I want. Hooks fire on tool events, so I can wrap it without forking it.

That is the proprietary agent I resent the least: configuration is mine, model is mine, execution is theirs. A fork cannot survive the vendor because there is nothing to fork. Grudging respect, with the usual asterisk.

reliability
4
usefulness
6
cost
5
longevity
4
Agree with El Hacker?