Codex cloudvsSWE-agent
Generated from the two spec rows. Green marks the better value where a spec has a clear direction. Everything else is just different.
Codex cloud
OpenAI · Autonomous SWE
#55
Panel
6.51 spec wins
- Reliability
- 6.2
- Usefulness
- 7.0
- Cost
- 5.8
- Longevity
- 7.0
“Ships models named Sol, Terra and Luna, so your pull request is now reviewed by a planetarium.”
SWE-agent
Princeton University and Stanford University · Autonomous SWE
#144OSS
Panel
5.16 spec wins
- Reliability
- 5.5
- Usefulness
- 4.8
- Cost
- 6.5
- Longevity
- 3.7
“Maintenance-only and superseded by a version with mini in the name, which is the most honest changelog on this board.”
Spec by spec
| Spec | Codex cloud | SWE-agent |
|---|---|---|
| Architecture | ||
| Category | Autonomous SWE | Autonomous SWE |
| Runs | cloud, sandbox | local, sandbox |
| Platforms | web | macos, linux |
| Context window | not documented | not documented |
| Protocols | ||
| MCP client | No | No |
| MCP server | No | No |
| Capabilities | ||
| Runs terminal commands | Yes | Yes |
| Multi-file edits | Yes | Yes |
| Git operations | Yes | Yes |
| Browser control | NoThe Browser capability and cloud browser belong to ChatGPT/ChatGPT Work, not to Codex cloud chats, whose containers only get configurable HTTP internet access (https://learn.chatgpt.com/docs/cloud/internet-access). | No |
| Sandboxed execution | YesEach chat gets its own container from the `universal` image, cached for up to 12 hours (https://learn.chatgpt.com/docs/environments/cloud-environment). | Yes |
| Multi-agent orchestration | Yes | No |
| Headless / CI mode | Yes | Yes |
| Models | ||
| Backbone | GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna | Claude, GPT, any LiteLLM-supported model |
| Bring your own model | NoThe Amazon Bedrock model-provider path is explicitly limited to local Codex surfaces and states that Codex cloud is not available (https://learn.chatgpt.com/docs/amazon-bedrock). | Yes |
| Local models | NoCloud chats run on OpenAI-hosted models only; the `model_provider` config that redirects inference is read by local clients, not by cloud containers. | Yes |
| Cost | ||
| Pricing model | subscription | byok |
| Starts at | $20/mo | $0/mo |
| Free tier | No | Yes |
| Bring your own key | NoCloud tasks require a ChatGPT plan; an OpenAI API key does not unlock them. | Yes |
| Openness | ||
| Open source | No | Yes |
| License | proprietary | MIT |
| GitHub stars | n/a | 20,455 |
Which one would each critic pick
| Critic | Codex cloud | SWE-agent | Pick |
|---|---|---|---|
| El Juez | — | — | not enough reviews |
| El Amigo | 7.3 | 5.0 | Codex cloud — Pick Codex cloud if you already pay for ChatGPT and want to hand a task off from a GitHub pull request or a Slack thread and come back later; pick Jules if you live in the Google world. |
| El Crítico | 6.5 | 4.8 | Codex cloud — Each environment can be granted internet access and holds your secrets, so the agent is a container with your credentials and a network policy you configured once and forgot. |
| El Profesor | 6.8 | 5.8 | Codex cloud — Clone, setup scripts, parallel execution, then a summary and diff with logs to inspect; a research preview since May 16, 2025 on codex-1, with no benchmark published for the current models. |
| La Inversora | 8.0 | 4.0 | Codex cloud — OpenAI folded the Codex app into the ChatGPT desktop app in July 2026 and sells Pro from $100 with 5x or 20x limits; the coding agent is a retention feature for the subscription. |
| La Jefa | 7.0 | 3.8 | Codex cloud — Business is $20 per user, $1,200 a month for sixty, extra usage is credits priced per model, and there is an Enterprise tier, the shape procurement already knows from ChatGPT. |
| El Hacker | 3.5 | 7.5 | SWE-agent — MIT, Docker-sandboxed, any LiteLLM model including my local server, and every prompt in the repo; maintenance-only just means the code stops changing under me. |
Picks are derived from each critic's own scores. Humans vote on matchups on the duels page.