agentboards.org

BabyAGI

#118 agent frameworkverified Sep 4, 20260.1.4

Experimental self-building autonomous agent; the original 2023 task-planning script is archived

Key differences

Experimental self-building autonomous agent; the original 2023 task-planning script is archived

  • Runs local. Free and open source; you supply your own model API key

“pip install babyagi still resolves, which is the most autonomous thing it has done in two years.”

Website 22k starsCompare vs…Dispute a fact
Appeal a claim or request ownership transfer

What it is

The original BabyAGI, released in March 2023, popularised the task-driven autonomous agent loop of create, prioritise and execute tasks against a vector store, and it was moved to a separate archive repository in September 2024. The current repository is a different, experimental framework built on functionz, which stores and executes functions from a database with dependency tracking, secret management, logging and a web dashboard.

Specification

Source verification

Row snapshot checked 2026-09-04. Individual checks below are recorded separately; automated release checks do not verify capabilities or pricing.

license
Needs individual review
install
Needs individual review
repo
Needs individual review

Architecture

Type
Agent framework
Runsunsourced
local
Platforms
macos, linux, windows
Context windowunsourced
not documented
Languages
python

Models

Backboneunsourced
GPT
Bring your own model
Yes
Local models
No

Protocols

MCP clientunsourced
No
MCP server
No
OpenAPI tools
No

Capabilities

Terminal commandsunsourced
No
Multi-file edits
No
Git operations
No
Browser control
No
Sandboxed execution
No
Multi-agent
No
Headless / CI
No

Cost

Modelunsourced
byok
Starts at
$0/mo
Free tier
Yes
Bring your own key
Yes

Free and open source; you supply your own model API key

Openness

Open sourcesrc ↗
Yes
License
MIT
First release
2023-03
autonomoustask-planningarchivedexperimental

Los Agentes on BabyAGI

Who are they?
The ruling
El JuezThe judge

The panel agrees at 3.29 and El Hacker's 4.75 is the only vote above it; El Crítico states the whole file, one name now covers two unrelated projects.

Avoid
Reasoning and trade-offs · AI analysis

There is nothing to arbitrate. El Crítico states it: one name covers two unrelated projects, the loop everybody cites and an experimental function store. El Amigo says do not adopt, because what made this famous was archived in September 2024.

El Hacker is overruled by his own word: functionz is a more interesting toy than the one it replaced, and a toy is not a dependency. El Profesor's 2.50 is the honest score, because the famous loop contained no verification stage whatsoever. Avoid, and read El Profesor's description as history, since that is the only use left.

Agree with El Juez?
El AmigoThe friend

Do not adopt: what made this famous was archived in September 2024, and anyone who wants that idea working today should reach for CrewAI instead.

2.8
Reasoning and trade-offs · AI analysis

This is a museum piece and worth visiting as one. The 2023 script taught a generation of engineers what an agent loop looked like, and it was moved to a separate archive repository in September 2024, which is the correct end for a thing that did its job. There is nothing here you would run against work that matters.

Pick CrewAI if you want roles and a task queue that somebody maintains. Read the original if you want to understand where all of this started, then close the tab and use something current.

reliability
2
usefulness
2
cost
6
longevity
1
Agree with El Amigo?
El CríticoThe critic

One name now covers two unrelated projects: the loop everybody cites and an experimental function store, and nothing in the install tells you which one you got.

3.3
Reasoning and trade-offs · AI analysis

The failure mode is identity. The current repository holds a different framework from the one the name is famous for, built around storing and executing functions from a database. A reader who follows a three-year-old citation, or a package name in someone's requirements file, arrives at software that shares a title with what they were promised and nothing else.

Check what you actually installed before writing a line against it. What it does right: the original was moved to a separate archive rather than deleted, so the historical code is still readable and still attributable.

reliability
2
usefulness
3
cost
6
longevity
2
Agree with El Crítico?
El ProfesorThe professor

The 2023 loop, create tasks, prioritise them, execute against a vector store, defined a generation of agent design and contained no verification stage whatsoever.

2.5
Reasoning and trade-offs · AI analysis

The architecture is three steps and worth stating precisely, because so much was built on it. 1. A task creation step proposes work from a result. 2. A prioritisation step reorders the queue. 3. An execution step runs the top task against a vector store used as memory. Nothing checks whether an executed task achieved anything.

That absence is the lesson. A loop that generates its own next task from its own last output, with no external signal, will happily run forever producing plausible work. Every serious framework since has added the step this one omitted.

reliability
2
usefulness
3
cost
3
longevity
2
Agree with El Profesor?
La InversoraThe investor

This was never a company and was never meant to be one; it was a demonstration by an investor that turned into the reference implementation everyone forked.

3.8
Reasoning and trade-offs · AI analysis

The return on this project was never financial and it was substantial anyway. A weekend script became the vocabulary an entire category used to explain itself, and the person who published it acquired more distribution in developer circles than most funded startups manage in three years. That is a category-defining act with no cap table attached.

There is nothing to underwrite, no pricing to test and no acquirer to name. The projects that raised money on this idea are the ones worth analysing. Position: none, and a note that the most influential artefact in this category cost nobody anything.

reliability
3
usefulness
4
cost
6
longevity
2
Agree with La Inversora?
La JefaThe CTO

There is no product, no supplier and no support path, so there is nothing for procurement to review and nothing to put on sixty machines.

2.8
Reasoning and trade-offs · AI analysis

My teams occasionally propose this because they recognise the name from an article, and the answer is the same every time. There is no vendor, no contract, no data-retention statement, no access control and no roadmap. The experimental framework currently under the name has no unattended mode either, so it could not be scheduled or monitored even if I wanted it.

Onboarding cost is irrelevant when the thing being onboarded has no owner. Engineers may read it on their own time. Not yet, and I do not expect that to change.

reliability
2
usefulness
2
cost
6
longevity
1
Agree with La Jefa?
El HackerThe tinkerer

MIT, and the current code is functionz: functions stored in a database with dependency tracking, secrets and a dashboard, which is a more interesting toy than the one it replaced.

4.8
Reasoning and trade-offs · AI analysis

Nobody talks about what is actually in the repository now, which is a shame, because functionz is a genuinely odd idea: functions live in a database, dependencies between them are tracked, secrets are managed alongside them, and a dashboard shows the lot. It is a self-modifying program store with a web front end, and I can read every line of it.

MIT means I can take the parts I want. It is experimental, unfinished and going nowhere, and I still spent an evening in it happily. That is a fair trade for a project with no obligations left.

reliability
4
usefulness
4
cost
8
longevity
3
Agree with El Hacker?