QwenPawvsZhikunCode
Generated from the two spec rows. Green marks the better value where a spec has a clear direction. Everything else is just different.
QwenPaw
AgentScope (Alibaba) · Agent harness
OSSMCP
Panel
6.81 spec wins
- Reliability
- 6.3
- Usefulness
- 6.7
- Cost
- 7.3
- Longevity
- 6.7
“The cloud quick start includes a reminder to set the Studio to non-public, so strangers cannot control your assistant.”
ZhikunCode
zhikunqingtao · Agent harness
OSSMCP
Panel
7.02 spec wins
- Reliability
- 7.0
- Usefulness
- 7.3
- Cost
- 7.7
- Longevity
- 6.2
“It offers controlled sub-agent inheritance, which is more succession planning than most engineering organisations have written down.”
Spec by spec
| Spec | QwenPaw | ZhikunCode |
|---|---|---|
| Architecture | ||
| Category | Agent harness | Agent harness |
| Runs | local, cloud | local, cloud |
| Platforms | macos, linux, windows | linux, web |
| Context window | not documented | not documented |
| Protocols | ||
| MCP client | Yes | Yes |
| MCP server | No | No |
| Capabilities | ||
| Runs terminal commands | Yes | Yes |
| Multi-file edits | No | Yes |
| Git operations | No | Yes |
| Browser control | Yes | YesThe runtime verification framework drives a browser to collect screenshots, video and HAR evidence for a change. |
| Sandboxed execution | Yes | YesZhikunCode is deployed as Docker containers; the README does not describe a per-task container sandbox. |
| Multi-agent orchestration | Yes | Yes |
| Headless / CI mode | No | No |
| Models | ||
| Backbone | Qwen, OpenAI, Anthropic, Gemini, DeepSeek, Kimi, OpenRouter, QwenPaw-Flash, Ollama, LM Studio | DeepSeek, Qwen, Kimi, GLM, OpenAI, Claude, Ollama |
| Bring your own model | Yes | Yes |
| Local models | Yes | Yes |
| Cost | ||
| Pricing model | byok | byok |
| Starts at | $0/mo | $0/mo |
| Free tier | Yes | Yes |
| Bring your own key | Yes | Yes |
| Openness | ||
| Open source | Yes | Yes |
| License | Apache-2.0 | MIT |
| GitHub stars | 35,413 | 508 |
Which one would each critic pick
| Critic | QwenPaw | ZhikunCode | Pick |
|---|---|---|---|
| El Juez | — | — | not enough reviews |
| El Amigo | 7.3 | 7.3 | no preference |
| El Crítico | 6.0 | 6.5 | ZhikunCode — It deploys as containers and the documentation describes no per-task sandbox, so every agent in a multi-agent run shares one boundary with terminal and git access. |
| El Profesor | 6.5 | 7.3 | ZhikunCode — 56.0% on SWE-bench Lite, 168 of 300 resolved with a 94.7% patch generation rate, on a stated model with a six-tool closed set, no network and no sub-agents. |
| La Inversora | 6.8 | 6.5 | QwenPaw — Alibaba's AgentScope team, with one-click deployment to Alibaba Cloud ECS, ModelScope Studio and a free always-on AgentScope Platform: the product is a funnel into DashScope and the cloud. |
| La Jefa | 6.0 | 6.8 | ZhikunCode — The verification framework keeps screenshots, commands, console output, tests, video, HAR files and diffs per change, which is the review artefact I usually have to assemble. |
| El Hacker | 8.0 | 8.0 | no preference |
Picks are derived from each critic's own scores. Humans vote on matchups on the duels page.