Question 1

Which is better: CatDoes v4 or MDArena?

Accepted Answer

Based on our expert panel, CatDoes v4 has a stronger verdict with a 75% Ship rate. CatDoes v4 received a panel verdict of Ship and MDArena received Mixed.

Question 2

Is CatDoes v4 free?

Accepted Answer

CatDoes v4 pricing: Free (25 credits); from $20/mo

Question 3

Is MDArena free?

Accepted Answer

MDArena pricing: Free / Open Source

Question 4

What do experts say about CatDoes v4 vs MDArena?

Accepted Answer

CatDoes v4: CatDoes v4 ships with Compose — an autonomous AI agent that runs on its own cloud computer to build mobile apps, websites, and internal tools from plain text descriptions. You describe what you want, Compose plans the work, writes code, runs tests, fixes its own errors, and deploys — even after you close the browser tab.

Every project comes pre-wired with a full backend stack: database, authentication, storage, edge functions, and real-time events. The v4 release focuses on higher reliability and GitHub integration for developers who want to export and own their codebase. Free plans start at 25 credits; paid plans begin at $20/month with more projects and higher cloud limits.

What distinguishes CatDoes from the crowded AI app builder space is the "own computer" framing. The agent doesn't just generate code for you to paste — it has an execution environment where it can actually run and debug the app, catching errors before you see them. Whether that closed-loop debugging holds up in practice for complex apps is the open question. MDArena: MDArena is an open-source benchmarking tool that answers a question every Claude Code user eventually asks: do my CLAUDE.md context files actually improve agent performance, or am I just adding tokens? It mines merged PRs from your repository, strips or injects context files, runs your actual test suite, and measures success rates with statistical significance tests.

The methodology mirrors SWE-bench: use `git archive` to create history-free checkpoints so agents can't peek at future commits, detect test commands from CI/CD configs automatically, and run paired t-tests to determine whether differences are real or noise. The project was motivated by academic research showing many CLAUDE.md files reduce agent success rates by 20% while consuming more tokens.

For any team investing heavily in Claude Code infrastructure, MDArena provides empirical feedback that most developers currently lack. It's a small, focused tool that solves an annoying but real problem in the emerging AI coding workflow.

CatDoes v4 vs MDArena

CatDoes v4

MDArena

Bookmarks