Question 1

Which is better: ClawBench or Talkie?

Accepted Answer

Based on our expert panel, ClawBench has a stronger verdict with a 75% Ship rate. ClawBench received a panel verdict of Ship and Talkie received Ship.

Question 2

Is ClawBench free?

Accepted Answer

ClawBench pricing: Free / Research

Question 3

Is Talkie free?

Accepted Answer

Talkie pricing: Free / Open Research

Question 4

What do experts say about ClawBench vs Talkie?

Accepted Answer

ClawBench: ClawBench is a browser agent evaluation framework built around 153 real-world tasks running on 144 live production websites — not simulated environments or curated sandboxes. Tasks span e-commerce, travel booking, SaaS dashboards, government portals, and developer tools. A built-in request interceptor blocks genuinely irreversible actions (payments, form submissions that send data) so evaluations can run safely on real sites.

The benchmark records five layers of data per run: session replays, screenshots at each decision point, raw HTTP traffic, agent reasoning traces, and browser action sequences. This makes failure analysis tractable — you can see exactly which DOM element the agent misidentified, not just a final score. The dataset is open and the evaluation harness is reproducible.

The headline finding is sobering: Claude Sonnet 4.6, the best performer, completes only 33.3% of tasks. GLM-5 is second at 24.2%. No model exceeds 50% on any individual task category. The implication is stark — current browser agents are far from autonomous on the open web, and the gap between benchmark performance and production performance is still enormous. Talkie: Talkie is a 13-billion parameter language model trained exclusively on English-language texts published before 1931 — the largest vintage language model built to date. Created by researchers Nick Levine, David Duvenaud (University of Toronto), and Alec Radford (of GPT and DALL-E fame), it represents a novel approach to understanding what training data really does to a model.

The research insight is elegant: modern LLMs are so thoroughly contaminated by modern internet data (directly or through distillation) that it's nearly impossible to isolate what the model "knows" from what it absorbed during training. Talkie solves this by hard-cutting the training corpus at 1931 — predating digital computers entirely. This lets the team run controlled experiments impossible with contemporary models, such as teaching the model to write Python from examples alone and measuring how quickly it generalizes.

Talkie was trained on ~260 billion tokens of historical text and fine-tuned using direct preference optimization with Claude as judge on structured historical documents (etiquette manuals, letter-writing guides). It's openly available on Hugging Face for research use. It also happens to produce wonderfully formal, slightly anachronistic prose.

ClawBench vs Talkie

ClawBench

Talkie

Bookmarks