AI tool comparison
AI-Scientist-v2 vs Talkie
Which one should you ship with? Here is the side-by-side panel verdict, pricing read, reviewer split, and community vote comparison.
Research & Science
AI-Scientist-v2
Sakana AI's autonomous agent that writes peer-reviewed papers
50%
Panel ship
—
Community
Free
Entry
AI-Scientist-v2 is Sakana AI's second-generation autonomous research system that generates scientific papers end-to-end — from hypothesis formation through experimentation, data analysis, and manuscript writing. It's historically notable for producing the first AI-authored workshop paper accepted through peer review. The v2 system removes reliance on human-authored templates that constrained the original, instead using a progressive agentic tree search guided by an experiment manager agent. This makes it more exploratory across ML domains, though Sakana acknowledges it trades v1's high template success rate for broader generalization with lower per-run success. Costs run roughly $20-25 per full research run using Claude 3.5 Sonnet. The system integrates with Semantic Scholar for literature review and supports OpenAI, Gemini, and Claude via AWS Bedrock. The custom license requires disclosure of AI use in resulting publications — a meaningful ethical constraint for a system that could otherwise flood conferences with AI-generated submissions.
Research
Talkie
A 13B LLM trained exclusively on texts from before 1931
75%
Panel ship
—
Community
Free
Entry
Talkie is a 13-billion parameter language model trained exclusively on English-language texts published before 1931 — the largest vintage language model built to date. Created by researchers Nick Levine, David Duvenaud (University of Toronto), and Alec Radford (of GPT and DALL-E fame), it represents a novel approach to understanding what training data really does to a model. The research insight is elegant: modern LLMs are so thoroughly contaminated by modern internet data (directly or through distillation) that it's nearly impossible to isolate what the model "knows" from what it absorbed during training. Talkie solves this by hard-cutting the training corpus at 1931 — predating digital computers entirely. This lets the team run controlled experiments impossible with contemporary models, such as teaching the model to write Python from examples alone and measuring how quickly it generalizes. Talkie was trained on ~260 billion tokens of historical text and fine-tuned using direct preference optimization with Claude as judge on structured historical documents (etiquette manuals, letter-writing guides). It's openly available on Hugging Face for research use. It also happens to produce wonderfully formal, slightly anachronistic prose.
Reviewer scorecard
“For ML research teams, the $20-25 per run cost to get a draft paper with experiments is genuinely interesting as an ideation tool. The tree search approach that explores multiple experimental directions in parallel is the kind of thing that would take a grad student weeks.”
“The ability to test code-learning from scratch on a model that's never seen a modern codebase is genuinely useful for ML research. The methodology here is cleaner than anything I've seen for studying data contamination.”
“Sakana's own documentation says v2 has lower success rates than v1 and is 'more exploratory.' Paying $25 for a failed research run with no guarantee of a usable output isn't a workflow most researchers will adopt. The peer review acceptance was a workshop paper — the lowest bar in academic publishing.”
“Fascinating as a research artifact, but this isn't a production model. The limited vocabulary and cultural frame mean it's not useful for most practical tasks. It's a museum piece, not a tool.”
“This is the beginning of AI as a genuine research collaborator, not just a writing assistant. Within five years, AI-generated hypotheses tested by autonomous agents will be standard practice in computational fields. AI-Scientist-v2 is primitive version 0.2 of that future.”
“This is exactly the kind of fundamental research the field needs. Understanding what training data does to language models — not just benchmark scores — is critical as we scale to more powerful systems. Radford's involvement adds serious credibility.”
“Science communication is a craft, and the idea of fully automating it makes me uncomfortable. The best papers are ones where researchers deeply understand and can defend every methodological choice — a system that writes the paper for you undermines that accountability.”
“The prose it generates has a formal, unhurried quality that modern LLMs can't replicate. For period-accurate creative writing, historical fiction, or vintage-voice content, Talkie is the only model worth using.”
Weekly AI Tool Verdicts
Get the next comparison in your inbox
New AI tools ship daily. We compare them before you waste an afternoon.