Question 1

Which is better: Darwin-4B-David or DeepSeek V4-Pro?

Accepted Answer

Based on our expert panel, Darwin-4B-David has a stronger verdict with a 75% Ship rate. Darwin-4B-David received a panel verdict of Ship and DeepSeek V4-Pro received Ship.

Question 2

Is Darwin-4B-David free?

Accepted Answer

Darwin-4B-David pricing: Open Source

Question 3

Is DeepSeek V4-Pro free?

Accepted Answer

DeepSeek V4-Pro pricing: Open Source (Apache 2.0) / ~$0.30/MTok API

Question 4

What do experts say about Darwin-4B-David vs DeepSeek V4-Pro?

Accepted Answer

Darwin-4B-David: Darwin-4B-David is a 4.5-billion-parameter model that achieves 85.0% on GPQA Diamond — outperforming Google's Gemma-4-31B (84.3%) at roughly 1/7th the parameter count. The kicker: it required no training whatsoever. It was built in 45 minutes on a single H100 using MRI-guided DARE-TIES model merging, a novel variant of the merge-and-trim technique.

The MRI-guided approach uses activation analysis to identify which parameters in each source model are most critical, then applies DARE-TIES merging only to the high-value weight regions. This avoids the catastrophic interference that usually degrades merged models. The result is a small model that inherits the strengths of multiple larger predecessors without any of the compute cost of fine-tuning.

For the AI community, this is a meaningful data point: model merging continues to close the gap with expensive training runs. Darwin-4B-David demonstrates that thoughtful merge strategies can extract benchmark-level performance from models that are a fraction of the size, making capable AI more accessible on consumer hardware. DeepSeek V4-Pro: DeepSeek just dropped V4-Pro and V4-Flash simultaneously — and it's a statement release. V4-Pro packs 1.6 trillion total parameters in a MoE architecture with only 49B active per token, a 1-million-token context window, and a hybrid attention system (Compressed Sparse Attention + Heavily Compressed Attention) that requires just 27% of single-token inference FLOPs compared to V3.2. Both models are Apache 2.0.

The hardware story is arguably the bigger news: V4 was trained entirely on Huawei Ascend 950PR chips, zero NVIDIA. That's a geopolitical and technical milestone — it validates China's domestic AI compute stack at frontier scale. The Engram Memory System gives V4 conditional context recall (94% at 128K tokens vs ~45% for V3.2), enabling genuinely long-context reasoning.

V4-Flash at 284B parameters (13B active) is the cheaper, faster sibling for production use. Pricing is expected around $0.30/M tokens for Pro. The timing — released to HN today with 99+ points within hours — confirms this as an immediate conversation in the developer community about whether open-weight frontier models have finally matched proprietary ones.

Darwin-4B-David vs DeepSeek V4-Pro

Darwin-4B-David

DeepSeek V4-Pro

Bookmarks