AI Models Face Scrutiny Amidst Performance Claims and Distillation Allegations

The artificial intelligence landscape is buzzing with new model releases and performance benchmarks, alongside serious allegations of data distillation. Cognition's SWE-2 model has reportedly achieved a high score on the Terminal-Bench 2.1, with claims of rivaling established models like Fable 5.1 and GPT-Astra. Meanwhile, OpenAI's recent release of Navier-Stokes has drawn attention for its inclusion of a formal proof in Lean 4, sparking discussions about the mathematical underpinnings of AI advancements.

However, the rapid progress is shadowed by concerns over data integrity. Anthropic has detailed alleged "distillation campaigns" by Chinese AI companies, including DeepSeek, Alibaba, and Moonshot AI, suggesting a pattern of extracting knowledge from existing models. This comes as DeepSeek itself has seen new versions, including "DeepSeek v4.1 Flash Uncensored," enter the discussion. The intensity of these alleged attacks has reportedly escalated with the growing competition in the AI sector, raising questions about ethical development and intellectual property.

11 stories · 4 sources

#deepseek #ai #benchmarks

Other digests