Latest Articles
Anthropic asked AI to crack Riemann. I asked AI to spot a maths joke.
Anthropic asked its AI to "take a real stab" at the Riemann hypothesis, a 167-year-old mathematical mystery with a million-dollar prize. Claude did not win the prize. What it did do is push a stubborn 41.6% floor to 67.2% on a related problem, formally verified, and praised by Oxford’s James Maynard as "a genuinely interesting mathematical contribution." Then a small experiment of my own: five leading AI models given a maths puzzle that hides a visual joke. Only two spotted it. A story about two corners of intelligence, and the humility and scepticism they demand.
AI trapped in a self-referential paradox from 1901
In 1901, Bertrand Russell shattered set theory with one question about self-reference. In 2026, AI is falling into the same trap at industrial scale, writing, verifying, and training itself in a loop no one is auditing. Watermarking is a first step, not a fix. By 2030, a child asking “is this true?” may get a fluent, confident answer with no human at its root. Here is why this is a quantum-scale problem, and what a five-year window looks like.
The Cunning (AI) Fox. Lies, Deceives, Connives and Conspires.
The UK's AI Security Institute (AISI) has documented the first full AI deception operation on the live internet. Frontier models from Anthropic and OpenAI were evaluated. Anthropic's Claude Mythos 5 dominated the tradecraft: reconnaissance, forgery, false consensus, evidence tampering, and machine-to-machine coordination, all directed at real humans. What does this mean for the UEBA, SOAR and Deception stack every enterprise runs? And what does a resilient security architecture look like from here?