No. 019

AlphaFold team disbanded, verification as the bottleneck, Tao's ICM caution

This week, verification crystallized as the central constraint for AI in science. Google Research shipped a prototype chaining every AI-generated claim to its evidence. A NeurIPS workshop adopted "the bottleneck is verification" as its thesis. A benchmark for comparing human and AI manuscript reviews landed on bioRxiv. And at the International Congress of Mathematicians, Terence Tao framed AI as a foundational crisis for the discipline, comparable to Russell and Godel. Meanwhile, Google disbanded the AlphaFold team. The juxtaposition says something: the industry is building verification infrastructure for AI-generated science while dismantling the team behind AI-discovered science.

Institutional Shifts

  • Google DeepMind disbands its Nobel-prize-winning AlphaFold team

    Engadget (Mariella Moon, reporting Financial Times), July 30 2026

    John Jumper left for Anthropic, remaining staff were reassigned to Gemini or Isomorphic Labs, and VP of Research Kohli said the strategy "has evolved" from grand challenges to concrete goals -- a sharp data point on whether big labs treat AI-for-science as a durable mission or a milestone to harvest.

  • arXiv welcomes inaugural CEO and Board of Directors

    arXiv blog, July 30 2026

    Penelope Lewis (formerly Chief Publishing Officer at AIP Publishing) named inaugural CEO of the newly independent nonprofit, effective August 17; the eight-person board includes Retraction Watch co-founder Ivan Oransky, signaling that integrity is a governance-level priority for the server that hosts the highest concentration of AI-generated preprints.

  • ChatGPT for Academic Researchers

    OpenAI, July 29 2026

    Free access to frontier models for 10,000 researchers this summer, scaling to 100,000 by 2027, with data excluded from training by default -- a frontier lab productizing directly for the academic market alongside the AI-scientist tools (Biomni, Co-Scientist, Robin) tracked in recent editions.

Verification

  • Science One Framework: A verifiable autonomous research framework via Chain-of-Evidence

    Google Research (Rui Meng and Tomas Pfister), July 30 2026

    An experimental prototype that chains every AI-generated claim to a recorded evidence trail: a Problem Investigator grounds references against up to 100 full-text PDFs via Semantic Scholar, a Discovery Engine logs parallel exploration branches, and a Claim Verifier tags each factual assertion with its source, eliminating hallucinated references entirely against baselines that fabricated up to 21 percent.

  • Verification in the Age of AI Scientists (NeurIPS 2026 workshop)

    AI4Science Community, NeurIPS 2026 (Sydney, December 2026)

    The AI4Science community's NeurIPS workshop is organized around a single thesis: "the bottleneck for AI for Science is no longer hypothesis generation, it is verification," with organizers from Harvard, MIT, Microsoft Research, and Google DeepMind, and submissions due August 29.

  • Benjamin Golub on the OpenAI-president math-slop incident and Refine.ink

    LinkedIn, July 2026

    After OpenAI's president shared an AI-generated math paper that a mathematician dismantled within a day (author conceded, tweet deleted), Golub argues heroes catching errors one at a time cannot scale against the volume arriving -- and positions Refine.ink as the automated independent verification layer the system now requires.

  • ReviewBench: An Extensible Framework for Benchmarking Human and AI Manuscript Review

    bioRxiv, April 2026

    First shared benchmark for comparing human and AI manuscript reviews, spanning three corpora across computer science (ICLR 2025, n=1,000), social science (Nature Human Behaviour, n=142), and life science (eLife, n=1,000) -- the empirical infrastructure the field's AI-in-peer-review claims have lacked.

  • Reviewer3 (PyPI package)

    PyPI (Robert Haase), July 28 2026

    Now a pip-installable command-line tool (v0.1.2, BSD-3) that runs LLM-based manuscript review locally via Ollama by default, moving from web service to scriptable, privacy-first infrastructure any lab can embed in its workflow.

What AI Cannot Do

  • LLMs Can't Jump: Why AI Masters the Proof but Misses the Premise

    Tom Zahavy (Google DeepMind), PhilSci Archive January 2026, presented at ICML 2026

    Argues LLMs handle deduction and induction but are structurally incapable of abduction -- the creative leap to a new explanatory hypothesis -- using Einstein's path to General Relativity as the case study. Zahavy is careful that this is his own position, not DeepMind's, and not a claim that AI can never discover. The narrower point: current models cannot invent new foundational axioms when data is scarce.

  • Mathematics in the age of AI (ICM 2026 public lecture)

    Terence Tao, International Congress of Mathematicians, July 24 2026

    A Fields medalist frames AI as a coming crisis in the foundations of mathematical values, analogous to the 1900-1930 foundations crisis, and reports that in the controlled First Proof challenge, 7 of 10 novel research-level problems were solved at publication quality by AI at $10 to $1,000 of compute per problem -- the question is not capability but what it means for the discipline's practices.