No. 020

Astra solves ten open math problems, AI agents fail at open-ended research, peer review under strain

Three editions running, AI-in-math results keep landing. This week: OpenAI claims ten open problems solved for under $2,000 of compute each, two independent teams used GPT-5.6 Sol Ultra to crack the same quantum cryptography problem hours apart, and a Fields medalist left his university post for OpenAI's safety team. Meanwhile, a Princeton-led study gave frontier agents $3,000 and six days to answer open research questions; the agents completed all the engineering and produced nothing publishable. The gap between "solve a known open problem when pointed at it" and "identify and pursue a novel research question" is where the discipline's actual uncertainty lives.

AI in Mathematics

  • Ten advances in mathematics and theoretical computer science

    OpenAI, August 1 2026

    Ten results on long-standing open problems, generated by an internal model OpenAI calls Astra, with humans writing up manuscripts and formalizing them in Lean -- areas span high-dimensional geometry, coding theory, group theory, operator algebras, circuit complexity, quantum complexity, lattice cryptography, and extremal combinatorics, including the first explicit construction of a non-sofic group (open since Gromov, 1999) and a superexponential lower bound for multicolor triangle Ramsey numbers (Erdos problem 183); OpenAI claims under $2,000 of compute per result, though formal peer review is still pending on several.

  • AI helped produce two proofs for the same cryptography problem

    Scientific American (via Yahoo News), July 31 2026

    Two independent teams -- Seyoon Ragavan at MIT and Ananth/Sahai at UCSB/UCLA -- used GPT-5.6 Sol Ultra to produce proofs addressing unclonable encryption in quantum cryptography, posting to arXiv three hours apart; neither paper was peer reviewed, and the culture shift is already felt: "Now the general mentality is: if someone mentions an open problem, the first thing is to see if GPT solves it."

  • Fields Medalist Jacob Tsimerman takes leave from Toronto to join OpenAI's AI safety team

    Wall Street Journal, August 2026

    Tsimerman, a 2026 Fields medalist who proved the Andre-Oort conjecture, announced the move at the ICM in Philadelphia on July 23, keeping his Toronto faculty post but telling the press he believes AI will surpass human mathematicians within two years -- a sharp data point on the labs-versus-academia pull, echoing the AlphaFold team departures tracked in edition 019.

Research Agents

  • Can AI agents conduct open-ended AI research? Early evidence from two case studies

    arXiv (Kirgis, Kapoor, Narayanan et al.), July 29 2026

    Shadow evaluations gave Claude Opus 4.8 on OpenClaw (with a GPT-5.6 Sol robustness check) six days, $3,000, and full VM access to tackle unpublished NeurIPS 2026 research questions; the agents ran hundreds of experiments, compiled camera-ready LaTeX, and completed all engineering autonomously, but both papers were rejected (2/6 and 1/6) -- the agents could not make substantial progress on the actual research questions.

  • Introducing Elicit Research Agent

    Elicit, August 4 2026

    A new AI environment for high-stakes evidence synthesis across scientific and public data sources, with sentence-level citations and a proprietary harness trained to avoid reasoning mistakes in decision-heavy contexts; on Elicit's new BioDecisionBench (40 pharmaceutical development scenarios), the Research Agent identified more key decision considerations than Claude Opus 5 Max or GPT-5.6 Sol Max, though no tool came close to identifying all of them.

Peer Review and Publishing

  • Do we need (r)evolution in peer review?

    The Scholarly Kitchen (Dmitry Kochetkov), August 7 2026

    Journals on ScholarOne received 33% more submissions in Q1 2026 than the same period in 2025, AI-assisted writing drove a 42% increase in manuscript volume, peer review reports themselves now show signs of AI drafting, and 10% of reviewers still do roughly 50% of the work -- Kochetkov argues for a Publish-Review-Curate model that decouples dissemination from gatekeeping and introduces standardized trust markers at the publication level.

  • The Published Voice: Whose Voice Are We Really Reading?

    The Scholarly Kitchen (Roohi Ghosh), August 5 2026

    Asks what "author voice" means in academic writing now that AI routinely drafts and edits manuscripts, arguing the concept was always a collective construct shaped by co-authors, reviewers, and disciplinary convention -- the practical concern is that editors may mistake competent non-native-English writing for AI generation, introducing a new form of bias.