No. 023

Co-Scientist, Terminal-Bench Science, AI in peer review

Google claims its Co-Scientist ran full closed-loop autonomous research across three scientific domains. Terminal-Bench Science, released the same week, puts the best agent at a 30 percent pass rate on tasks drawn from real research workflows. Both can be true: narrow, well-defined problems are yielding to autonomous methods, while general scientific competence remains far off.

AI Agents as Researchers

  • Terminal-Bench Science: a benchmark for AI agents on real scientific workflows

    Steven Dillmann et al. (Stanford, Laude Institute), August 2026

    A community benchmark of 70 terminal-based research tasks across five scientific domains, each contributed by a working researcher from their own workflow, where the best agent (Claude Opus 5) passes about 30 percent.

  • Accelerating Scientific Research with Gemini in the Real-World

    Samuel Schmidgall, Tao Tu et al. (Google), arXiv, August 27 2026

    An updated multi-agent Gemini system extending the May Nature Co-Scientist paper to full closed-loop research cycles, including physically driving lab equipment, with a claimed autonomous discovery of an inference-time scaling architecture that outperformed six frontier models on HealthBench.

  • Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment

    Stephen Chung, Wenyu Du, William J. Wesley, arXiv, August 24 2026

    AI agents from different model families, with no central coordinator, produced new results on Kakeya sets and kissing configurations in dimension 11, outputting theorems and explanations rather than just numerical solutions.

Research Practice

  • The future of peer review requires AI support, not AI bans

    Nature (Xuegong Zhang), August 25 2026

    Zhang argues that blanket AI bans in peer review address real concerns imperfectly and are impossible to monitor, calling instead for structured AI support that preserves reviewer accountability.

  • AI Use in The Scholarly Kitchen

    The Scholarly Kitchen (David Crotty), August 27 2026

    The Scholarly Kitchen adopted a formal AI disclosure policy for its own contributors, requiring disclosure of how, when, and which model was used, modeling the transparency standards it has long advocated for across scholarly publishing.