No. 023
Co-Scientist, Terminal-Bench Science, AI in peer review
Google claims its Co-Scientist ran full closed-loop autonomous research across three scientific domains. Terminal-Bench Science, released the same week, puts the best agent at a 30 percent pass rate on tasks drawn from real research workflows. Both can be true: narrow, well-defined problems are yielding to autonomous methods, while general scientific competence remains far off.
AI Agents as Researchers
-
Terminal-Bench Science: a benchmark for AI agents on real scientific workflows
Steven Dillmann et al. (Stanford, Laude Institute), August 2026
A community benchmark of 70 terminal-based research tasks across five scientific domains, each contributed by a working researcher from their own workflow, where the best agent (Claude Opus 5) passes about 30 percent.
-
Accelerating Scientific Research with Gemini in the Real-World
Samuel Schmidgall, Tao Tu et al. (Google), arXiv, August 27 2026
An updated multi-agent Gemini system extending the May Nature Co-Scientist paper to full closed-loop research cycles, including physically driving lab equipment, with a claimed autonomous discovery of an inference-time scaling architecture that outperformed six frontier models on HealthBench.
-
Autonomous Mathematical Discovery in an Open-World Multi-Agent Environment
Stephen Chung, Wenyu Du, William J. Wesley, arXiv, August 24 2026
AI agents from different model families, with no central coordinator, produced new results on Kakeya sets and kissing configurations in dimension 11, outputting theorems and explanations rather than just numerical solutions.
Research Practice
-
The future of peer review requires AI support, not AI bans
Nature (Xuegong Zhang), August 25 2026
Zhang argues that blanket AI bans in peer review address real concerns imperfectly and are impossible to monitor, calling instead for structured AI support that preserves reviewer accountability.
-
AI Use in The Scholarly Kitchen
The Scholarly Kitchen (David Crotty), August 27 2026
The Scholarly Kitchen adopted a formal AI disclosure policy for its own contributors, requiring disclosure of how, when, and which model was used, modeling the transparency standards it has long advocated for across scholarly publishing.