Researchers Tried Letting AI Do Science. It Failed

Summary

A new study tested whether frontier AI agents can independently do open-ended AI research. Researchers from Princeton, Stanford, Toronto, and others gave agents the core questions from two unpublished NeurIPS 2026 papers, plus six days, API credits, GPUs, internet access, and a virtual machine. The agents handled much of the workflow well: literature review, debugging, experiments, resource management, and writing full papers. But both outputs were rejected by the original authors as lacking original scientific contributions and not meeting top-conference standards. The result suggests current agents can automate many engineering parts of research but still struggle to produce novel, publishable ideas. The study is limited by its small sample and author-based evaluation, but it aims to measure scientific reasoning better than task-based benchmarks.