A multi-institution research group led by Peter Kirgis and Sayash Kapoor at Princeton University has found that AI agents currently lack the judgment and creativity required to conduct original, open-ended AI research. While the agents proved capable of handling research engineering—such as reviewing literature, running experiments, and compiling results—they failed to produce work of the caliber required by top machine-learning conferences.

To test these qualitative skills, researchers employed a "shadow evaluation" method, requiring AI to answer research questions from high-quality, unpublished papers. Using Anthropic’s Claude Opus 4.8 via OpenClaw software, the agents attempted to answer questions from papers submitted to the NeurIPS 2026 conference. In both instances, the original authors rejected the agents' papers.

The study noted that while the agents could solve engineering problems, they struggled with the fundamental aspects of the scientific process. They failed to explore diverse ideas, often committed to unpromising approaches too quickly, and lacked the ability to fundamentally rethink their methodology after a failure. Instead of revising their approaches based on feedback from subagents or external tools, they often narrowed their claims or added caveats.

Kapoor suggested that these limitations may stem from current training methods. While reinforcement learning is effective for tasks with automatically checkable success criteria, creating training environments for open-ended research remains a challenge. This finding suggests that the timeline for "recursive self-improvement"—a process where AI systems accelerate their own development—may be longer than some industry forecasts suggest.


Sources: