A Breakthrough in Autonomous Science
A Chinese artificial intelligence system has officially topped an international ranking for autonomous scientific research. The Zhejiang University-led Qiushi Engine recently pulled ahead of Anthropic’s Claude Code and other top global agents. Currently, this advanced system holds the top overall spot on the ResearchClawBench leaderboard. The international benchmark explicitly tests the ability of AI agents to independently carry out complex research tasks. It compares their final results against reference papers written by humans to see if they can match or outdo the original authors. Qiushi Engine is designed to perform scientific discoveries within real physical environments. Furthermore, its developers claim the system is fully capable of end-to-end autonomous scientific discovery.
Leaderboard Performance and Current Limitations
To evaluate these systems, ResearchClawBench scores agents across ten distinct scientific disciplines using a 100-point scale. A score of 50 indicates that an agent has successfully matched the conclusions of the original human paper. As of Tuesday, only the Qiushi Engine achieved an overall score above 30 by utilizing OpenAI’s GPT-5.5 model. Meanwhile, all other competing agents scored in the 20s or lower. Consequently, developers emphasized that even the strongest systems still remain far from reliable rediscovery. Common failures occur when the generated conclusions are mismatched to critical evidence. In addition, agents frequently fail to identify the core scientific mechanism of the reference paper.
A Future Paradigm for Research
Despite these limitations, the Qiushi Engine represents a significant leap forward for independent research tools. Previously, alternative platforms like the experimental Open Science Desktop and Claude Code held top positions on the leaderboard. However, prior systems could not demonstrate end-to-end autonomous discovery starting from an open-ended theme. The Zhejiang University team successfully proved this new capability using an open-ended prompt on optical computing. The agent independently identified a research direction, formulated a physical mechanism, and validated the final results. Therefore, the developers believe the system represents a future paradigm where autonomous systems actively drive scientific breakthroughs.
Reference
Bela, V., & Bela, V. (2026b, julio 21). Chinese AI agent Qiushi Engine outperforms Anthropic’s Claude Code in autonomous research. South China Morning Post. https://www.scmp.com/news/china/science/article/3361370/chinese-ai-agent-outperforms-anthropics-claude-code-autonomous-research?share=GqwoUehosyxYcuOzMfDmIdctGV%2BW3Ya2%2FbtkOZebW7QsYK3Nsslrr20EhHGtsZmnoYPxhbVLfxvz5jBikkC1%2FLzQHk926rc6IOJGVvhc5tY%3D&utm_campaign=social_share
