Chinese AI agent outperforms Anthropic’s Claude Code in autonomous research
Zhejiang University’s Qiushi Engine topped the ResearchClawBench leaderboard, but it can’t reliably make new discoveries yet

ResearchClawBench tests the ability of AI agents to independently carry out research and compares their results against reference papers written by humans to see if they can reach the same conclusions or even outdo the original authors.
The benchmark, created by a team led by the Shanghai Artificial Intelligence Laboratory, was designed to assess whether agents can really conduct the kind of tasks their creators say they can handle.
Qiushi Engine, which was officially launched by a Zhejiang University-led team last week, is a large language model-based agent designed to perform scientific research in real physical environments.
Its developers said that unlike some other existing systems that could be limited to performing specific tasks, Qiushi Engine was capable of “end-to-end autonomous scientific discovery”.