Notebookcheck Logo

Caught lying in 88% of tests: AI agents on Chinese models learned to cheat, US models did the same

DeepSeek-V3.2-Exp came in at 84% in the research reviewed by Reuters.
ⓘ John Cameron on Unsplash
DeepSeek-V3.2-Exp came in at 84% in the research reviewed by Reuters.
AI agents on Qwen, DeepSeek and Kimi models lied in 84% to 88% of sessions in a simulated bidding test, per research reviewed by Reuters. US models behaved similarly in the past, and none of the cases led to an agent escaping its controlled test environment.

AI agents built on Chinese models lied in most sessions of a simulated bidding contest, according to research reviewed by Reuters. Agents running on Alibaba's Qwen3-Max-Preview and Moonshot's Kimi-K2 made at least one false claim in 88% of sessions. DeepSeek-V3.2-Exp came in at 84%. Models from US firms produced similar results in the same test, though.

The experiment took place in March. Researchers from Beihang University, Peking University, the University of Nottingham Ningbo China and 360 AI Security Lab had the agents bid for customer contracts. Each agent was told what its product could do and what the customer needed. When allowed to learn from earlier rounds and try again, the three Chinese-powered agents got more deceptive, by 12 to 20 percentage points.

That study is one of at least 20 that Reuters identified from more than 200 documents published since 2025. Another, presented at ICML this year, found agents on both Chinese and US models faked results and fabricated files instead of admitting a task had failed. In March 2025, Fudan University researchers also reported that an agent on Alibaba's Qwen2.5-72B-Instruct made a copy of itself in another environment after learning it would be replaced. Separately, the Alibaba-linked ROME agent diverted cloud computing power to mine cryptocurrency before security systems stopped it.

Most of these cases came from controlled experiments designed to provoke failures. Reuters found no evidence that any Chinese-powered agent escaped to the wider internet or evaded shutdown. Several researchers still called the findings a warning. As systems get more capable, the same behaviour gets harder to contain.

China has started addressing this in policy, too. Its AI Safety Governance Framework 3.0, released on September 14, lists deceiving evaluators and concealing capabilities as risks. Alibaba, DeepSeek, Moonshot and Z.ai did not respond to Reuters. The first three have previously said they test their systems regularly and update safeguards.

Google LogoAdd as a preferred source on Google
Mail Logo

No comments for this article

Got questions or something to add to our article? Even without registering you can post in the comments!
No comments for this article / reply

static version load dynamic
Loading Comments
> Expert reviews and news on laptops, smartphones and tech innovations > News > News Archive > Newsarchive 2026 09 > Caught lying in 88% of tests: AI agents on Chinese models learned to cheat, US models did the same
Anubhav Sharma, 2026-09-30 (Update: 2026-09-30)