Notebookcheck Logo

The cloned voice on the phone: how you will not spot it

An older woman with white hair and glasses on the phone in a bright living room, holding her hand to her temple.
ⓘ Kaboompics.com / Pexels
Wait until the voice sounds tinny and you have already transferred the money.
Six signs are supposed to give away a synthetic voice on the phone. Two of them appear in no source at all, three are technically obsolete, and the one that still works has nothing to do with sound. The largest listening study to date also shows that people are not getting better at catching fakes. They are getting worse at believing real voices.

On a Friday afternoon in July, the phone rang at the home of an 80-year-old woman in Kühlungsborn on Germany’s Baltic coast. The caller claimed to be a doctor: her granddaughter had caused a serious accident, someone was dead, and 30,000 euros in bail were needed or she would go to prison. Then a woman took over the call, presenting herself as the victim’s daughter. Rostock police wrote in their 27 July statement, translated from the German: “Her voice resembled that of her actual daughter so closely that the 80-year-old initially believed the account.” The pensioner went to get the money. On her way to Rostock she ran into her real grandson, told him about it, and he called the police. It never got past an attempt.

Whether a cloned voice was involved, nobody knows. The police do not say so, and nobody checked. That is exactly where the problem this article is about begins.

Because what every advice column has promised for years is an answer to the question “how could she have known?”. Six signs circulate for that purpose. They sound plausible, they are everywhere, and most of them are no longer true.

The six signs, checked one by one

The source of the list is the deepfake page published by the BSI, Germany’s Federal Office for Information Security. It is the text everyone has been copying from ever since.

The page gives away its own age. In the section on countermeasures it still points to the EU Commission’s “draft regulation on AI systems” that would require labelling. That draft became law a long time ago. The AI Act entered into force on 1 August 2024, and its transparency obligations have applied since 2 August 2026.

Overview of the six common detection signs for cloned voices with their status: four obsolete or only partly valid, two without any source, one that holds up.
ⓘ Notebookcheck
Of six signs, one holds up. And it has nothing to do with sound.

Two things stand out. First, two of the most frequently cited signs do not appear on the BSI page at all. Neither “no background noise” nor “no breathing sounds” is there. They emerged through repetition, without any authority or study ever testing them.

Second, of the four that remain, the metallic sound is the best known and the weakest.

Why the delay disappeared

The BSI lists “noticeable delay” as a sign and explains it technically: a system has to receive part of the content before it can convert it. That used to be true. It is not any more.

To see why, two designs have to be kept apart, both of which get called “AI voice” in everyday language.

With text to speech, or TTS, someone types text and the system speaks it. With voice conversion, or VC, a human speaks and the system only swaps the timbre. Intonation, emotion, breathing and the ability to answer a question all come from a real person.

For a call in which the victim asks questions and cries, only the second design is usable. And it runs in real time. A measurement by the company Koe AI on an ordinary desktop CPU without a graphics card puts the LLVC method at 19.7 milliseconds of latency, QuickVC at 97.6 and RVC at 189.8. An ordinary mobile connection sits at 100 to 200 milliseconds. A conversion layer of 20 milliseconds vanishes into the line before anyone could hear it.

That leaves the question of source material. “Three seconds are enough” is the sentence attached to this everywhere. It comes from two sources that get mixed up. The serious one is VALL-E, a Microsoft Research system from January 2023 that synthesises from a three-second sample, though measured on clean audiobook and studio recordings. The more popular one is a figure from a May 2023 study by the security software vendor McAfee whose methodology was never published.

What current research measures on spontaneous everyday speech looks more sober. A 2026 study by the Barcelona Supercomputing Center loses roughly 31 percent of speaker similarity with the F5-TTS system when the reference material shrinks from 7.7 to 5.5 seconds. With StyleTTS2 the output additionally becomes partly unintelligible.

So the honest distinction is this. A few seconds are enough for “recognisably similar”. For “holds up through an entire phone call” it takes more, on current evidence something like ten to sixty seconds of clean connected speech. The London study whose clones genuinely could not be told apart from real voices used around four minutes of reference material per voice.

The BSI, incidentally, tried this itself, and that detail is more remarkable than any number. For a demonstration video the agency cloned the voice of its own president at the time, Arne Schönbohm. It used around ten minutes of audio taken from publicly available videos that were, in its own words, “only of medium quality”. The sentence spoken by the fake voice reads: “I am not real, and for now, you can recognise that fact. As technology matures, however, you will find this very, very difficult.” The second half of that sentence is what the rest of this article is about.

The point cuts the other way

In May 2026 Nicolas M. Müller and Wei Herng Choong of the Fraunhofer Institute for Applied and Integrated Security published the largest listening study on the subject so far. 1,768 participants, 35,532 individual judgments, 138 speech synthesis systems, directly comparable with a 2021 baseline study.

The expected result would be that people have become worse at spotting fakes. That is not the result.

On fakes, accuracy is almost unchanged at 71.2 percent against 72.9. What collapsed is the other side. On real recordings it fell from 72.7 to 64.1 percent. People have not become worse at hearing synthesis. They have become worse at believing real voices. The authors call it a “skepticism shift”.

For a phone call at half past eleven at night that means something uncomfortable. Anyone who resolves to listen more carefully next time mostly shifts the threshold at which they distrust their own daughter. The advice to listen to the sound makes things measurably worse, not better.

A second finding from the same study finishes off the checklists. The hardest to detect were not the exotic research systems but the commercial services and the language-model based systems, at 61.3 to 65.9 percent. The easiest were the older, technically simpler methods, at 75.4 to 76.8 percent. What anyone can book with a credit card sounds the most real.

A paper by Queen Mary University of London published in PLOS ONE in September 2025 arrives at the same picture. Voice clones were judged human in 58 percent of cases there, real voices in the same setup in 62 percent. The researchers stress explicitly that clones do not surpass real voices. Their paper is titled “Voice clones sound realistic but not (yet) hyperrealistic”. They do not sound more real than real. They sound exactly as real, and that is entirely sufficient.

An automated detector in the Fraunhofer study stayed above 94.5 percent across all conditions. Machines still hear the difference. People do not.

Bar chart. Detection of fakes stays nearly unchanged between 2021 and 2026 at around 72 percent. Detection of real recordings falls from 72.7 to 64.1 percent.
ⓘ Notebookcheck
People still hear the fakes. They believe real voices less.

What the German figures show, and what they do not

On 6 August Germany’s Federal Criminal Police Office released figures from the 2025 police crime statistics on request from the news agency dpa. The finding runs in two directions. For grandparent and shock calls the number of cases fell from 6,658 to 4,798, yet the losses still rose. For fake police officers both figures rose, with losses up by two thirds. That points to fewer but more targeted crimes with higher individual amounts.

Table from the 2025 German police crime statistics. Grandparent and shock calls: 4,798 cases and 49.0 million euros in losses. Fake police officer: 4,646 cases and 49.5 million euros.
ⓘ Notebookcheck
Fewer grandparent scam cases, bigger losses all the same.

And now the sentence that matters: the crime statistics do not record whether AI was involved. For Germany there is not a single reliable figure on how many shock calls were made with a cloned voice. Anyone quoting one did not measure it.

That does not stop such figures circulating. The most quoted, total losses of 10.6 billion euros, comes from a survey of 2,000 consumers by the Global Anti-Scam Alliance from June 2025 and covers all forms of fraud, not AI fraud. Along the way it turned into an estimate by the Federal Criminal Police Office and Europol, which neither agency ever gave. Similarly with the supposed half a million fraudulent calls the Federal Network Agency is said to have recorded in January 2026. That figure comes from a caller ID app and counts reports from its own user base.

What the authorities actually say is more cautious. The Federal Criminal Police Office told dpa that simulating a convincing grandchild’s voice is, translated from the German, “possible even without deep technical knowledge, using generally accessible AI systems”. That is a statement about possibility, not a case report. On the German police crime prevention page for shock calls, artificial intelligence and voice cloning do not appear at all to this day.

One recent case shows why restraint is warranted. On 29 April 2026 the Upper Palatinate police reported three completed shock call frauds in a single afternoon in Regensburg and Amberg. All three conversations were conducted in Russian, deliberately targeting Russian-speaking pensioners. Someone who negotiates freely in Russian and reacts flexibly to an 89-year-old’s questions is a human being on the phone.

What actually helps

If listening does not work, the only option left is to change the channel. Everything effective follows from that single idea.

Hang up and call back yourself. The German police put it this way on their prevention site, translated from the German: “Hang up immediately. That is not rude, it gives you room to breathe and to gather your thoughts.” And: “Contact the caller, or someone you trust, on a number you already know.” The last part is what matters. Not the redial key, but the number from your own address book, because displayed numbers can be spoofed.

A code word, or a question only the real person can answer. The police recommend agreeing on a “code word for identification on the phone”, and the Upper Palatinate force adds that it should be easy to remember and hard to guess. On its grandparent scam page the same authority recommends the more robust variant: “Ask the caller about things only the real relative or acquaintance could know.” A code word can be forgotten under shock. The family dog’s nickname from 1998 cannot, and no language model can guess it, because that information exists nowhere.

One caveat belongs here. Callers respond to questions with crying, time pressure and a handover to a supposed lawyer. Asking alone is not enough. The sequence has to be hang up, then call back.

Change your phone book entry. This is the most effective and the least known piece of advice. On its grandparent scam page the German police are unusually blunt, translated from the German: “Fraudsters use phone book entries to select victims for telephone fraud. Older first names, or short phone numbers, tell them that an elderly person is behind the entry.” Anyone who wants to stay listed should have their first name abbreviated. A ready-made form for the phone provider sits on the same page. For victims of phone fraud, changing the number is usually free.

Hand over nothing, and call 110 if in doubt. The police sentence is shorter than any explanation: “Never hand money or valuables such as jewellery to strangers, not even to the police.” Report it even if it stayed at an attempt. That is the only way patterns get spotted.

The new labelling duty does not help here

The AI Act’s transparency obligations have applied since 2 August 2026. On 5 August the Bavarian consumer advice centre summarised what that means and named this exact case, translated from the German: “This expressly includes synthetic voices. Anyone receiving a call or a voice message created with a cloned voice must be informed of that.”

Formally the duty falls on the caller, and the exemption for purely private, non-professional use does not cover a gang operating commercially. In practice it changes nothing for a victim. Nobody obeys a disclosure duty that would destroy the point of the crime. It is enforced administratively, with fines up to 15 million euros, against companies with an address and a turnover, not against anonymous callers abroad. In criminal law it remains fraud under section 263 of the German criminal code, exactly as before.

Tatjana Halm, head of law and digital affairs at the Bavarian consumer advice centre, says it herself, translated from the German: “The new rules create more transparency. But fraudsters do not follow rules, so a missing label is no guarantee of safety.”

That is the sentence that replaces the entire checklist. There is no label to wait for and no sound to listen for. There is only the second channel.

Google LogoAdd as a preferred source on Google
Mail Logo

No comments for this article

Got questions or something to add to our article? Even without registering you can post in the comments!
No comments for this article / reply

static version load dynamic
Loading Comments
> Expert Reviews and News on Laptops, Smartphones and Tech Innovations > Reviews > The cloned voice on the phone: how you will not spot it
Steffen Zahn, 2026-08-17 (Update: 2026-08-16)