WhatsApp Scam Alert: The AI only checks strangers, not your contacts

On August 12, Meta published a technical overview of Scam Alert, a new feature for WhatsApp. It is designed to warn users about scam messages before anyone falls for them. The verification process takes place entirely on the user’s phone.
Once enabled, WhatsApp downloads a small AI model to the device. This model analyzes incoming messages and compares them to patterns from reported scams, meaning conversation threads and phrasing typical of known scams. No message content leaves the phone for this purpose. End-to-end encryption, where only the sender and recipient see the plain text, remains intact.
Only the recipient sees the warning
If the model detects a likely scam attempt, a notification appears directly in the chat. The sender is not notified. From there, the user can block the sender, report them, or simply continue the conversation. WhatsApp does not automatically report anything to Meta. This only happens if the user actively taps “Report.”
Only messages from senders not in the address book are checked
This is the limitation, and it’s stated in Meta’s own description. The model exclusively classifies messages from senders who aren’t saved in the user’s contacts. Messages from people already in the address book aren’t checked.
This means that, ironically, the scam that works best is left out. If a friend’s or relative’s WhatsApp account is hijacked, the scam message comes from a real, familiar account. Scam Alert remains silent in such cases. The case of a fake app that hijacks WhatsApp accounts demonstrated how such a takeover works. Even the shocking call with a cloned voice bypasses this check because it doesn’t come through WhatsApp at all.
Mark it as trustworthy once, and you’ll never see a warning again
If you believe a warning is false, you can mark the chat as trustworthy. The warning then disappears permanently for that contact. This is convenient, but it’s also the weak spot: this is exactly the kind of decision you make under pressure when a skilled scammer has already established a conversation. Anyone who marks the chat this way can also voluntarily submit the last five messages received to WhatsApp.
Built for verifiability
When it comes to transparency, Meta goes further than usual. Each model version is recorded in a public registry along with its checksum before it is deployed. The signature comes from Cloudflare; Meta does not possess the key required for this. The model weights are published so that researchers can verify whether the model truly detects only scams. Users themselves can find a log under “Account” > “Request Info” > “Scam Alert Activity” showing which messages were checked and which model version was running at the time.
Statistics on how often warnings were issued and how users responded are processed in isolated computing environments. Before Meta sees them, the data is aggregated and obfuscated so that no conclusions can be drawn about individual users.
It’s not available yet
Scam Alert is currently running as part of a limited beta alongside Meta’s bug bounty program. The company has not announced a date for the full launch or the countries where it will be available. There has been no statement regarding Germany so far. Until then, the old rule applies: If you receive a request for money or a supposedly new phone number, contact the person in question through a second channel, preferably by calling the old number. If you want to know how quickly an account can be compromised even without a password, check out the report on the stolen session cookie.





