Deepfake production has gone from roughly 140,000 pieces a month in 2023 to more than 900,000 in 2026, and the operational damage is already documented. In 2026 the Singapore Police confirmed a case in which fraudsters ran a Zoom conference populated by synthetic government officials and walked a single victim out of at least S$4.9 million. The question for security teams is no longer whether generated video is convincing enough to fool someone. It is what a payment process does once video evidence stops counting as evidence at all.
The Singapore call is worth reading closely, because the attackers did not win on picture quality. They won on roster. The AI-generated video and audio impersonated Prime Minister Wong, President Tharman Shanmugaratnam, Minister Indranee Rajah and representatives of the Monetary Authority of Singapore. Foreign officials were added for depth: Canada's foreign minister and a senior diplomatic adviser from the UAE. Then came the private sector, including BlackRock and the Dubai International Financial Center. After the call ended, someone presenting himself as a lawyer asked the victim to move the money.
Every one of those names is a credential the victim had no way to check in the moment, and checking them afterwards would have confirmed that each person exists and holds the office claimed. That is the design. The fraud does not ask you to believe an impossible thing; it asks you to believe a plausible meeting.
The economics behind this shift are simple and they favour the attacker. Impersonating a voice or a face once required particular skills, large volumes of training data and time. Now the same operation runs quickly and cheaply on publicly available tools. The result is an asymmetry that no amount of staff training fully closes: the attacker needs to succeed once, the defender needs to be right every time.
The attack surface widened on its own, without anybody's help. Voice clones can be built from a few seconds of publicly available recording and deployed against relatives, executives and officials over phone calls and live streams. Face-swap and lip-sync systems support real-time video calls in which every participant except the victim is synthetic. There are already known cases of multimillion-dollar bank transfers approved by employees after video conferences with a deepfaked chief financial officer and deepfaked colleagues. The same toolkit shows up in phishing, business email compromise, romance scams and political influence operations.
The training material is free and voluntarily published. LinkedIn posts, conference talks, podcasts and social platforms supply the raw audio and video; the dark web accelerates distribution of the tools and the methods. The pandemic shift to remote work moved a large share of professional and personal contact onto video platforms, which means the channel most people now treat as proof of presence is the channel most exposed.
For now, the tells still exist. Stiff or constrained mannerisms. Unnatural eye contact or blinking. Flat or excessively linear intonation. Imperfect lip sync. Inconsistent lighting or reflections. Delayed responses to unexpected questions.
Those signals work today, and the window in which they work is closing. The models improve quickly, and real-time systems are becoming harder to separate from a real call by eye and ear alone.
The technical countermeasures on the table are watermarking, cryptographic provenance for content, commercial detection tools and AI-versus-AI analysis. None of them offers a full guarantee, particularly at the volume of material moving through platforms, and steganography and other data-embedding techniques make the problem harder still. Quantum computing is not an immediate threat to most organisations, but it will eventually strain existing cryptographic mechanisms, which is the argument for planning the move to post-quantum cryptography now: data stolen today should not be readable later.
Here is what stands out to me about the whole framing. Almost every defence that actually works in this account is procedural rather than technical — confirm identity through a second channel, keep an escalation path, do not let one video call authorise a wire. That is a quiet admission that detection has already lost the arms race, or is losing it fast enough that betting a treasury function on it is negligent. The 900,000-a-month figure carries no counting method with it, and a number that large and that round should be treated as a direction rather than a measurement. The Singapore roster is the more instructive datum: the attackers scaled by adding institutions, not by improving pixels, which means the defence has to sit at the point where authority is accepted, not at the point where video is rendered. The post-quantum paragraph is the odd one out in this argument — a real risk, but on a different clock, aimed at a different attacker, and it does not help the employee on a call with a synthetic finance chief this afternoon.
The recommended protocol is old and unglamorous: trust, but verify, applied to people, processes and technology. Confirm identity before any critical interaction involving money, credentials or confidential decisions, using an independent trusted channel — a known phone number, a mutual contact, a process agreed in advance. Do not rely on a video call or a voice message alone. Run behavioural and contextual checks: naturalness of reaction, hand movement, eye contact, answers to unexpected questions. Verify profiles over time rather than at a glance — posting history, mutual connections, stated employer, internal consistency.
The rest is hygiene and governance. Multifactor authentication, strong unique passwords or passkeys, network segmentation, encrypted channels, regular backups. Zero-trust principles extended past networks into human interactions and supply chains. Deepfake training treated the way phishing training is treated. Clear escalation routes for suspicious requests, and AI governance that makes the defensive tools themselves transparent and auditable. Until proven otherwise, a new digital relationship should be treated as potentially synthetic. Cybersecurity at this point is a board-level risk and a fiduciary duty, not a support function.
All of that is friction, and friction is precisely what a decade of remote collaboration was sold as removing. Organisations are about to learn the price of digital trust either as slower approvals and more phone calls back to known numbers, or as transfers they cannot reverse. Seeing is no longer sufficient grounds for believing, and the cost of verification is now a line item whether or not anyone budgets for it.