Deepfake-as-a-Service Is a Twelve Billion Dollar Industry
From: Lurk More
Contents 8 sections
Lurk More Newsletter
In January 2024, a finance worker at Arup – a multinational engineering firm headquartered in London – joined a video call with his chief financial officer and several colleagues. The CFO asked him to execute a series of confidential fund transfers. The employee complied, wiring HK$200 million – approximately $25.6 million – across fifteen transactions to five separate bank accounts.
Every person on that call was fake. The CFO was fake. The colleagues were fake. The entire meeting was a deepfake, generated in real time using publicly available video and audio of the real executives scraped from online conferences and company presentations.
The employee was not stupid. He was not careless. He was suspicious at first – the initial email had looked like phishing. But then he joined the call and saw his colleagues’ faces, heard their voices, watched them nod and respond in real time. He did what any reasonable person would do: he trusted the evidence of his senses.
His senses were wrong. In 2024, that was a $25.6 million mistake. By 2025, it is a rounding error.
The Numbers
The growth curve is not gradual. It is vertical.
In 2023, there were approximately 500,000 deepfakes circulating online. By 2025, that number reached 8 million. That is not a percentage increase. That is a sixteen-fold explosion in two years. The annual growth rate is approaching 900%.
Europol – the European Union Agency for Law Enforcement Cooperation, not a think tank, not a vendor selling detection software, but the actual police – estimates that 90% of online content may be generated synthetically by 2026. Let that number sit for a moment. Nine out of ten things you encounter online may not have been created by the person or entity that appears to have created them.
Global losses from deepfake-enabled fraud exceeded $200 million in the first quarter of 2025 alone. Some major retailers report receiving over 1,000 AI-generated scam calls per day. Not per month. Per day.
The content moderation industry – the people whose job it is to detect and remove this material – is worth $12.48 billion as of 2025, projected to grow to $42 billion by 2035 at a compound annual growth rate of 13%. That is the defensive side of the arms race. The offensive side is cheaper, faster, and does not need to be right every time. It just needs to be right once.
The Threshold
In late 2025, researchers began using a specific phrase: voice cloning has crossed the “indistinguishable threshold.” The term is precise and its implications are catastrophic.
A few seconds of audio – a voicemail, a conference recording, a podcast appearance, a YouTube video – now suffice to generate a voice clone that reproduces not just the words but the intonation, rhythm, emphasis, emotion, pauses, and breathing patterns of the original speaker. For low-resolution video calls, for media shared on social platforms, for the vast majority of everyday communication contexts, synthetic voices and faces are now indistinguishable from authentic recordings by ordinary people.
Not “hard to distinguish.” Not “occasionally convincing.” Indistinguishable. The test subjects cannot tell. The technology has crossed the line where human perception is no longer a reliable detection mechanism.
This is a phase transition. Everything before the threshold is a problem. Everything after it is a different kind of problem entirely. Before the threshold, you could tell people to look for artifacts – the weird hands, the uncanny valley, the audio glitches. After the threshold, there is nothing to look for. The fake is perfect. The only way to detect it is with other AI, and that AI is losing the race.
The Business Model
Deepfake-as-a-Service is not a metaphor. It is a literal business model.
The tools are commercially available. Some are marketed for legitimate purposes – accessibility, entertainment, dubbing, voice preservation for people with degenerative diseases. The same tools, with no modification, produce weapons-grade synthetic media. The distinction between “legitimate voice synthesis platform” and “fraud toolkit” is a marketing decision, not a technical one.
The economics are straightforward. Creating a convincing deepfake used to require significant computational resources and technical expertise. It now requires a consumer laptop, a free or low-cost software tool, and a few seconds of source material. The marginal cost of producing a deepfake has fallen to approximately zero. The marginal cost of detecting one has not.
This is the structural problem. Defense is expensive. Offense is cheap. Every detection system that is built must be maintained, updated, and scaled across billions of pieces of content. Every attack needs to succeed once. The economics favor the attacker absolutely and permanently.
The Arup Pattern
The Arup case is instructive because it is not exotic. It is not a nation-state operation. It is not a sophisticated intelligence agency running a years-long influence campaign. It is criminals using commercially available tools to impersonate executives on a video call.
The attack surface is every organization in the world. Every company whose executives have ever appeared on video – which is every company – now has a voice and face library available to attackers. Every earnings call, every conference keynote, every webinar, every podcast appearance is training data. The executives cannot undo this. The data is public. The horse is not merely out of the barn. The barn has been demolished and the horse is running free across the continent.
By March 2025, a similar attack hit a multinational firm in Singapore. The pattern is replicating because it works. The return on investment for deepfake fraud is extraordinary. A few hours of preparation can yield millions of dollars. Compare this to traditional social engineering, which requires weeks or months of relationship-building. The deepfake collapses the entire attack timeline into a single meeting.
The Post-Evidence Era
In the final chapter of Lurk More, I describe what I call the post-evidence era. The argument is not that evidence stops existing. It is that the relationship between evidence and belief breaks down.
When any piece of media – any photograph, any video, any audio recording – could be synthetic, the evidentiary value of media approaches zero. Not because all media is fake, but because all media could be fake, and there is no way for a non-expert to tell the difference. The result is not that people believe nothing. The result is that people believe whatever they wanted to believe already, because the external check on belief – the photographic proof, the recorded confession, the video evidence – no longer functions as a check.
This is already happening. In courtrooms, defendants are challenging video evidence by claiming it could be a deepfake. In politics, authentic recordings of politicians saying embarrassing things are dismissed as AI-generated. The deepfake does not need to be deployed to do damage. The possibility of the deepfake is sufficient. It creates universal plausible deniability for anything captured on any recording device.
The phrase “seeing is believing” governed human epistemology for the entire history of recorded media. Photography. Film. Video. Audio recording. For roughly 180 years, a recording was proof. That era is over. It ended not with a dramatic announcement but with a threshold crossing that most people have not yet noticed.
The Detection Problem
The detection industry is real, well-funded, and losing.
Every detection model trains on a corpus of known deepfakes. The generative models then train on the detection models. The cycle repeats. This is an adversarial arms race with no stable equilibrium. The detection side must be perfect – a 99% detection rate means 80,000 undetected deepfakes out of the current 8 million. The generative side must be good enough, which is a lower bar that keeps getting lower.
The fundamental asymmetry is computational. Generating a deepfake requires a single forward pass through a neural network. Detecting a deepfake requires analysis of multiple signal layers – visual artifacts, audio spectral analysis, temporal consistency, metadata verification. Generation is fast. Detection is slow. Generation is cheap. Detection is expensive. Generation scales linearly. Detection scales quadratically.
There is no technological solution to this problem. There are technological mitigations – watermarking, provenance tracking, cryptographic signing of authentic media. All of them require universal adoption to work. None of them have achieved it. The C2PA content provenance standard is technically sound and almost nobody uses it. The watermarking solutions can be stripped. The metadata can be forged.
What This Means
The content moderation industry exists because humans produce an enormous volume of harmful content and other humans have to sort through it. That was the old problem. The new problem is that machines produce an effectively infinite volume of harmful content and the humans cannot keep up.
Twelve billion dollars. That is what the world currently spends trying to moderate human-generated content. The synthetic content wave will make that number look quaint. When 90% of online content is machine-generated – and Europol says we are one year away from that – the moderation problem becomes mathematically unsolvable through human review. There are not enough people. There will never be enough people.
The solution, presumably, is AI-powered moderation of AI-generated content. Machines policing machines. The human removed from the loop entirely. If that sounds like a stable equilibrium to you, I have a compliance framework to sell you.
In Lurk More, I trace the arc from trolling to deepfakes as a story about the democratization of deception. Trolling was deception performed by humans, at human scale, with human limitations. Deepfakes are deception performed by machines, at machine scale, with no limitations that matter. The skill that used to require talent – the ability to convincingly pretend to be someone you are not – has been automated. The cost has been driven to zero. The volume has been driven to infinity.
The internet did not create deception. It scaled it. AI did not create synthetic media. It made it free. The question is not whether the technology can be controlled. The question is what happens to a civilization that can no longer distinguish real from fake and has decided, as a matter of policy, to punish the people who were trying to help.
This essay draws from Lurk More, coming fall 2026.
Sources
- Fortune: 2026 will be the year you get fooled by a deepfake, researcher says (December 27, 2025)
- Deepstrike: Deepfake Statistics 2025: The Data Behind the AI Fraud Wave
- Keepnet Labs: Deepfake Statistics & Trends 2026
- CFO Dive: Scammers siphon $25M from engineering firm Arup via AI deepfake ‘CFO’
- ScamWatch HQ: The $200 Million Deepfake Disaster
- [European Parliamentary Research Service: Deepfakes Briefing (2025)](https://www.europarl.europa.eu/RegData/etudes/BRIE/2025/775855/EPRS_BRI%282025%29775855_EN.pdf)
- Darknet.org.uk: Deepfake-as-a-Service 2025
- Research Nester: Content Moderation Services Market Size to 2035
Prefer RSS? Subscribe here.