Resemble AI's CEO tells LDS what the deepfake numbers miss
Zohaib Ahmed, CEO and co-founder of Resemble AI, answered Lets Data Science's questions at the launch of the company's Deepfake Threat Report and DETECT-World detector, and spent most of the exchange explaining what the headline numbers do not say. A single domain seizure carries roughly 14,000 of the 15,736 documented victims, the $2.24 billion figure is statutory exposure rather than losses, and 97.7 percent of verified attacks disclose no loss figure at all. He also gave externally validated false positive rates for the new detector and said which numbers remain internal.
Resemble AI published its Deepfake Threat Report today, alongside DETECT-World, a detector the company says judges whether audio, images and video behave like the physical world instead of only hunting generator artifacts. The report's headline is $2.24 billion in liability exposure across 15,736 documented deepfake victims. Before publishing, Lets Data Science put seven questions about the methodology to Zohaib Ahmed, the company's CEO and co-founder, and he answered all of them in writing. What makes the exchange worth reading is that he spent most of it explaining what his own numbers cannot say.
One seizure carries most of the victim count
Of the 15,736 documented victims, roughly 14,000 come from a single row: the June 12 seizure of CFAKE.com and SOCFAKE.com by the Department of Justice and Homeland Security, the first domain seizures announced under the TAKE IT DOWN Act. The sites published nonconsensual AI-generated sexual imagery of adult public figures, and the 14,000 figure is the enforcement agencies' own count, drawn from French investigators who identified about 300,000 images and 7,000 videos depicting roughly 14,000 people.
Ahmed volunteered the sensitivity rather than defending the headline. Without that row, the report is left with 1,736 documented victims, and statutory exposure of $137.2 million instead of $2.24 billion. "On every consequence axis in this dataset, a single incident carries most of the total," he told Lets Data Science. "One incident is 46% of documented losses, one is 87% of documented files, one is 89% of documented victims. That is not a quirk of our methodology, it is what disclosure-based data looks like."
The seizure also cuts against the report's own central finding. Excluding CFake, 61% of the remaining victims are private individuals rather than public figures. "Incidents cluster on famous people, but the harm mostly lands on ordinary ones," Ahmed said. "CFake cuts against that finding and we report it anyway."
A ceiling, not a damage estimate
The $2.24 billion is 14,915 documented NCII and CSAM victims multiplied by the $150,000 statutory election a victim can make under 15 U.S.C. 6851. Asked who actually carries that exposure, Ahmed declined to say. "We don't take a position on who carries it. That question is being litigated right now," he said, adding that the report labels the figure "exposure, never losses" and never adds the two together.
He did note which direction the statutes are moving. Under the DEFIANCE Act's $250,000 ceiling, passed by the Senate in January and pending in the House, the same victim count implies $3.73 billion. Courts and legislatures are already attaching real numbers: an Amsterdam court issued an injunction against an AI image generator at 100,000 euros per day per defendant, and Minnesota's nudification statute carries $500,000 per incident.
What the database structurally cannot see
The database is built from public disclosure, and Ahmed was blunt about the blind spots: completed frauds that companies never announce, sectors like healthcare that surface through enforcement rather than news, and abuse in closed channels that may never surface at all. In the small set of fraud cases where coverage states an outcome, he said, most of the frauds completed, which is why the company refuses to publish an interception rate.
The starkest gap is financial. 97.7% of verified attacks carry no stated loss figure, so the report documents $6.95 million in losses while the FBI logged $893 million in AI-enabled fraud complaints in 2025, a figure the bureau itself considers understated. Resemble publishes no multiplier of its own; where it extrapolates, it uses published reporting rates of 14.3% to 26% and labels the result implied rather than documented. The same discipline ran through the Grok analysis: the company reproduced the widely cited figure of 23,000 images appearing to depict children as CCDH's modeled estimate, carrying two layers of sampling uncertainty, and rejected an "80 million Grok images" number it could not trace to a primary source.
The detector's numbers, labeled unevenly on purpose
Asked for the false positive rate behind DETECT-World's 99.5% audio accuracy claim, Ahmed gave it: 0.66%, externally validated on a Podonos benchmark published today, against 2.50% for the company's previous detector. Video and image are a different story, at 2.01% and 4.31% respectively, and both remain internal measurements. "I will label them that way rather than present them as equivalent," he said.
On day-zero coverage, the company ran a fixed evaluation across more than 260 generator models it had not trained on, with at least 50 samples each. Detection on deepspeak2 moved from 32.0% under pattern-matching to 96.0% with the world-model approach, with similar jumps on sadtalker and wav2lip. Asked whether world models simply arm both sides, Ahmed did not argue. "We presume detection and generation will keep improving against each other indefinitely," he said. The number that matters, in his view, is how much performance survives in the minutes and hours after a new generator ships.
What practitioners should change on Monday
Ahmed's advice for engineers building identity verification, independent of buying anything: prioritize image and multimodal detection, because image-only attacks now make up 37.1% of incidents against 28.6% for video-only. Stop relying on repeat-target signals for intimate-imagery abuse, since NCII in the data shows zero repeat targets, meaning hash-matching catches recirculation but not the first-generation attacks NCII now mostly is. And put verification friction in-line, before money moves: in every reported fraud case with a stated save, the interception happened at the transfer approval, the callback or the verification step. "Detection after the fact gives you evidence rather than your money back."
Key Points
- 1Resemble AI CEO Zohaib Ahmed told Lets Data Science that one seizure row carries 89% of documented victims, and that the report publishes every headline number with and without it: $2.24B in statutory exposure becomes $137.2M.
- 2The company documents only $6.95M in losses because 97.7% of verified attacks disclose no figure, against $893M in AI-enabled fraud complaints the FBI logged in 2025, a gap Ahmed says is what disclosure-based data looks like.
- 3DETECT-World's audio false positive rate is 0.66%, externally validated by Podonos; video and image FPRs are internal numbers, and Ahmed labels them that way rather than presenting them as equivalent.
Scoring Rationale
Exclusive written Q&A provided directly to Lets Data Science by the CEO of Resemble AI at the lift of the Deepfake Threat Report and DETECT-World launch, covering the sensitivity of the headline numbers, the limits of disclosure-based data, and externally validated false positive rates.
Sources
Original reporting, with the public references used alongside it.
LDS Exclusive
Reporting based on written answers given directly to Let's Data Science by Zohaib Ahmed, CEO and co-founder, Resemble AI.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems

