BENCHMARKS

Benchmarks for Deepfake Detection and Watermarking

Every result comes from public leaderboards or reproducible evaluations. We show methodology, source, and test conditions — not just the number.
99.5%
Audio detection accuracy • Podonos Audio DFD Bench
#1
DFBench • Speech and Image detection
54
languages detected from Arabic to Spanish
250+
Generative models tested against for detection
Model

DETECT-World ranked #1 across audio, image, and video detection

DETECT-World is the first deepfake detection model built on a world model architecture with the highest accuracy on third-party benchmarks.
99.5%
Pooled detection accuracy, Podonos Audio DFD Bench
Last updated August 2026
IMAGE
95.8%
Average accuracy on images
False negative rate of 4.3% and false positive rate of 4% when benchmarked against 1,000 image files.
VIDEO
98.2%
Average accuracy on video
False negative rate of 2% and false positive rate of 0.95% when benchmarked against 1,000 video files.
RTF
0.12
Real-time factor
shows ultralow latency that provides results almost instantly.
Deepfake detection benchmarks
RESEMBLE AI

COMPETITOR
Audio deepfake detection — Accuracy
Pooled accuracy across all audio files. Higher is better.
Resemble DETECT-World
99.5%
Whispeak
97.7%
Aurigin AI
96.8%
Pella Research
95.8%
Pindrop
95%
0%
Accuracy % (higher = better)
100%
Audio DFD Benchmark · Podonos · August 2026 · Source
Audio deepfake detection — False Negative Rate (FNR)
Misidentification of AI files as authentic. Lower is better.
Resemble DETECT-World
0.4%
Whispeak
1.7%
Pella Research
2.8%
Reality Defender
3.6%
Pindrop
3.7%
0%
False negative rate (lower = better)
5%
Test sets include mp3, wav, flac, ogg, m4a, webm across 4,524 files. Source
Accuracy vs Real-time Factor (RTF)
RESEMBLE AI

COMMERCIAL COMPETITOR
COMPETITOR
Audio deepfake detection
Accuracy, higher = better, RTF: lower speed = better
#
Audio DFD Benchmark · Podonos · June 2026 · Source
LANGUAGE COVERAGE
Over 50 languages detected
ALSO SUPPORTED BY CHATTERBOX TTS
Arabic
Chinese
Danish
Dutch
English
Finnish
French
German
Greek
Hebrew
Hindi
Italian
Japanese
Korean
Malay
Norwegian
Polish
Portuguese
Russian
Spanish
Swahili
Swedish
Turkish
Ukranian
Vietnamese
Thai
Indonesian
Romanian
Czech
Slovak
+ over 20 additional languages
Validated against MLAADv10. Detection relies on generation artifact patterns, not language-specific features — enabling generalization to languages not seen during training. EER and accuracy figures from the Hugging Face Speech DF Arena and DFBench Speech/Image leaderboards, March 2026. Resemble AI does not control test set composition. Image figures from DFBench Image 2025. Video EER from internal evaluations on held-out test sets.
VERIFY

PerTh watermarker: survives compression, re-encoding, and attack

Detection accuracy across 18 real-world attack conditions. PerTh V2 ships with improved robustness across reverb, pitch shift, and spectral manipulation.
~100%
Detection on clean and compressed audio
PerTh V2 · No-attack + standard codecs
STRONG (>90%)
MODERATE (70-90%)
WEAK (<70%)
PerTh — Open Source, Audio Only
wav_dither_attack
100%
random_wav_wavelet_noise
90%
random_wav_reverb_attack
95%
random_wav_resample_attack
100%
random_wav_precision_attack
100%
random_wav_pitch_shift_attack
10%
random_wav_mulaw_attack
100%
random_wav_high_pass_attack
100%
random_wav_gaussian_noise_clipped
45%
random_spec_time_mask
100%
random_spec_stretch
100%
random_spec_scale
100%
random_spec_lowclip
100%
random_spec_highclip
100%
random_spec_gaussian_noise_clipped
100%
random_spec_contiguous_band_mask
72%
no_watermark
100%
no_attack
100%
0%
Accuracy
100%
PerTh Multimodal
wav_dither_attack
100%
random_wav_wavelet_noise
100%
random_wav_reverb_attack
88%
random_wav_resample_attack
100%
random_wav_precision_attack
100%
random_wav_pitch_shift_attack
94%
random_wav_mulaw_attack
100%
random_wav_high_pass_attack
100%
random_wav_gaussian_noise_clipped
100%
random_spec_time_mask
100%
random_spec_stretch
98%
random_spec_scale
100%
random_spec_lowclip
100%
random_spec_highclip
100%
random_spec_gaussian_noise_clipped
100%
random_spec_contiguous_band_mask
98%
no_watermark
100%
no_attack
100%
0%
Accuracy
100%
Watermark robustness = detection accuracy after each attack transform (1.0 = 100%). "no_attack" = clean audio. "no_watermark" = false positive rate on unwatermarked audio. PerTh V2 figures are pre-release internal evaluations.
DEPLOYMENT / INFRA

Deploy on the infrastructure that meets your needs.

Resemble AI models — DETECT-3B Omni and PerTh run on cloud API, on-prem, or air-gapped. Pick the deployment model that fits your security posture and latency requirements.

CLOUD API
Managed service

RESTful API, 99.9% uptime SLA, auto-scaling. Fastest time to value.

ON-PREMISES
Your infrastructure

Docker / Kubernetes. No data leaves your network. NVIDIA GPU optimized.

AIR-GAPPED
Zero internet dependency

Run Resemble AI in your own AWS, GCP, or Azure VPC. Isolated compute, private networking, and compliance with your cloud security policies.

Infra partners include

All infra partners + integrations
Get complete generative AI security
Join thousands of developers and enterprises securing with Resemble AI