Back
Research
May 18, 2026

MLAAD: The Multi-Language Audio Anti-Spoofing Dataset

CONTENTS
Active heading
Section heading
CONTRIBUTORS
Piotr Kawa
Deep Learning Researcher
Nicolas Müller
Machine Learning Engineer for Audio Deepfake Detection

This research introduces MLAAD, a large-scale, multi-language dataset for training and evaluating audio deepfake detectors, authored by Resemble AI's Nicolas Müller and Piotr Kawa alongside researchers at Fraunhofer AISEC, Wrocław University of Science and Technology, Coqui.ai, and Thorsten-Voice. Peer reviewed and published at IJCNN 2024, with the underlying dataset continuously expanded since.

Overview

Almost every audio deepfake detector is only as good as the data it was trained on, and most training data is English (or Chinese). This means detection systems are systematically less reliable for the majority of the world's languages, at exactly the moment deepfakes are showing up in non-English media, elections, and fraud attempts globally.

MLAAD solves this. It takes real human speech in eight languages and extends it, via translation and 175 different text-to-speech systems, into synthetic audio spanning 54 languages, giving researchers and companies a way to train and test detectors that don't just work for English speakers. Two of Resemble AI's researchers, Nicolas Müller and Piotr Kawa, are core authors on this work.

Key Findings
54 languages spanned by the synthetic audio in the dataset
175 TTS systems used to generate the synthetic speech
105 architectures underlying model designs represented across those systems
1,002.9 hours total synthesized audio
456,000 individual utterances in the dataset
4 of 8 cross-dataset tests where MLAAD outperformed ASVspoof19 (complementary, not redundant)


Download the full research PDF on arXiv
.

Methodology

MLAAD starts from real human speech, audiobooks and public speeches in eight languages (English, French, German, Italian, Polish, Russian, Spanish, and Ukrainian), pulled from the existing M-AILABS Speech Dataset. For languages outside that original set, the team translates the same source text using machine translation, then synthesizes it using 175 different text-to-speech systems, everything from older baseline models to current state-of-the-art voice cloning tools.

Every synthetic clip is paired with metadata: which model made it, what architecture it's based on, the transcript, and (as of the dataset's 8th version) which reference speaker was used.

To prove the dataset is actually useful, the team trained three different detection models on four training datasets (including an early snapshot of MLAAD) and tested all of them across eight datasets total. That cross-dataset design matters: a detector that only performs well on the exact dataset it was trained on isn't proving much, the real test is whether it generalizes to audio it's never seen.

Results

The headline result: no single dataset wins across the board. MLAAD and ASVspoof19, the field's long-standing benchmark, each produced the best-performing detector on four of the eight test datasets. That's a meaningful finding in itself, it means MLAAD isn't just "more data," it's data that teaches detectors something ASVspoof19 doesn't, and vice versa.

A second, subtler finding demonstrates that detectors trained on some datasets developed shortcuts that backfired, in one case scoring effectively at random (51%) on an unfamiliar test set, and in several cases scoring below 50%, worse than a coin flip, meaning the model had learned something actively misleading. Detectors trained on MLAAD were the only ones that never dropped meaningfully below chance, suggesting the language and model diversity in MLAAD makes it harder for a detector to latch onto a misleading shortcut.

Why this matters: A detector that's only trained and validated on English audio is a detector that will quietly underperform for most of the world, at the exact moment deepfakes are showing up in non-English elections, fraud attempts, and media manipulation. MLAAD is infrastructure-level work: it doesn't detect anything itself, but it's the training and evaluation resource that makes better multi-language detectors possible, for Resemble and for the field generally.

Limitations

  • This is a continuously updated dataset, not a fixed one-time release. Any specific numbers on this page reflect a snapshot in time (v10, May 2026) 
  • MLAAD is spoof-data only, so in order to train on it, one needs to pair it with bona-fide sources (the authors recommend M-AILABS)

How to cite this paper

APA Müller, N. M., Kawa, P., Choong, W. H., Casanova, E., Gölge, E., Müller, T., Syga, P., Sperl, P., & Böttinger, K. (2024). MLAAD: The multi-language audio anti-spoofing dataset. 2024 International Joint Conference on Neural Networks (IJCNN), 1–7.

BibTeX @inproceedings{muller2024mlaad, title={MLAAD: The Multi-Language Audio Anti-Spoofing Dataset}, author={M{"u}ller, Nicolas M. and Kawa, Piotr and Choong, Wei Herng and Casanova, Edresson and G{"o}lge, Eren and M{"u}ller, Thorsten and Syga, Piotr and Sperl, Philip and B{"o}ttinger, Konstantin}, booktitle={2024 International Joint Conference on Neural Networks (IJCNN)}, pages={1--7}, year={2024}, organization={IEEE} }

Sources

Curated subset, the full paper cites 68

  1. Wang et al. ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech. Computer Speech & Language, 64, 2020.
  2. Müller et al. Does audio deepfake detection generalize? INTERSPEECH, 2022.
  3. Yamagishi et al. ASVspoof 2021: accelerating progress in spoofed and deepfake speech detection. arXiv:2109.00537, 2021.
  4. Frank and Schönherr. WaveFake: A data set to facilitate audio deepfake detection. NeurIPS Datasets and Benchmarks Track, 2021.
  5. Reimao and Tzerpos. FoR: A dataset for synthetic speech detection. SpeD, IEEE, 2019.
  6. The M-AILABS Speech Dataset. caito.de, 2019.
  7. Tak et al. End-to-end spectro-temporal graph attention networks for speaker verification anti-spoofing and speech deepfake detection (RawGAT-ST). arXiv:2107.12710, 2021.
  8. Kawa et al. Improved DeepFake Detection Using Whisper Features. INTERSPEECH, 2023.
  9. Müller, Sperl, and Böttinger. Complex-valued neural networks for voice anti-spoofing. INTERSPEECH, 2023.
  10. Müller et al. Speech is silver, silence is golden: What do ASVspoof-trained models really learn? arXiv:2106.12914, 2021.
  11. Ardila et al. Common Voice: A massively-multilingual speech corpus. LREC, 2020.

FAQs

Why does audio deepfake detection need a multi-language dataset? Most existing training data is overwhelmingly English or Chinese, so detectors trained on it tend to perform worse on other languages. MLAAD extends real speech into 54 languages using 140 different voice-cloning systems, giving researchers a way to train and test detectors that generalize beyond English.

Does more training data automatically mean a better detector? Not necessarily. MLAAD and the industry-standard ASVspoof19 dataset each produced the best detector on four of eight test datasets in this study, meaning the two datasets teach detectors different, complementary things rather than one simply being "more."

Try Resemble AI free
Generate with confidence. Verify ownership. Detect deception. Only with Resemble AI.
Get started
Generate and verify assets. Detect deception.
Start building now with a free account. Full API access. No credit card required.