Back
Blog
Aug 13, 2026

Why a former voice AI company went all in on deepfake detection

CONTENTS
Active heading
Section heading
CONTRIBUTORS
Zohaib Ahmed
Co-Founder and CEO

In 2019, my co-founder Saqib and I started Resemble AI to fix a problem we kept hitting in gaming: voice work couldn't keep pace with how fast studios shipped, and human actors didn't scale. The answer was generative voice, and building it well meant going deep, into the architecture, the training data, and every layer of how synthetic speech actually gets produced.

That turned out to be the most valuable thing we ever did, but not for the original reason. Understanding how voice AI is built at the model level gives us a major advantage at identifying clones and gen AI voices in the wild. Detection was also in our minds early on - we had a sneaking suspicion that the better the voices got, the more they were likely to be used in malicious ways, with bad intent or with positive intent without knowing the origin. So we researched, built and open sourced:

  • DramaBox, a highly expressive voice model built on a novel architecture, with watermarking built-in by default, as an example of TTS with native marking
  • PerTh, our audio-only watermarking model, so any piece of audio (including the audio our models generated) could be traced to its source
  • Resemblyzer, our speaker verification and matching model, so any piece of audio can be matched for an enrolled identity

And, unfortunately, we started to see deepfake incidents and corresponding harm grow. Our H1 2026 Deepfake Threat Report found at least 15,736 people victimized across 821 documented attacks in six months, one in six of them involving non-consensual intimate imagery of adults or children. The files behind those attacks numbered around 3.46 million, and the majority of them were image and video. The detection work we'd started as a companion to voice had become the thing the world most needed from us, and for every modality, including video, and image.

Which is why we decided to commit to producing the best deepfake detection models possible. We're not selling voice AI to new customers, we have put the force of the whole company behind detection, for audio, video, and image as well as watermarking multimodal content, multimodal identity matching, fraud signaling and human-readable explainability.

What this means for the voice work

The voice models we built are open and free to use, but we won't be producing new ones for commercial sale.

If you came for What to know
Chatterbox, DramaBox, and our generative voice models They live on under Voice AI Research, still open, still free to build on.
Voice generation as a paid product We're supporting our existing customers but not taking new ones.
Deepfake detection Our core mission, in the form of five core detection APIs that create a complete enterprise trust stack.

And here’s how the five core APIs work and what questions they answer:

API The question it answers What it does
Detect Is this real? A verdict on audio, video, and image in under 300 milliseconds. Runs on DETECT-World, which checks whether media is physically consistent rather than matching signatures from generators it has already seen, so it holds up against models released after it was trained.
Watermarker Where did this come from? Embeds an imperceptible watermark at the point of generation and reads it back after compression, editing, and re-encoding. C2PA compatible, and the mechanism behind EU AI Act Article 50 disclosure.
Identity Whose voice or face is this? Enrolls a known voice and likeness, then matches incoming media against the enrolled identity. Tells you whether the person on the call is the person you registered.
Signal Have we seen this attack before? Fraud pattern matching against preset and custom libraries, closer to Shazam than to a wiretap. No transcription, no stored content, no PII exposure.
Intelligence Why should anyone believe the verdict? Converts a deterministic score into plain-language reasoning and structured JSON, so an analyst, a regulator, or a court can act on the result instead of taking it on faith.

We've also retired more than 300 older posts on voice cloning, text to speech, and audio editing. Most redirect here. If you followed one to this page, that's probably why.

We’ll keep researching voice in addition to image and video

Everything we learned about voice AI is why our detection model is the most accurate on the market. Our open source models have been downloaded more than 17 million times on Hugging Face which proves we can build foundational models that hold up under real scrutiny, in public, where other researchers test them. That same capability is the foundation under DETECT-World, today the most accurate audio deepfake detector in the world by public benchmarks at 99.5% detection accuracy. You can follow our experiments and advancements as we continue to publish datasets, research and content on our work.

We started in audio because that's where our expertise ran deepest but the threat didn't stay there. Deepfakes now cross every format, so we carried the same insight into video and image, building multimodal detection on the understanding that catching a fake in any medium is well served by knowing how it’s built. 

We're laser focused on one problem now, the one we happen to be best in the world at. It's the same company it was in 2019, built on the same insight, we've just chosen which half of it the world needs most.

Deepfakes are everywhere. So are we.

Try Resemble AI free
Generate with confidence. Verify ownership. Detect deception. Only with Resemble AI.
Get started
Generate and verify assets. Detect deception.
Start building now with a free account. Full API access. No credit card required.