In 2019, my co-founder Saqib and I started Resemble AI to fix a problem we kept hitting in gaming: voice work couldn't keep pace with how fast studios shipped, and human actors didn't scale. The answer was generative voice, and building it well meant going deep, into the architecture, the training data, and every layer of how synthetic speech actually gets produced.
That turned out to be the most valuable thing we ever did, but not for the original reason. Understanding how voice AI is built at the model level gives us a major advantage at identifying clones and gen AI voices in the wild. Detection was also in our minds early on - we had a sneaking suspicion that the better the voices got, the more they were likely to be used in malicious ways, with bad intent or with positive intent without knowing the origin. So we researched, built and open sourced:
- DramaBox, a highly expressive voice model built on a novel architecture, with watermarking built-in by default, as an example of TTS with native marking
- PerTh, our audio-only watermarking model, so any piece of audio (including the audio our models generated) could be traced to its source
- Resemblyzer, our speaker verification and matching model, so any piece of audio can be matched for an enrolled identity
And, unfortunately, we started to see deepfake incidents and corresponding harm grow. Our H1 2026 Deepfake Threat Report found at least 15,736 people victimized across 821 documented attacks in six months, one in six of them involving non-consensual intimate imagery of adults or children. The files behind those attacks numbered around 3.46 million, and the majority of them were image and video. The detection work we'd started as a companion to voice had become the thing the world most needed from us, and for every modality, including video, and image.
Which is why we decided to commit to producing the best deepfake detection models possible. We're not selling voice AI to new customers, we have put the force of the whole company behind detection, for audio, video, and image as well as watermarking multimodal content, multimodal identity matching, fraud signaling and human-readable explainability.
What this means for the voice work
The voice models we built are open and free to use, but we won't be producing new ones for commercial sale.
And here’s how the five core APIs work and what questions they answer:
We've also retired more than 300 older posts on voice cloning, text to speech, and audio editing. Most redirect here. If you followed one to this page, that's probably why.
We’ll keep researching voice in addition to image and video
Everything we learned about voice AI is why our detection model is the most accurate on the market. Our open source models have been downloaded more than 17 million times on Hugging Face which proves we can build foundational models that hold up under real scrutiny, in public, where other researchers test them. That same capability is the foundation under DETECT-World, today the most accurate audio deepfake detector in the world by public benchmarks at 99.5% detection accuracy. You can follow our experiments and advancements as we continue to publish datasets, research and content on our work.
We started in audio because that's where our expertise ran deepest but the threat didn't stay there. Deepfakes now cross every format, so we carried the same insight into video and image, building multimodal detection on the understanding that catching a fake in any medium is well served by knowing how it’s built.
We're laser focused on one problem now, the one we happen to be best in the world at. It's the same company it was in 2019, built on the same insight, we've just chosen which half of it the world needs most.
Deepfakes are everywhere. So are we.




