Voice Cloning and the Delivery Style Engine

Speech Deep Dive

Flat text-to-speech has been around for years. You paste text, you get audio, and it sounds like a GPS giving directions. Technically correct, emotionally dead. Foundry takes a different approach.

The delivery style system

Every voice in Foundry, whether it is a built-in preset or a cloned voice, supports 38 delivery styles with 5 intensity levels. The styles cover emotion, drama, narration, and character states. That includes familiar emotions such as happy, sad, and angry, dramatic direction such as heroic or villainous, narration such as storytelling or documentary, and states such as exhausted, breathless, whispering, drunken, or singing.

You can assign different delivery styles to different parts of your script. The narrator starts calm and measured. By the climax, the voice carries tension. In the resolution, it softens. This is not post-processing or pitch tricks. The generation model produces the performance directly in the audio.

Some examples of what this enables:

  • An audiobook narrator who genuinely sounds heartbroken during a character's loss
  • A podcast host whose excitement sounds real, not performed
  • A game villain whose cold calm shifts to rage at exactly the right moment
  • A children's story narrator who sounds warm and playful throughout

Voice cloning

The 60 built-in speaker presets cover a lot of ground. Male, female, young, old, deep, bright, gravelly, smooth. But sometimes you need a specific voice.

Voice cloning creates a new voice profile from a short audio sample. Record yourself, or use an existing recording. Foundry captures the vocal characteristics and builds a reusable voice that works with the full delivery style engine.

The cloned voice stays consistent. If you generate a 30-minute audiobook chapter with a cloned voice, it sounds like the same person from start to finish. Different delivery styles, different energy levels, but always recognizably the same voice.

Direction inside the rich document

Scripts rarely have one mood throughout. A conversation shifts. A story builds. An advertisement has energy peaks and quiet moments.

Foundry lets you direct a short sentence or an entire chapter inside one rich document. Delivery styles can change wherever the performance calls for it. Multiple speakers can overlap, interrupt each other, and use natural crosstalk. Music, effects, and custom delays sit in the same document, and the exact arrangement is shared by preview and export.

All of this runs locally

Voice cloning samples stay on your computer. The emotion model runs on your GPU. Nothing is uploaded to any server. For anyone working with client voices, sensitive scripts, or unreleased material, this matters. Your voice data never leaves your machine.

More from Echoes

AI Dubbing Software That Runs on Your Own PC

AI dubbing software that runs on your Windows PC. Translate and re-voice video in 10 languages in the original voice, with no minute limits or uploads. DubbingVideo Translation 10 min

Demodokos Foundry 2.0: Speech v4, the Biggest Speech Update Since Launch

Demodokos Foundry 2.0 introduced the Speech v4 model generation, better AMD/Vulkan support, a 4GB VRAM mode, and a rebuilt mixer. All local, still $12.00/month when billed annually. Product UpdateAI Voice 7 min

Why AI Voices Lose Emotion in Long Audio (And the Fix)

AI voices drift from warm to flat over long audio. Here is why delivery consistency breaks across audiobooks and how Foundry keeps a voice steady inside one rich narration document. AI VoiceTTS 8 min

What GPU Do You Need for Local AI Audio?

Local AI audio needs the right GPU. Here's exactly how much VRAM you need for voice cloning, music generation, and TTS in June 2026, with specific card picks at every budget. Local AIHardware 9 min

You Run LLMs Locally. You Generate Images Locally. Why Is Your Audio Still in the Cloud?

You went local for text and images. But every time you need a voiceover, a soundtrack, or a sound effect, you are back in a browser uploading files to someone else's GPU. Here is why local AI audio deserves a spot in your stack. Local AIPrivacy 9 min

The Best ElevenLabs Alternatives in 2026 (Especially If You're Tired of the Bill)

Looking for ElevenLabs alternatives in 2026? We compare the top AI voice generators by price, privacy, and features, including one that runs entirely on your own computer. ElevenLabsComparison 5 min

How to Pick a TTS Tool for Production Use (Not Just Demos)

Every TTS tool sounds good on a demo. This is the version for people who actually need to ship something — covering consistency, per-character pricing at scale, API reliability, and when cloud vs. local is the right answer. TTSVoice Production 5 min