Foundry Is Now a Music and Speech Studio

Features Speech AI Music

When Foundry launched, it was a music generation studio. You typed a description, it produced a full song with vocals and instruments, and then you could edit, layer, and mix the result. That part is still there and stronger than ever.

What changed: Foundry now generates speech too.

What speech generation looks like in practice

You write or paste a complete script into the rich Speech Editor. You assign voices to roles and direct the performance where it changes. Speakers, delivery styles, music, effects, delays, overlaps, and crosstalk all stay inside the same readable document.

Foundry has 38 delivery styles across emotion, drama, narration, and character states, each with 5 intensity levels. You can direct joy, anger, heroic calm, documentary narration, exhaustion, whispering, singing, and much more. The voice stays consistent across the entire read without flattening every performance into the same neutral tone.

60 built-in speaker presets cover a wide range of vocal types. If none of them match what you need, voice cloning lets you create a custom voice from a short audio sample. That cloned voice then works with the same delivery controls as everything else.

Multi-speaker scenes

Audiobooks have multiple characters. Podcasts have co-hosts. Dialogues need distinct voices that interact naturally. Foundry handles all of this in one document. Speakers can alternate, overlap, interrupt each other, or speak in natural crosstalk while each voice stays separate and recognizable.

A full audiobook chapter with three characters, each with their own emotional arc, generates at up to 15x real-time speed. On a 12 GB NVIDIA card, that means a 10-minute scene is ready in under a minute.

Music and speech in the same narration document

This is where the combination gets interesting. You can place background music, effects, and exact timing directly around the narration. Preview, Play All, and export use the same sample-accurate arrangement, so the finished result matches what you heard while editing.

When a production needs deeper timeline work, you can still send a speaker, paragraph, or complete document to Tracks. Normal narration no longer requires manually arranging spoken clips there.

Podcast intros with custom music beds. Game trailers with character dialogue over an original score. Audio dramas with layered sound design. All built inside one application, all running on your hardware.

Everything else is still here

Text-to-music generation with the caption builder. Patch editing to fix sections without starting over. Stem separation into 7 tracks. Cover and restyle workflows. 32 base DSP effects with hundreds of presets, sub-effects, and audio modifications that can be combined. The Creative AI for writing assistance. The full timeline editor with multi-track mixing.

Speech did not replace any of that. It just made Foundry into something broader: a complete local audio production environment for both music and voice.

More from Echoes

AI Dubbing Software That Runs on Your Own PC

AI dubbing software that runs on your Windows PC. Translate and re-voice video in 10 languages in the original voice, with no minute limits or uploads. DubbingVideo Translation 10 min

Demodokos Foundry 2.0: Speech v4, the Biggest Speech Update Since Launch

Demodokos Foundry 2.0 introduced the Speech v4 model generation, better AMD/Vulkan support, a 4GB VRAM mode, and a rebuilt mixer. All local, still $12.00/month when billed annually. Product UpdateAI Voice 7 min

Why AI Voices Lose Emotion in Long Audio (And the Fix)

AI voices drift from warm to flat over long audio. Here is why delivery consistency breaks across audiobooks and how Foundry keeps a voice steady inside one rich narration document. AI VoiceTTS 8 min

What GPU Do You Need for Local AI Audio?

Local AI audio needs the right GPU. Here's exactly how much VRAM you need for voice cloning, music generation, and TTS in June 2026, with specific card picks at every budget. Local AIHardware 9 min

You Run LLMs Locally. You Generate Images Locally. Why Is Your Audio Still in the Cloud?

You went local for text and images. But every time you need a voiceover, a soundtrack, or a sound effect, you are back in a browser uploading files to someone else's GPU. Here is why local AI audio deserves a spot in your stack. Local AIPrivacy 9 min

The Best ElevenLabs Alternatives in 2026 (Especially If You're Tired of the Bill)

Looking for ElevenLabs alternatives in 2026? We compare the top AI voice generators by price, privacy, and features, including one that runs entirely on your own computer. ElevenLabsComparison 5 min

How to Pick a TTS Tool for Production Use (Not Just Demos)

Every TTS tool sounds good on a demo. This is the version for people who actually need to ship something — covering consistency, per-character pricing at scale, API reliability, and when cloud vs. local is the right answer. TTSVoice Production 5 min