Where Words
Become Sound

A local AI audio studio for music, speech, editing, and automation.

Demodokos Foundry is a Windows desktop app that lets you generate music and speech, separate stems, patch sections, mix on a timeline, and export finished audio - all on your own machine.

Start Trial $0 today via PayPal
Download Foundry Windows installer
Create Anything Unlimited & offline

100% Local Generation Windows 6GB+ VRAM Create without cloud credits 50 music / 10 speech languages

Signed Installer Defender Verified Generation runs locally

Hear What Foundry Creates

Music, voices, audiobooks, narration - all from a simple text description.

Cozy Night Lounge
Music
Desert Night Jazz Jazz
Music
Focus Study Music
Music
Forgotten Metal
Music
Less of You Rock
Music
Neon Reverie Synthwave
Music
Noche De Fuego Reggaeton
Music
Silence of the Night Ambient
Music
Storm Trance
Music
The Cold Side of the Bed Singer-Songwriter
Music
Winter Soliloquy Classical
Music
Investigator Crime Audiobook, Two Speakers
Speech
Barnaby Bear Kids Story, Three Voices
Speech
Anonymous Caller Distorted Voice, Blackmail Threat
Speech
History Podcast Male Narrator, Background Score
Speech
Guided Meditation Soothing Female, Mindfulness
Speech
Fantasy Audiobook Epic Fantasy, Book Intro
Speech
Learning Colors Educational Kids Podcast
Speech
Advertisement Woman Narrator, Upbeat Music
Speech

Everything you heard above was created inside Foundry - from songs to narration to multi-voice scenes.

Speech generation in 10 languages

Open the dedicated speech language page directly from the flags below, or hit play to hear a real sample first.

Speech generation supports 10 languages. Music and lyrics support 50+ languages inside the same product.

How it works

It reads the whole thought before it ever makes a sound.

Most voice tools read words one at a time. Demodokos Foundry v4 is built on acoustic-token synthesis: a language transformer takes in the whole sentence first, modelling context, intent and delivery instead of pronouncing words in isolation. The voice is then generated from fine-grained acoustic tokens, roughly one decision every 40-80 milliseconds, which is what gives it control over timing, emphasis, rhythm, emotion and pacing down to the smallest beat.

Precomputed style embedding hybrid global style embedding
Spectral identity control speaker-similarity monitor + correction
1Text context words, intent, punctuation, instructions
2Acoustic token generation dual-track LM over discrete speech tokens
3Speaker identity conditioning speaker embedding / timbre prior
4Style / emotion conditioning prosody, pacing, affect
5Waveform decode codec / code2wav synthesis
6Final speech voice acting synthesis, natural and expressive

The v4 speech model builds on a substantially reworked Qwen3-TTS foundation, delivering stable long-form synthesis, fine-grained expressive control, and consistent speaker identity.

A full hour of finished speech in under five minutes on an RTX 5090. And all of it consistent in the correct emotions and voices.

It stays itself

Each cloned or synthetic voice is grounded in learned neural representations that capture identity, timbre, accent and vocal texture, with no hundreds of hours of fine-tuning. As it generates, identity-preserving latent conditioning and spectrogram-based consistency control keep every line locked to the same speaker.

It feels every line

On top of that identity sits an independent performance layer: more than 40 emotions and speaking styles, each in five intensity levels, on generated and cloned voices alike. Emotion, rhythm, pacing, accent and texture can change completely while the person underneath never does.

Why Foundry Changes the Game

Most cloud tools make you use separate products for music, voice, editing, and automation. Foundry brings it all together in one local studio.

4 Studios in 1

Music creation, expressive voice acting synthesis, real mixing/mastering, and pro automation.

Create Without Credit Anxiety

No cloud credits burning through your budget. Your hardware, your rules, your pace.

Voices with Range and Identity

Voice acting performance that is stable, and recognizable across styles and scenes.

Built for Speed

Generate one hour of speech in as little as 4 minutes. Create entire music tracks in under a minute.

Real Production Depth

Timeline editing, stem separation, patching, cover/extend workflows, and pro DSP effects.

A Creative AI Agent

Helps analyze, segment, refine, direct, narrate, and create.

Speech with Direction

30+ emotions and styles in multiple intensities, and consistent voice identity.

Batch Workflows & Agentic Control

Batch workflows, agentic control, and serious automation.

This is not another generator. It is a complete local AI audio production environment.

From blank page to finished audio in 5 simple steps

No studio booking. No per-character fees. The entire pipeline, voice and music, runs on your Windows machine in under a minute per page.

First time here? Watch the 4-minute install and first-launch tutorial before you start.
1/ 5
Step 1

Pick or build a voice

Choose from built-in voice presets, clone yourself or a reference sample, or generate a brand-new voice. Cloned voices stay on your disk, they never touch a server.

2/ 5
Step 2

Paste your script, cast every speaker

Drop in a chapter, a full book, a video outline, or a dialogue file. The redesigned script editor splits it into lines the moment you paste, so you can hand each character its own voice and direct a whole cast, not just a single narrator.

3/ 5
Step 3

Direct emotion line by line

Tag any paragraph as calm, excited, whispered, angry, sarcastic, or anything in between, with 5 intensity levels per emotion. The voice stays the same character; the feeling changes.

4/ 5
Step 4

Generate music in any style, in 50 languages

Original scores, ambient loops, full songs with vocals, instrumental beds. Any genre, any mood, sung in any of 50 languages. Score your narration or write a standalone track, then drag it straight onto your timeline. No second tool, no separate subscription.

5/ 5
Step 5

Sculpt any line with DSP and voice effects

Drop studio-grade effects on a single line or the whole track. Turn a narrator into an alien, a demon, or a vintage radio broadcast in one click, then reach for reverb, EQ, auto-tune, formant shifting and glitch, all non-destructive. A full DSP rack, built in, no plugins and no external DAW.

Music, Voice & Everything In Between

Generate music in 50 languages and speech in 10.

Type It. Hear It.

Describe a song, a voice-over, or a narration. Foundry generates complete audio with vocals, instruments, or spoken word, ready in seconds, entirely on your machine.

Generate music or speech from one prompt

Caption Builder

Type what you imagine: music or speech. Adjust mood, tempo, voice style, press Generate. A full track or spoken read, done in seconds.

Audiobooks & Podcasts

Turn scripts, chapters, and show segments into polished spoken audio. Multi-speaker scenes with distinct voices, emotional direction, and 15× real-time generation speed.

Turn scripts into multi-speaker narration

Script to Speech

Assign voices to roles, steer emotion per line, and produce full audiobook chapters or podcast episodes without a recording session.

Voice & Music, One Timeline

Layer narration over original scores, blend dialogue with sound design, and mix spoken word with music beds, all in the same editor. Build trailers, immersive stories, and rich audio productions without switching apps.

Layer voice, music, and effects on one timeline

Unified Timeline

Drag voice, music, and effects onto the same visual tracks. The spectrum analyzer identifies BPM, time signature, and key. Mix everything, then export your finished production.

Voices That Actually Feel

A narrator who breaks with sadness at just the right moment. A villain whose calm whisper turns to fury. A podcast host who sounds genuinely thrilled. With 40 emotions and 5 intensity levels, every cloned or preset voice stays perfectly in character.

Direct emotion line by line

Emotion Engine

Pick any emotion, from whisper and rage to heartbreak and storytelling, then dial the intensity. Assign different moods per line or paragraph. 60 speaker presets, voice cloning, and it all sounds like the same person. Not a robot.

Separate the Instruments

Take any song apart: vocals, drums, guitar, piano. Each clean on its own track. Use Karaoke mode and mix in your own vocals or an AI-generated track.

Split any song into separate instruments

Stem Separation

Splits any song into up to 7 separate tracks. Each instrument and voice gets its own channel.

Fix One Part, Keep the Rest

Something sounds off? Select just that area, generate a patch in seconds, and DSP-blend it seamlessly. Everything else stays exactly as it is.

Patch blending workflow

Patch & Blend

Select any region and regenerate just that section. Before and after stay untouched. Spectral blending ensures seamless boundaries.

Cover, Extend & Transform

Feed Foundry any audio and restyle it completely, or extend a 30-second idea into a full song. Same structure, with an entirely new character or a seamless continuation.

Cover and extend workflow

Cover & Extend Modes

Load audio as a foundation and apply a new style with Cover mode. Or select any part and press Extend to continue naturally from where it left off.

32 Studio-Grade Effects

EQ, reverb, chorus, tape warmth, voice transformation, and glitch, with 200+ presets across 7 groups. Stack freely on your timeline, non-destructive, one click.

DSP rack overview

Post-Processing Suite

From surgical 24-band EQ to cathedral reverb, from drum punch to spectral crossfading. Auto-Tune, voice effects, glitch, and granular stretch, all non-destructive.

Automate Everything

120+ commands to generate, compose, separate stems, mix, and export. Build pipelines, batch-produce content, or let an AI agent drive Foundry for you.

Automation console

Automation / API

120+ commands. Create, compose with Creative AI, manage tracks, split stems, export. Build anything.

Unlimited Generation. Voice, Music, Everything.

Start with a 7-day free trial, then pick the plan that fits. Everything runs on your computer - no cloud, no queues, no per-song limits.

Most Popular

Creator

For your own projects - YouTube, games, podcasts, personal work

$12.00 per month billed annually
  • Unlimited music and voice generation
  • Generate full songs, extend them, remix existing audio, fix any section
  • Up to 3.5 min per music generation
  • 8 tracks per project
  • Separate vocals, drums, bass and more from any track
  • Narration (~13 min/script, unlimited)
  • Up to 5 speakers per script
  • Create custom voices or clone your own
  • AI assistant - describe what you need, it handles the settings
  • Commercial use for your own projects
  • 16 DSP effects with dozens of presets
  • No API / automation

For 1 person - use it for your own content, not for client work

Professional

Do work for clients - freelance, agency, studio

$39.20 per month billed annually
  • Everything in Creator, plus:
  • Up to 10 min per music generation
  • 32 tracks per project
  • Narration up to ~7 hour/script
  • Up to 16 speakers per script - full cast productions
  • API and command-line automation
  • Batch processing - Up to 6 parallel batches
  • Lossless FLAC export
  • Integrate into your pipeline
  • 33 DSP effects with 200+ presets

1 seat per person - produce work for your clients

Enterprise

For teams of 5 or more

$3,000+ per year
  • Multiple seats for your whole team
  • Full automation and API access
  • Dedicated support channel
  • Custom licensing and compliance terms
  • Centralized deployment across workstations

Required for 5+ seats, 10+ employees, or $2M+ revenue

Everything runs on your computer. Private, fast, no cloud queues.

Cancel anytime, no further charges. Start with the free trial - full access, pay only if you love it.

What You'll Need

Foundry runs on Windows GPUs, with full support on NVIDIA and limited support on AMD. Here's a quick overview.

Operating System

Windows 10 or 11, 64-bit
Optimized for Windows workstations

SSD/NVME disk recommended

Graphics Card

NVIDIA GPU with 4 GB+ VRAM
any RTX series card, incl. GTX 1080, 12 GB+ recommended

AMD GPU with 8 GB+ VRAM
Full Music and Speech quality since v2.1

VRAM Tiers

up to 6 GB
reduced quality
6 to 8 GB
reduced performance
16 GB+
+ multilingual Creative AI
24 GB+
+ brilliant Creative AI

Good to Know

Setup: 70 MB
Latest app version: v2.2.33
First-run model pack: ~20 to 25 GB

Stop Scrolling. Start Creating.

Download Demodokos Foundry setup here.

Download for Windows Version 2.2.33 · Windows 10/11 · NVIDIA 6 GB+ · AMD partially supported
Windows Defender Verified Digitally Signed
SHA-256: A5589A6DAF1B879BC773D62F0E2A19E2D4F3DE369AAF88FE319DF0DEF525CC95

Windows release available now

Demodokos Foundry v2.2

A new, significantly improved Speech Editor.

Foundry 2.2 rebuilds narration around rich documents, with overlapping speech, crosstalk, background music, precise timing, and a faster workflow.

Entirely rebuilt Speech Editor with rich-document narration
Overlapping speech, crosstalk, music, effects, and custom delays
What you preview is exactly what you hear in playback and exports
Keep editing and saving your work even when it exceeds plan limits

Frequently Asked Questions

Does Demodokos Foundry work offline?

Yes. All AI generation, voice cloning and audio processing happen locally on your GPU. Your prompts, scripts, voice samples and generated audio are not uploaded to a cloud service. An internet connection is still required for login and license verification. Foundry may sign you out if it cannot reach the authentication server for an extended period.

Does Demodokos Foundry upload my audio files to the cloud?

No. Your audio files, voice samples, prompts, scripts, project data and generated content stay on your computer. Foundry does not upload creative content to a cloud AI service for generation or processing.

What GPU do I need to run Demodokos Foundry?

An NVIDIA GPU with at least 4 GB of VRAM can run selected Speech models at reduced quality. 6 GB is a more practical starting point, while 12 GB or more is recommended for the best overall Music and Speech performance. AMD GPUs with at least 8 GB of VRAM can run both Music and Speech through Vulkan, although AMD support remains experimental. NVIDIA with CUDA is recommended.

I have an AMD GPU. Will that work?

Yes. AMD GPUs with at least 8 GB of VRAM can run both Music and Speech through Vulkan. AMD support remains experimental and may be less consistent than NVIDIA/CUDA, which provides the most mature Foundry experience.

Does Demodokos Foundry run on macOS or Linux?

Foundry is currently available for 64-bit Windows 10 and Windows 11. Native macOS and Linux versions are not currently available.

How much does Demodokos Foundry cost?

The Creator plan is $15.00/month and the Professional plan is $49.00/month. Discounted annual billing is also available. Both plans include unlimited local AI music and voice generation without per-generation credits; plan limits apply to project size, duration and advanced features. A free 7-day trial is included.

Is there a free trial?

Yes. Demodokos Foundry offers a 7-day free trial with the full capabilities of your selected license. It is processed through PayPal with a $0 authorization, and you are not charged unless you keep the subscription after the trial ends. Cancel any time before the trial expires.

Can I cancel my subscription anytime?

Yes, any time, under Billing > Cancel. Cancelling stops your next renewal and takes effect at the end of your current billing period - the current month for monthly plans or the current year for annual plans. You keep full access until then. Payments already made are not refunded, and there are no cancellation fees. During a free trial, you can cancel any time with no charge.

Can I clone my own voice with Demodokos Foundry?

Yes. Import a short recording of your voice (or another voice you are authorized to use) and Foundry creates the cloned voice locally on your machine. Cloned voices can speak all 10 languages and use more than 40 emotions and speaking styles, each with five intensity levels, without per-character or per-generation charges.

What languages does Demodokos Foundry support?

Demodokos Foundry supports music generation and lyrics in 50 languages, and speech and voice generation in 10 languages. The Creative AI agent is optimized for English conversation. Larger Creative AI models available on higher-VRAM systems understand additional languages and provide stronger multilingual lyric writing, text analysis and narration support.

Does Demodokos Foundry use open-source or proprietary AI models?

Demodokos Foundry uses proprietary AI systems and adapted models developed from open-source foundations. Demodokos extensively modifies, extends and further adapts those foundations before integrating them into its proprietary music, speech and audio-production architecture.

The Music v4 engine builds on a modified ACE-Step 1.5 foundation, with spectral-flux-guided stabilization integrated directly into the generation process. Speech v4 advances Qwen3-TTS through further adaptation and extensive architectural changes, while Creative AI uses constrained Qwen3 language models within a custom agentic pipeline. Demodokos has also re-engineered Metas AudioSeal technology into Tonotope, a lightweight AI-origin marking system embedded directly into generated audio while remaining inaudible.

All AI inference runs locally through a custom C++ inference stack built around GGML and ONNX.

Is Demodokos Foundry just a wrapper around open-source AI models?

No. A wrapper typically places a new interface over an existing model while leaving the underlying technology and workflow largely unchanged. Foundry goes much further.

Demodokos modifies and extends the models themselves, then builds proprietary generation, voice-consistency, orchestration and audio-processing systems around them. Music v4 and Speech v4 include changes that directly affect musical stability, speaker identity, expression and long-form consistency.

The difference is especially visible in the rich Speech Editor. It combines a familiar document-writing experience with tools built specifically for spoken production: visual speaker and delivery cues, background music and sound inserts, overlapping dialogue, natural crosstalk, sample-accurate timing and visually adjustable DSP effects. Every sentence can be previewed, edited or regenerated in context, with individual control over pace, pitch, volume and delivery.

Foundry´s agentic pipelines can process entire books, split them into chapters, summarize and analyze their content, inspect images and OCR text, identify speakers, create voices from AI-generated character descriptions and prepare narration. A separate narration pipeline can adapt text, choose the right speaker and select fitting emotions line by line.

Projects can then move into music generation, stem separation, section repair, timeline arrangement, mixing and final export. Foundry is not a front end for third-party models; it is an integrated local audio-production environment built to take long-form text from document to finished sound.

How is Demodokos Foundry different from ElevenLabs?

ElevenLabs is a cloud platform built around monthly usage credits. Demodokos Foundry is a local Windows audio-production environment: generation runs on your own GPU, your scripts, voices and audio stay on your machine, and output is not metered by characters or per-generation credits.

Foundry is also designed around complete productions rather than isolated generations. Its rich Speech Editor lets you write in a familiar document interface, assign speakers, direct emotion and delivery line by line, insert music and sound, build overlapping dialogue and crosstalk, regenerate any sentence in context, and adjust pace, pitch, volume, timing and DSP without leaving the document. Agentic workflows can prepare whole books by segmenting chapters, analyzing text and images, identifying speakers, designing voices and directing narration.

Music generation, stem separation, section repair, timeline mixing and automation are built into the same desktop application. The key difference is a private, integrated local studio instead of a cloud service whose usage is measured in credits.

Is Demodokos Foundry safe for business and GDPR compliant?

Yes. Demodokos Foundry is designed to support GDPR-compliant workflows. All AI generation, voice cloning and audio processing run on your own hardware, so business content, client audio, proprietary voice recordings, internal scripts and confidential narration stay on your premises instead of being sent to a cloud AI provider. No cloud AI provider receives or retains your creative content. This gives your organization direct control and data sovereignty over the generation workflow, making Foundry especially well suited for businesses, legal teams, healthcare, media agencies and other privacy-sensitive work.

Is Demodokos Foundry AI Act compliant?

Yes. Demodokos Foundry is designed to meet the applicable transparency requirements of the EU AI Act. Generated audio is identified as AI-generated in its metadata and carries Tonotope, Demodokos´ robust, inaudible AI-origin watermark.