Voice cloning creates a private voice model from an authorized audio sample. For product teams, the value is consistency: the same approved voice can be used across demos, onboarding, explainers, training content, and product updates.
The most important rule is simple: clone only voices you have permission to use. A good workflow makes consent, sample quality, storage, and credit cost visible before generation begins.
Key takeaways
- Voice cloning should start with consent, not with audio.
- A short clean sample is better than a long noisy one.
- Product teams should create private voice models for repeatable brand or narrator voices.
- Miso One charges a fixed clone creation cost plus sample speech generation cost.
- A new account can test a short cloning workflow before moving to paid usage.
What is voice cloning?
Voice cloning is the process of creating a reusable AI voice model from a reference recording. The model does not need a full studio session. It needs a clean, representative sample that captures tone, pacing, pronunciation, and recording quality.
In Miso One, the voice cloning workflow is designed around a private model:
- Upload or record an authorized voice sample.
- Add sample text for the generated preview.
- Confirm consent.
- Create the clone.
- Generate a short speech sample.
- Manage the result in My Voice Models.
Consent comes first
Consent is not a checkbox for legal decoration. It is the foundation of a safe voice product.
Before creating a clone, answer these questions:
- Do you own the recording or have permission to use it?
- Does the speaker know the voice may be used for AI generation?
- Is the intended use clear?
- Is the voice model private, unlisted, or public?
- Who inside the team can generate speech with it?
If any answer is unclear, pause. Use a public voice model instead, or collect a fresh authorized recording.
What makes a good voice sample?
The best sample is short, clean, and representative. You do not need dramatic acting. You need useful signal.
Aim for:
- quiet room
- no background music
- no overlapping speakers
- natural speaking pace
- consistent microphone distance
- clear pronunciation
- a sample that reflects the tone you want later
Avoid:
- heavy reverb
- clipped audio
- phone calls with compression artifacts
- music under the voice
- copyrighted media clips
- speeches from people who did not grant permission
A practical sample script
For a first clone, record a simple script like this:
Hello everyone. I am recording a short voice sample for a private AI voice model. I will speak clearly, at a natural pace, with a steady tone and clean pronunciation.
This is enough to test whether the workflow works. After that, create a more tailored sample for your brand voice.
How credits work for cloning
Voice cloning has two cost parts in Miso One:
- a fixed cost to create the private voice model
- a text-to-speech cost for the generated sample speech
That separation is important. Creating a model is different from generating speech with that model. If a provider creation step fails, the clone creation cost should be refunded by the backend workflow. If sample speech fails, the sample generation cost should also be handled as a generation failure.
For a new user, the best first test is a short sample text. That keeps the full clone experience within the welcome credit budget.
When should a team create a private voice model?
Create a private voice model when you need a repeatable voice identity.
Good use cases:
- product walkthrough narrator
- internal training voice
- founder-approved update voice
- character voice for a game prototype
- consistent support audio for tutorials
- voice style for social video templates
Use a public voice instead when:
- you only need one short clip
- brand consistency does not matter
- you are still exploring tone
- you do not have consent to clone a specific speaker
Quality checklist before generating
Use this checklist before pressing create:
- The speaker has permissioned the use.
- The recording is clean.
- The audio is not too short or too long for the uploader.
- The sample text is within the current character limit.
- The account has enough credits for clone creation and preview speech.
- The model name is descriptive enough to find later.
How to organize cloned voices
Do not name models "test 1" or "final final." Use names that explain the use case:
- Product Demo Narrator
- Calm Support Voice
- Founder Update Voice
- Training Module Voice
- Podcast Intro Voice
Then review the model later in My Voice Models. Keep only voices that have a clear owner, use case, and consent trail.
FAQ
Can I clone any voice I find online?
No. You should clone only voices you own or have permission to use. If you do not have consent, use a public voice model or record a new authorized sample.
How long should the sample be?
Short and clean is better than long and noisy. Follow the uploader guidance in the product, and focus on clear pronunciation and natural pacing.
Why does voice cloning use more credits than basic TTS?
Cloning creates a private voice model before generating speech. That model creation step has different compute and storage considerations than a one-off text-to-speech generation.
Where do I find my cloned voices?
Use My Voice Models to manage private voice models and Generation History to review generated audio.
Next step
Open the Voice Cloning tab, record a clean authorized sample, and generate one short preview before using the voice in production.

