Voice Cloning

elevenlabs logo
In this guide:

Voice cloning lets you create a digital replica of any voice and use it to generate realistic speech from text. ElevenLabs offers two ways to do this: Instant Voice Cloning (IVC) for quick results from a short sample, and Professional Voice Cloning (PVC) for a hyper-realistic, custom-trained model. This guide walks you through the entire process — from recording clean audio to using your cloned voice in the dashboard.

Prerequisites

  • An ElevenLabs account (free tier works for Instant Voice Cloning; Creator+ plan required for Professional Voice Cloning)
  • A clear audio recording of the voice you want to clone (1–3 minutes recommended)
  • Legal right and consent to clone the voice in question
  • Optional for best results: a condenser microphone, pop filter, and a quiet recording space

Step 1: Understand Your Options — IVC vs PVC

Before you start, pick the right cloning method for your needs:

 

Instant Voice Cloning (IVC) works from a short audio sample and produces results near-instantly. It does not train a new AI model — instead, it uses ElevenLabs’ existing training data to approximate the voice. This works great for most voices, but may struggle with very unique accents or unusual vocal styles the model hasn’t encountered before. Available on all plans.

ElevenLabs-instant-voice-cloning

 

Professional Voice Cloning (PVC) trains a dedicated AI model specifically on your voice data, producing results that are near-indistinguishable from the original. It takes longer — roughly 3 hours for English, ~6 hours for multilingual — and requires a Creator+ plan. Use this when accuracy and realism are non-negotiable.

ElevenLabs-professional-voice-cloning

💡 Start with IVC to test the waters. Switch to PVC only if the results aren’t good enough for your use case.

Step 2: Record High-Quality Audio

The quality of your clone is almost entirely determined by the quality of your input audio. The AI will replicate everything it hears — including background noise, mouth clicks, uneven pacing, and breathing artifacts. Here’s how to get it right:

Location: Record in the quietest, most acoustically “dead” space you can find. A closet full of clothes works surprisingly well as a DIY vocal booth. Avoid rooms with hard walls and high ceilings.

Microphone: A professional XLR microphone in the $150–$300 range is sufficient for most use cases. A popular affordable setup is a Focusrite interface paired with an Audio-Technica AT2020 or Rode NT1. Always use a pop filter to prevent plosive sounds.

DAW: Use any recording software that can export WAV files at 44.1kHz or 48kHz, 24-bit. REAPER is the industry standard; Audacity is a solid free option. No post-processing, noise reduction, or effects — record raw and clean.

Positioning: Keep your mouth approximately 20cm (about two fists) from the microphone, with the pop filter in between. Speak at a slight angle to the mic to reduce direct breath impact.

Levels: Aim for peaks of -6 dB to -3 dB and an average loudness of -18 dB. For IVC specifically, the ideal RMS is between -23 dB and -18 dB with a true peak of -3 dB.

Performance: Be consistent. If you’re recording an upbeat, expressive voice — keep it that way throughout. Mixing a flat monotone with an animated delivery confuses the model and leads to unstable output. Avoid filler words, stutters, and deep audible breaths unless those are genuinely part of the voice you’re cloning.

Step 3: Prepare Your Audio File

Once you’ve finished recording, prep your files before uploading:

  • Length: Aim for 1–2 minutes of usable audio. You can upload multiple shorter clips — ElevenLabs uses the total combined runtime, not the number of files. Don’t go beyond 3 minutes; more audio rarely improves quality and can sometimes hurt the clone.
  • Format: WAV is ideal, but MP3 at 128 kbps or higher also works fine. Higher bitrates beyond that won’t meaningfully improve your results.
  • Clean it up: Trim silence at the start and end. Remove any sections with background noise, coughing, or technical artifacts. Do not apply noise reduction or EQ — keep the recording natural.
  • Consistency check: Listen back through and make sure the tone, pace, and energy are consistent throughout. Wide fluctuations in pitch or volume will produce unpredictable cloning results.

Step 4: Create Your Voice Clone in the Dashboard

With your audio ready, here’s how to create the clone:

  1. Log in to your ElevenLabs account and navigate to the Voices section in the left sidebar.
  2. Click “Add a new voice”.
  3. From the modal, select “Instant Voice Clone” (or “Professional Voice Clone” if you’re on a Creator+ plan and want PVC).
  4. Upload your audio file(s) or use the in-browser recording option to record directly.
  5. Give your voice clone a name and label. This helps you identify it later, especially if you’re managing multiple clones.
  6. Confirm that you have the right and consent to clone the voice, then click “Save voice”.

For IVC, the clone will be ready almost immediately. For PVC, expect a wait of approximately 3 hours (English) or 6 hours (multilingual) depending on queue times.

Step 5: Use Your Cloned Voice

Once your clone is ready, go back to the Voices section in the dashboard and open the “Personal” tab. Your new voice clone will appear there. Click on it to select it and start generating speech.

From here you can use your cloned voice in the Text to Speech playground, the Studio product, or via the ElevenLabs API by referencing the voice ID assigned to your clone. The voice ID is visible in the voice details panel — copy it for use in any API-based workflow.

💡 The cloned voice will replicate the performance style of your original recording. If you recorded at a calm, measured pace, that’s how the generated speech will sound. Plan your recording with the end use case in mind.

Common Mistakes

Mistake Why It’s a Problem & What to Do Instead
Recording in a reverberant room Reverb gets cloned too, making generated speech sound hollow. Use a closet, blanket tent, or acoustically treated space.
Uploading more than 3 minutes of audio Extra length rarely helps and can hurt quality. 1–2 minutes of clean audio is the sweet spot.
Inconsistent tone or energy in the recording The AI can’t tell which version of the voice to clone, leading to unstable output. Keep performance consistent throughout.
Applying noise reduction or EQ before uploading Processing can introduce artifacts that confuse the model. Upload the raw, unprocessed recording.
Expecting IVC to perfectly clone a unique accent IVC relies on pre-existing training data. Highly unusual accents may not clone well — use PVC (Creator+ plan) for those cases.
Cloning a voice without consent ElevenLabs requires you to confirm consent before saving a clone. Always ensure you have the legal right to clone the voice you’re uploading.

Quick Recap

  • ElevenLabs offers two cloning methods: IVC (fast, any plan) and PVC (hyper-realistic, Creator+ only)
  • Audio quality is everything — record in a quiet space, use a decent mic, and skip all post-processing
  • Aim for 1–2 minutes of clean, consistent audio; don’t exceed 3 minutes
  • Keep your performance consistent in tone, pace, and energy throughout the entire recording
  • In the dashboard: Voices → Add a new voice → Instant or Professional Voice Clone → Upload → Save
  • Find your cloned voice under the Personal tab in the Voices section
  • Use the voice ID from the dashboard to call your clone via the ElevenLabs API
  • Always have legal consent before cloning any voice

Quick FAQ

Marco-profile-pic

Written by

Marco Sansalone

Founder of AI Tool Curator. UX/UI Designer & strategist with 20+ years in the design field.
elevenlabs

Want to know more?

View ElevenLabs

Also in ElevenLabs Guide

Something Not Working? Tell Us What’s Wrong.

Popular requests move to the top of our queue.
You'll be notified when the tutorial goes live and join our newsletter on AI tools and tutorials.

Find Your Perfect AI Tool