Skip to main content

ElevenLabs Voice Cloning: A Hands-On VoiceLab Guide

Learn how to clone your voice with Instant and Professional Voice Cloning in ElevenLabs, from recording clean audio to generating speech via the Python SDK.
Aug 3, 2026  · 13 min read

Explore with AI

Open in ChatGPTOpen in ClaudeOpen in Perplexity

You’ve probably found a great text-to-speech voice before, only to realize full usage is locked behind a more expensive tier, you cannot make it sound like your brand, and definitely cannot call it your own. That’s usually where most people hit a wall with voice AI. ElevenLabs VoiceLab is where that changes.

In this guide, I’ll walk you through all three tools inside ElevenLabs VoiceLab: Voice Design, Instant Voice Cloning (IVC), and Professional Voice Cloning (PVC). We’ll go step by step, and I’ll show you exactly what to expect at each stage.

Before we start, here’s what you need:

  • A free ElevenLabs account is enough for Voice Design
  • Starter plan ($6/month) is required for Instant Voice Cloning
  • Creator plan ($22/month) is required for Professional Voice Cloning

By the end, you’ll have at least one custom voice saved in your library, ready to use in Speech Synthesis (Text to Speech), Studio, or even via API.

What is ElevenLabs VoiceLab?

At a high level, VoiceLab is ElevenLabs’ dedicated suite for creating and managing custom voices. It’s where you go when you want something original, not just something prebuilt.

Here’s a quick comparison of the three VoiceLab tools:

Tool

What It Does

Quality

Input Required

Time to Create

Plan Required

Best Use Case

Voice Design

Generates synthetic voices

Good

None

Instant

Free

Experimentation

IVC

Clones voice from short audio

Good

~1–5 min audio

Instant

Starter+

Prototypes

PVC

High-fidelity voice cloning

Excellent

30 min to 3+ hours

1–2 days

Creator+

Production

If you want to go deeper into using voices programmatically, I recommend our guide to the ElevenLabs API.

Working With the OpenAI API

Start your journey developing AI-powered applications with the OpenAI API.
Explore Course

VoiceLab vs. the Voice Library

It’s important not to confuse this with the Voice Library.

  • Voice Library: Browse and use pre-made voices
  • VoiceLab: Create and manage your own voices

Once you create a voice in VoiceLab, you can optionally publish it to the Voice Library for others to use. But by default, it stays private in your account.

Setting Up Your ElevenLabs Account

Getting started is straightforward. 

  1. Go to elevenlabs.io
  2. Click Sign Up
  3. Use Google or email
  4. Choose your initial platform. Go with Eleven Creative to focus on the voice creation tools.

Choose between ElevenCreative and ElevenAgents

  1. Verify your account

Once you’re in, navigate to the VoiceLab area:

  • Click Voices in the left sidebar.
  • Click + Create Voice

ElevenLabs VoiceLab Voices

You’ll see four options:

  • Voice Design
  • Instant Voice Cloning
  • Professional Voice Cloning
  • Voice Remixing

Create voice options

Note that not all these features are available for all tiers. Check the table below to know what is available at each tier:

Plan

Voice Design

IVC

PVC

Commercial License

Free

Yes

No

No

No

Starter

Yes

Yes

No

Yes

Creator

Yes

Yes

Yes

Yes

Creating a Synthetic Voice with Voice Design

Voice Design is the easiest place to start. You don’t need any recordings. It generates a voice from scratch. It’s also the only option available on the free plan, which makes it perfect for beginners.

The workflow is simple:

  1. Write a prompt with some key settings
  2. Generate
  3. Preview
  4. Save

One tip from experience: generate at least 5 to 10 variations before picking one. Even with the same settings, results vary quite a bit. To get started, let’s click on the Voice Design button in the Create Voice menu. This will open a menu that allows you to write a prompt for your voice.

Voice Design prompt

You can tap Settings to adjust two details:

  • Loudness: how loud the voice will be when generated. 
  • Guidance level: similar to temperature in other models, it indicates how closely the AI should follow your prompt. Higher guidance means it sticks more strictly to what you say. Lower guidance means you give it more freedom.

Adjust loudness and guidance scale

You can choose to let the AI agent have its own preview text, or you can provide it with your own script by turning off the setting and putting your own text in the preview text box.

Preview text

Prompting voice parameters

You’ll want to use the prompt area to generate your configurations. Try using the following prompt template to key in on some specific parameters you want:

Native <Language>.  <Accent>, <Gender>, <Age range>, <Quality level>.
Persona: <2–5 words>. Emotion: <2–3 adjectives>.
<1–2 sentences about timbre, pacing, delivery>

Each parameter influences the output:

  • VoiceLabs is multilingual, so language determines the language tonality of the voice
  • Accent changes pronunciation patterns
  • Age affects pitch and tone
  • Gender influences vocal range
  • The quality level determines how clean the audio sounds

The interesting part is how these combine. Sometimes unexpected combinations produce the best results, so it’s worth experimenting.

Generating, previewing, and saving

Click Generate voice, and ElevenLabs produces a short sample.

Sample generated voices

You can regenerate as many times as you want. Each version will be slightly different.

Once you find one you like:

  • Click Save

  • Use a descriptive name like narrator_us_male_30s

After saving, the voice will appear under My Voices and in the Speech Synthesis dropdown.

Cloning a Voice with Instant Voice Cloning (IVC)

Instant Voice Cloning lets you replicate a voice using just a small audio sample. It’s fast and surprisingly effective, though not perfect. Remember, you’ll need the Starter ($6/mo) plan or higher for this.

Important note: only clone voices you have permission to use. The platform requires you to confirm this, and you should take it seriously. The full pipeline is pretty straightforward:

  • Prepare your audio
  • Upload the audio
  • Name and save your voice clone
  • Test in Text to Speech

Preparing your audio sample

So you want your audio to be of the right length, the right format, and of decent quality. The minimum length is about a minute, but I’ve found that 3 to 5 minutes gives much better results.

You should upload either MP3 or WAV files under 50MB, and you can always upload multiple files at once. 

To ensure the best cloning experience and good audio quality, I recommend the following when recording:

  • Quiet room
  • No background noise
  • Consistent microphone distance

Make sure to avoid the following:

  • Multiple speakers
  • Phone recordings
  • Heavy post-processing

Uploading and creating the clone

Navigate back to your Voices dashboard and do the following:

  1. Click + Create Voice
  2. Select Instant Voice Cloning
  3. Upload your audio and click Next

Instant Voice Clone

  1. Name your voice, optionally add labels and a description
  2. Confirm consent
  3. Click Save Voice

The following labels can be chosen from a dropdown menu:

  • Language
  • Accent (opens up options depending on the language)
  • Gender
  • Age (options: young, middle-aged, or old)

The voice is ready immediately, and you can test it right away in the text-to-speech generator. Try a few different sentences and pay attention to where it sounds natural and where it struggles.

To do this, click on Text to Speech on the left sidebar. (Note that some versions have referred to this as Speech Synthesis.)

Text to Speec

This opens a window with a similar text prompt window. 

Text to Speech window

On the right side, click on the Voice dropdown and select your voice from the My Voices window.

Voice selection

Write a test sentence and click on Generate speech.

Generate speech

This will generate an audio sample with your chosen voice. Feel free to adjust the settings on the right side to suit your preferences. You can adjust settings such as:

  • Speed
  • Stability
  • Similarity
  • Style Exaggeration 

Choosing different models can also provide different options!

Adjusting generated speecj

To save your file, click the download button on the right to save the generated audio.

Save generated speech

Building a High-Fidelity Clone with Professional Voice Cloning (PVC)

If IVC gives you a rough approximation, PVC gives you something much closer to the real thing. This is where you capture nuances like pacing, breath, and emotion.

It requires the Creator plan ($22/mo). According to ElevenLabs, processing typically takes 2 to 6 hours, though the total training time can be up to 24 hours depending on server load. 

It’s the most relevant option if we’re trying to build production-grade voice assets.  Think things like narration, audiobooks, and branded voice agents. Things that might see commercial or widespread use.

Recording and curating your training data

Since we’re trying to create the best possible model, the requirements have gone up. For instance, the minimum recording time is 30 minutes of clean audio, but for best results, aim for 3+ hours. The best way to submit this much audio is to break it up into 30-minute samples.

Here are some tips to make the most of your recording:

  • Use a quiet recording environment
  • Try to use a good microphone
  • Include varied content, don’t just read from a single source document

For example, you might want to mix up your narration style:

  • Use different kinds of dialogue. Read a few instructional texts. 
  • Go through some numbers and a list. This provides a lot of variety in terms of pronunciation and syntax. 
  • On top of that, try recording the voice at different paces, with different emotions, and so on. Adding more flair and style helps build a better model.

As always, avoid bad quality audio like the following:

  • Background noise
  • Heavy compression
  • Fillers like “um” and “uh”

Submitting and waiting for training

Once you get done with preparing the audio, the steps are pretty easy:

  1. Select Professional Voice Clone
  2. Click on the Create new clone button

Create a Professional Voice Clone

  1. Upload all recordings, fill in the name, language, and other labels

PVC details

  1. Fill in consent details
  2. Submit

Before ElevenLabs trains your clone, they make you prove the voice is actually yours.

This is the big difference from IVC: professional cloning only lets you clone your own voice, so you can't clone someone else's, even with their consent. If a colleague wants to lend theirs, they have to create and verify the PVC on their own account, then share it with you through a private sharing link.

Verification is a short live recording. You'll read a few on-screen lines into your mic. Two things make or break it:

  • Use the same or very similar equipment you used to record your samples, and maintain a similar tone and delivery. A mismatch here is the most common failure.
  • Read each line only once, then hit Stop. Reading it more than once can cause verification to fail.

If it fails, you can wait 24 hours and retry, or reach out to support. One heads-up: once you reach the verification stage, the clone is locked in. You can't delete an unverified PVC yourself, so make sure your samples are right before you get here.

You’ll get notified when your voice clone is ready! Once that is the case, a useful trick is to compare your IVC and PVC outputs using the same sentence. The difference is usually very clear.

Using Your VoiceLab Voices in Projects

Once your voice is saved, it becomes available across the platform everywhere ElevenLabs generates audio.

You can use it in:

  • Text to Speech (Speech Synthesis) for quick generation
  • Studio for long-term projects
  • Python API for programmatic usage

I’ll go over how we can use your voice for each of these tools. The Beginner’s Guide to the ElevenLabs API is the best way to get started with the Python API.

Using custom voices in Text to Speech and Studio

In Text to Speech you just need to select your voice from the Voices -> My Voices dropdown, as we mentioned above.

Some tuning tips for the Text to Speech tool:

  • Start with Stability at 40 to 50 percent
  • Set Similarity around 75 percent
  • Adjust based on output

It might take a few tries to get the exact voice and recording you want, but the more you experiment, the better you’ll understand the speech generation process.

In Studio, it’s somewhat similar. You start by navigating to Studio on the left sidebar, then select + New Blank Project. Next, you can choose the kind of project:

  • Audio Project or
  • Video Project

An audio project is just that; it allows you to create long-form conversational content with multiple voices. A video project requires you to upload a video first, where you can then add voices to voice-over a video.

Audio project navigation in the Studio

The Studio audio project allows you to create a multi-speaker audio file. This can be great for generating podcasts or conversational content. Even audiobook content where you want different voices or different characters. 

The Studio’s audio dashboard is a blank canvas for you to work on. We’ll focus on the key voice component here, but feel free to play around with it.

Audio project in the Studio

On the left, you will see a voice section selected for you. As you type in the prompt area, it will highlight the character’s speech for you.

Creating new paragraph

When you go to the next line, you can then select a different voice and have a different character. Each of these lines is a “paragraph”:

Selecting voice for a paragraph

The limits are quite high for the Studio. Each project can have up to 500 chapters, each with up to 400 paragraphs. Each paragraph can have up to 5000 characters.

The number of projects you can have will depend on your tier:

  • Free: 5 projects
  • Starter: 20 projects
  • Creator: 1,000 projects

Video project navigation in the Studio

Let’s follow the same steps, but this time choose a video project. You are taken to a blank canvas and asked to upload a video:Video project in Studio

Once you’ve uploaded a video clip, you can see it on the left and hit play to preview. You can add multiple videos if needed.

Then go to the left and choose Speech.

Video project files

Similar to an audio project, you can choose multiple voices and type the sentences of interest. As you type, you can hit Enter to get to a new line. It’s possible to choose different voices for each line to get a conversation going.

Add voice to a video

At the bottom, you can easily adjust the timeframes for each line by dragging and dropping them on the timeline.

Video editing with generated speech

If you do not have photos or video, you can generate them using the Video tab on the left. You simply write a prompt, which then generates the image or video of choice. 

Please note that you must first accept the Image and Video beta terms of service. Also, free members are only allowed to generate images; you must have at least a Starter membership ($5/mo) to generate video.

Generate video

Accessing custom voices via the ElevenLabs Python SDK

To use the Python SDK, you must first have Python installed. Then you want to create a Python environment where you can then install the ElevenLabs SDK using the following: pip install "elevenlabs[pyaudio]"

Next, go to your ElevenLabs developers dashboard by clicking on the bottom of the left sidebar and selecting API Keys -> Create Key.

ElevenLabs developer console

You will then select what tools you want to give access to the API and click Create Key.

Create API key

This will pop up an API Key where you must copy or else you will lose it forever.

Created API key

Next, save the API key to a .env file in a safe place or a project folder.

After that, go to your Voices dashboard and select My Voices. You’ll want to copy the Voice ID of your voice. I would save it somewhere easy for you to access later.

Copy voice ID

Now you can create a script to run your ElevenLabs voice! Here’s a simple example:

import os
from dotenv import load_dotenv
from elevenlabs.client import ElevenLabs

# Loads the .env file from the folder
load_dotenv()

# uses the API key to get access to the ElevenLabs Client
client = ElevenLabs(api_key=os.getenv("ELEVENLABS_API_KEY"))
# Sets the voice ID for your voice
custom_narrator_voice_id = "your_voice_id_here"

# Calls the text to speech API using your custom voice ID
# Creates an audio file with the specific text
# Uses the model and output format defined for the voice
audio = client.text_to_speech.convert(
    voice_id=custom_narrator_voice_id,
    text="Thanks for getting through this DataCamp article!.",
    model_id="eleven_v3",
    output_format="mp3_44100_128",
)

# Takes the output audio and joins it into a byte object
audio_bytes = b"".join(audio)

# Saves the audio bytes to a file 
with open("output.mp3", "wb") as f:
    f.write(audio_bytes)

Conclusion

VoiceLab offers three paths to choose from, depending on your needs. Voice Design is perfect for experimenting, Instant Voice Cloning is great for quick prototypes, and Professional Voice Cloning is where you go for production-level quality.

If you’re just getting started, I recommend sticking with Voice Design first. Upgrade only when you have a clear use case.

If you want to go deeper into building AI-powered applications, check out our Developing AI Applications skill track, which teaches you to use the latest AI developer tools, including the OpenAI API, Hugging Face, and LangChain.

ElevenLabs VoiceLab FAQs

Is ElevenLabs free to use?

Yes. The free plan gives you about 10,000 credits a month (roughly 10 minutes of audio) plus Voice Design and the stock voice library. But it can't be used commercially and doesn't include voice cloning; you'll need at least the Starter plan for Instant Voice Cloning and a commercial license.

Do I need a paid plan to use ElevenLabs voices commercially?

Yes. The free plan requires ElevenLabs attribution and grants no commercial rights, so you can't monetize that audio. A commercial license starts with the Starter plan (around $6/month, or $5 billed annually), while Professional Voice Cloning for production-grade voices requires the Creator plan ($22/month) or higher.

Can I clone someone else's voice with ElevenLabs?

Only with explicit permission, and the rules differ by method. Instant Voice Cloning lets you clone any voice you're authorized to use, but Professional Voice Cloning is restricted to your own voice and requires a live verification recording (even with consent). Using a cloned voice for impersonation or fraud is illegal in most jurisdictions.

What's the difference between Instant and Professional Voice Cloning?

Instant Voice Cloning (IVC) produces a usable clone in seconds from just 1–2 minutes of audio and is ideal for prototypes, though it can miss subtle nuances. Professional Voice Cloning (PVC) trains a dedicated model on 30 minutes to 3 hours of audio and takes a few hours to process, giving a far more accurate, production-ready result. Use IVC to test quickly and PVC when quality matters.

Can I use my custom ElevenLabs voice outside the platform?

Not directly. Voice clones can't be exported and only work inside ElevenLabs. You can still use them anywhere the platform generates audio, including Text to Speech, Studio, and programmatically via the ElevenLabs API and Python SDK, which is how most people wire a custom voice into their own apps or pipelines.


Tim Lu's photo
Author
Tim Lu
LinkedIn

I am a data scientist with experience in spatial analysis, machine learning, and data pipelines. I have worked with GCP, Hadoop, Hive, Snowflake, Airflow, and other data science/engineering processes.

Topics

Learn AI with DataCamp!

Track

Developing AI Applications

21 hr
Learn to create AI-powered applications with the latest AI developer tools, including the OpenAI API, Hugging Face, and LangChain.
See DetailsRight Arrow
Start Course
See MoreRight Arrow
Related

blog

Voxtral TTS: A Guide With Practical Examples

Learn how Mistral's first text-to-speech model works, how it compares to existing alternatives, and how to generate speech using the Python SDK with step-by-step code examples.
Khalid Abdelaty's photo

Khalid Abdelaty

9 min

Tutorial

A Beginner’s Guide to the ElevenLabs API: Transform Text and Voice into Dynamic Audio Experiences

Harness the capabilities of the ElevenLabs API, a powerful AI voice generator. Learn how to transform text into speech and clone voices with this technology.
Stanislav Karzhev's photo

Stanislav Karzhev

Tutorial

OpenAI's Audio API: A Guide With Demo Project

Learn how to build a voice-to-voice assistant using OpenAI's latest audio models and streamline your workflow using the Agents API.
François Aubry's photo

François Aubry

Tutorial

Grok Voice Agent Builder: A Hands-On Guide in Python

Build a Python voice agent with the same API used by Grok Voice Agent Builder: WebSocket setup, audio streaming, tool calling, cost tracking, and a FastAPI endpoint.
Khalid Abdelaty's photo

Khalid Abdelaty

Tutorial

NVIDIA PersonaPlex Tutorial: Run a Natural, Real-Time Local Voice Assistant

What if real-time AI voice conversations felt natural, interruptible, and genuinely human? Learn how to run NVIDIA PersonaPlex locally and experience true full-duplex conversational AI.
Abid Ali Awan's photo

Abid Ali Awan

code-along

Create a Deepfake AI Puppet of Yourself with Gemini & Elevenlabs

Rhys Phillips, Content Operations Manager at DataCamp, will show you how to generate AI videos from reference images and AI audio from sound clips, then combine them into a controllable avatar.
Rhys Phillips's photo

Rhys Phillips

See MoreSee More