> ## Documentation Index
> Fetch the complete documentation index at: https://smallestai-ff1e543d.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Voice Cloning Best Practices

> Guidelines for recording reference audio and achieving high-quality voice clones.

High-quality reference audio is the single most important factor in clone quality. These guidelines cover recording environment, speaking style, multi-lingual cloning, and expressive control.

<Card title="Try Voice Cloning" icon="microphone" href="https://app.smallest.ai/dashboard/voice-cloning">
  Clone a voice directly in the console. 5-15 seconds of audio, no code required.
</Card>

***

## Recording Reference Audio

### Environment

* Record in a quiet room with minimal background noise. Ambient noise, hiss, or rumble will be captured in the clone.
* Use a dedicated microphone when possible. MacBook and mobile device microphones are acceptable if positioned at an appropriate distance to avoid distortion.
* Avoid rooms with echo (large empty spaces, outdoor areas). Small treated rooms produce the best results.
* After recording, listen back to the audio before uploading. Verify it is free of interruptions, clipping, or background interference.

### Speaking Style

* Speak naturally in your normal conversational voice. The model captures timbre, accent, emotional tone, rhythm, and pacing automatically.
* Maintain a consistent pace throughout the recording. Avoid long pauses, as they can degrade clone quality.
* Do not exaggerate emotion unless a specific tone is the intended output (see [Expressive Cloning](#expressive-cloning) below).

### Audio Length

* Provide **5 to 15 seconds** of clean, continuous speech.

***

## Multi-Lingual Voice Cloning

### Language Matching

For best results, record reference audio in the same language as your intended output. The model supports cross-lingual cloning (e.g., English reference audio used for Spanish output), but a language-matched reference will always produce higher fidelity.

| Scenario                                      | Expected Quality                                                               |
| --------------------------------------------- | ------------------------------------------------------------------------------ |
| Reference and output in the same language     | Best results. Highest phonetic accuracy.                                       |
| Reference in a different language than output | Functional. Voice characteristics transfer, but the source accent is retained. |

### Accent Retention

When synthesizing in a different language than the reference audio, the original accent is preserved. A clone from a South Indian English speaker will retain that accent when generating Hindi or Tamil output. This is by design: the clone reproduces *your* voice, including accent characteristics.

If accent-neutral output is required for a specific language, provide reference audio recorded by a native speaker of that language.

### Language Group Constraints

Cloned voices follow the same language group routing rules as standard synthesis. See [Code-Switching](/v4.0.0/content/text-to-speech/model-cards/lightning-v3-1#code-switching) for details on Indic and Global group restrictions.

***

## Expressive Cloning

The model captures emotional and prosodic characteristics from the reference audio. The tone, pace, and volume of the reference directly influence the synthesized output.

### Emotional Control

The emotion conveyed in the reference audio (e.g., calm, happy, angry) is reflected in the generated speech. To produce an angry-sounding clone, provide an angry reference. To produce a neutral clone, provide a neutral reference.

### Speed Control

The pace of the reference audio determines the output speed. A fast-paced reference produces faster delivery; a slower reference produces more measured output.

### Volume Control

The volume level in the reference audio carries over to the output. A soft-spoken reference produces quieter output; a louder, more energetic recording produces bolder output.

***

## Reference Audio Examples

<Note>
  Audio samples are embedded as video due to platform constraints.
</Note>

### Good Reference Audio

Clear, consistent tone with no background noise.

<video controls autoplay>
  <source src="https://mintcdn.com/smallestai-ff1e543d/DmcebTQzZ_bVLafo/video/good_ref_t.mp4?fit=max&auto=format&n=DmcebTQzZ_bVLafo&q=85&s=ff088a98953be5a76b706d5de2b56182" type="video/mp4" data-path="video/good_ref_t.mp4" />
</video>

### Bad Reference Audio

**Background noise present.**

<video controls autoplay>
  <source src="https://mintcdn.com/smallestai-ff1e543d/DmcebTQzZ_bVLafo/video/bg_ref_t.mp4?fit=max&auto=format&n=DmcebTQzZ_bVLafo&q=85&s=7dfb6fb1a3b7793a5ebad335ba27379d" type="video/mp4" data-path="video/bg_ref_t.mp4" />
</video>

**Inconsistent speaking style.**

<video controls autoplay>
  <source src="https://mintcdn.com/smallestai-ff1e543d/DmcebTQzZ_bVLafo/video/inconsistent_ref_t.mp4?fit=max&auto=format&n=DmcebTQzZ_bVLafo&q=85&s=9f61687bc102b2db1fc4bd4c00644259" type="video/mp4" data-path="video/inconsistent_ref_t.mp4" />
</video>

**Overlapping voices.**

<video controls autoplay>
  <source src="https://mintcdn.com/smallestai-ff1e543d/DmcebTQzZ_bVLafo/video/overlap_ref_t.mp4?fit=max&auto=format&n=DmcebTQzZ_bVLafo&q=85&s=44ee0823e3f447a3894adb08e3528a85" type="video/mp4" data-path="video/overlap_ref_t.mp4" />
</video>

***

## Expressive Audio Examples

### Angry Tone

**Reference:**

<video controls autoplay>
  <source src="https://mintcdn.com/smallestai-ff1e543d/DmcebTQzZ_bVLafo/video/angry_ref_t.mp4?fit=max&auto=format&n=DmcebTQzZ_bVLafo&q=85&s=bf6048cabda11cf271fae7f6da291e2e" type="video/mp4" data-path="video/angry_ref_t.mp4" />
</video>

**Output:**

<video controls autoplay>
  <source src="https://mintcdn.com/smallestai-ff1e543d/DmcebTQzZ_bVLafo/video/angry_gen_t.mp4?fit=max&auto=format&n=DmcebTQzZ_bVLafo&q=85&s=174cadce56a01583903b61e849a2ab69" type="video/mp4" data-path="video/angry_gen_t.mp4" />
</video>

### Whisper Tone

**Reference:**

<video controls autoplay>
  <source src="https://mintcdn.com/smallestai-ff1e543d/DmcebTQzZ_bVLafo/video/whisper_ref_t.mp4?fit=max&auto=format&n=DmcebTQzZ_bVLafo&q=85&s=e86716c62c72176b80c262e0a4fd10ba" type="video/mp4" data-path="video/whisper_ref_t.mp4" />
</video>

**Output:**

<video controls autoplay>
  <source src="https://mintcdn.com/smallestai-ff1e543d/DmcebTQzZ_bVLafo/video/whisper_gen_t.mp4?fit=max&auto=format&n=DmcebTQzZ_bVLafo&q=85&s=e7cf7d7a4f5f0ab92ff1904158f9a943" type="video/mp4" data-path="video/whisper_gen_t.mp4" />
</video>

### Fast-Paced Tone

**Reference:**

<video controls autoplay>
  <source src="https://mintcdn.com/smallestai-ff1e543d/DmcebTQzZ_bVLafo/video/fast_ref_t.mp4?fit=max&auto=format&n=DmcebTQzZ_bVLafo&q=85&s=635bba8650b7523255f8c8a782d7487f" type="video/mp4" data-path="video/fast_ref_t.mp4" />
</video>

**Output:**

<video controls autoplay>
  <source src="https://mintcdn.com/smallestai-ff1e543d/DmcebTQzZ_bVLafo/video/fast_gen_t.mp4?fit=max&auto=format&n=DmcebTQzZ_bVLafo&q=85&s=55b8a4204ec8e28698814148956c70ed" type="video/mp4" data-path="video/fast_gen_t.mp4" />
</video>
