# Voice cloning failure signs: 15-dB test, denoise vs re-record

Dylan Cooper · September 11, 2026

> Spot voice cloning failures early with the 15-dB test, plus clear rules for when to denoise audio versus re-record for natural results every time.

| Takeaway | Detail |
| --- | --- |
| Total pension contributions must meet the statutory minimum. | 8% |
| Employers are legally required to pay a specific portion of qualifying earnings. | 3% |
| Workers contribute the remaining balance, including tax relief. | 5% |
| The annual earnings trigger for automatic enrolment remains unchanged. | £10,000 |

The UK pensions landscape for 2026/2027 reveals that the earnings trigger for automatic enrolment holds steady at £10,000 annually. This threshold translates to £768 every four weeks, ensuring that workers earning above this amount are automatically enrolled in workplace pension schemes. The Lower Qualifying Earnings Level is set at £6,240 annually, or £480 every four weeks, establishing the baseline for contribution calculations.

Contribution requirements mandate a total minimum of 8% of qualifying earnings. Employers must contribute at least 3%, while workers contribute 5%, which includes basic-rate tax relief paid by HMRC. These figures apply to workers aged at least 22 and below State Pension age who ordinarily work in the UK. Non-eligible jobholders earning between £6,240 and £10,000 can opt in, requiring employer contributions if they do.

Entitled workers earning below £6,240 may request to join, though employers are not required to contribute. The Upper Qualifying Earnings Level remains at £50,270 annually. These thresholds are reviewed yearly by the DWP under the Pensions Act 2008, with changes taking effect from April 6 following an announcement. Payroll duties require ongoing monitoring of age and earnings to ensure compliance.

![Empty home recording studio with acoustic foam wooden](https://static.mm-ais.com/article-images-ai/voice-cloning-failure-signs-15-db-test-d-ai-4b2f256d.jpg)
Empty home recording studio with acoustic foam wooden

## Why Sub-15-dB Audio Starves ECAPA-TDNN Embeddings and

Sub-15-dB enrollment audio starves ECAPA-TDNN embeddings by corrupting the spectral features they rely on for identity extraction. The failure is not merely a reduction in quality; it is an irreversible collapse of speaker-specific information. According to the UIC-AIHealth4All result, modern systems like ArchEHR-QA 2026 require high-fidelity inputs for grounded question answering, and voice cloning faces identical fidelity constraints. When enrollment SNR drops below 15 dB, the model cannot distinguish the speaker from the noise floor.

| Component | Failure Mechanism | Impact on Identity |
| --- | --- | --- |
| ECAPA-TDNN Embedding | Averages over 3-second windows with HVAC hum filling 80 mel-bins | Pulls vector toward generic noise centroid |
| Formant Region (F1/F2) | Masked in 500-2000 Hz range | Flattens accent and prosody cues |
| Pitch Tracking | PYIN octave errors when harmonics are within 10 dB of noise | Destroys fundamental frequency contour |
| Spectral Gating | Deletes fricatives /s,f,th/ above 6000 Hz | Leaves decoder to hallucinate muffled sibilants |
| Reverberation | 300-ms reflection tails blur phoneme boundaries | Smears temporal precision even if energy is loud |

The 15-dB test defines this threshold as 10*log10(speech-power/noise-power) computed with Python librosa 0.10 on a 2-second leading silence plus voiced speech at 16kHz, failing if the result is under 15 dB. This metric exposes the vulnerability of the ECAPA-TDNN architecture, which computes a 192-dim x-vector by averaging over 3-second sliding windows. In sub-15-dB conditions, broadband HVAC hum at 50-500 Hz fills 80-bin mel-spectrogram bins, pulling the embedding toward a generic noise representation rather than capturing unique vocal tract characteristics.

Formant masking in the 500-2000 Hz F1/F2 region further degrades identity. This critical band contains the primary resonances that define vowel space and accent. When noise obscures these frequencies, the PYIN pitch-tracker suffers octave errors because the harmonics sit within 10 dB of the noise floor. The resulting flattened prosody removes the rhythmic and intonational markers that listeners use to identify a speaker. Additionally, aggressive spectral gating deletes unvoiced fricatives /s,f,th/ above 6000 Hz and transient plosives. This leaves the decoder to hallucinate muffled sibilants from 44.1kHz-degraded input, creating artifacts that sound like a different person entirely.

Additive noise failure must be distinguished from late reverberation smearing. While additive noise reduces signal clarity, late reverberation blurs phoneme boundaries through 300-ms reflection tails. This smearing occurs even when frame-level energy looks loud enough to pass basic thresholds, making it a deceptive failure mode. Re-recording in a quiet under-35-dBA space is mandatory because AI denoising cannot reverse these structural losses. The only way to preserve speaker identity is to capture clean data at the source.

![Empty windswept train platform with steel beams concrete](https://static.mm-ais.com/article-images-ai/voice-cloning-failure-signs-15-db-test-d-ai-78a8c38e.jpg)
Empty windswept train platform with steel beams concrete

## The 15-dB Cliff

At the 15-dB threshold, neural voice cloning does not merely degrade; it undergoes a phase transition into irrecoverable identity collapse. This is not a linear loss of fidelity but a structural failure where the model can no longer distinguish speaker-specific timbre from acoustic noise. The mechanism is deterministic: below this SNR, the spectral features required for identity extraction are corrupted beyond the recovery capacity of current AI denoising architectures.

The perceptual reality of this collapse was quantified in January 2026 by Stanford HAI. In a controlled panel with n=212 listeners evaluating XTTS-style clones, the Mean Opinion Score (MOS) dropped precipitously from 4.21 at 20 dB to 3.08 at 10 dB. According to the Stanford HAI Speech Perception Study 2026, this delta represents a shift from "highly intelligible" to "noticeably artificial," confirming that the human ear detects identity drift long before the audio becomes unintelligible. The study attributes this directly to the erosion of high-frequency formant structures that carry unique speaker signatures.

This perceptual decay is mirrored by objective similarity metrics. Resemble AI’s Robustness Report 2026 documents a mean Speaker Embedding Cosine Similarity (SECS) of 0.84 at 18 dB, which collapses to 0.69 at 12 dB. According to Resemble AI, an SECS drop below 0.70 indicates that the cloned voice has lost its core identity markers, rendering it unsuitable for professional use regardless of post-processing. The data suggests that once the enrollment SNR dips below 15 dB, the embedding space becomes too sparse for the model to anchor the synthetic voice to the original speaker.

| Metric | Threshold | Value | Implication |
| --- | --- | --- | --- |
| Perceptual MOS | 20 dB vs 10 dB | 4.21 → 3.08 | Stanford HAI: Identity drift detected by humans |
| Speaker Similarity (SECS) | 18 dB vs 12 dB | 0.84 → 0.69 | Resemble AI: Embedding collapse below 0.70 |
| Transcription WER | Clean vs 10 dB | 7.2% → 19.6% | VCC 2024: Whisper-large-v3 fails on noisy enrollments |
| Audit Rejection Rate | Under 15 dB | 68% | Mozilla: Automatic rejection due to identity mismatch |
| Regeneration Attempts | Sub-15 dB | 3.4x more | Descript: ElevenLabs v2 requires excessive retries |

The downstream impact on automated pipelines is severe. The Voice Conversion Challenge 2024 organizers reported that Whisper-large-v3 transcription Word Error Rates (WER) rose from 7.2% on clean enrollments to 19.6% on 10-dB enrollments. According to VCC 2024 organizers, this increase in transcription error directly feeds garbage into the cloning model, amplifying artifacts. Furthermore, Mozilla’s Common Voice Noisy Clone Audit 2025 found a 68% automatic rejection rate for enrollments metering under 15 dB due to identity mismatch. According to Mozilla, these systems flag low-SNR inputs as "non-human" or "corrupted" because the statistical distribution of the speech features falls outside the training manifold.

Operational efficiency also suffers. Descript’s 2026 summary of ElevenLabs Multilingual v2 internal evaluations revealed that sub-15-dB enrollments required 3.4 times more regeneration attempts to produce one acceptable take. According to Descript, this inefficiency stems from the model’s inability to converge on a stable identity vector, forcing repeated sampling until a lucky artifact-free generation occurs. The consensus across these independent audits is clear: re-recording in a quiet environment (

Canonical: https://clonemyvoice.io/blog/voice-cloning-failure-signs-15-db-test-denoise-vs-re-record.php
Markdown: https://clonemyvoice.io/blog/voice-cloning-failure-signs-15-db-test-denoise-vs-re-record.php/index.md
