2026
Journal Article
Human Detection of Voice-Cloned Speech Under GSM, VoLTE and VoIP Conditions
Acoustics 2026, 8(2), 41;
The rapid progress of generative speech synthesis and voice-cloning technologies has enabled the creation of highly natural synthetic voices that pose a serious threat to telecommunication security. While most prior studies evaluate human ability to detect audio deepfakes using high-quality, studio-grade recordings, little is known about how real-world telecommunication channels affect perceptual detection. This study investigates the influence of three transmission scenarios—GSM (AMR-NB), VoLTE (AMR-WB), and VoIP with packet-loss modeling—on the human ability to distinguish natural speech from AI-generated speech. A custom speech corpus was developed, consisting of natural recordings from nine speakers and corresponding synthetic utterances generated using a state-of-the-art voice cloning system (ElevenLabs). All samples were processed through simulated telecommunication channels using real codec implementations. A listening test with 95 participants was conducted, involving binary classification (human vs. synthetic) and confidence ratings. Results show an overall detection accuracy of 54.8%, confirming that humans are poorly equipped to identify synthetic speech. Surprisingly, the highest accuracy was achieved for the narrowband GSM channel (63.7%), while VoLTE yielded the lowest performance (44.0%). The findings suggest that restricted bandwidth may emphasize prosodic irregularities typical of generative models, whereas high-quality channels mask synthetic artifacts, increasing susceptibility to voice spoofing. The results highlight the necessity of deploying additional security mechanisms in telecommunication systems relying on voice identity verification.