Detecting AI-Generated Speech: Challenges and Advances in Audio Forensics
Authors: Smar Ahmad Sohail, Aqsa Nawaz, Muhammad Zohaib, Iftikhar Rasheed, Umar Fayyaz
Abstract
Recently, people and organizations have been having problems distinguishing real speech from AI-generated speech, as voice-generation methods and deep learning developments make AI voices increasingly similar to natural-sounding voices. This extreme level of naturalness allows for synthesizing an individual’s target voice and regenerating someone’s voice with very high accuracy. Audio forensics is essential for enhancing speech intelligibility and conducting analysis to establish whether audio recordings are authentic and to determine if any manipulation has been performed on them. This review explores the methods and technologies used to verify voice authenticity, specifically in identifying AI-generated voices, discussing algorithms, techniques, and tools applied for voice authentication — including preprocessing, feature extraction (such as chromagram, spectrogram, Mel-spectrum, and MFCC), voiceprint databases, and matching logic using CNNs. It also examines the limitations and challenges of these systems, especially as AI continues to advance in generating highly realistic fake voices.
