Detecting AI-Generated Speech: Challenges and Advances in Audio Forensics — ICACNC 2025 | TechShield Publications
ICACNC 2025 · Conference Article

Detecting AI-Generated Speech: Challenges and Advances in Audio Forensics

Authors: Smar Ahmad Sohail, Aqsa Nawaz, Muhammad Zohaib, Iftikhar Rasheed, Umar Fayyaz

Abstract

Recently, people and organizations have been having problems distinguishing real speech from AI-generated speech, as voice-generation methods and deep learning developments make AI voices increasingly similar to natural-sounding voices. This extreme level of naturalness allows for synthesizing an individual’s target voice and regenerating someone’s voice with very high accuracy. Audio forensics is essential for enhancing speech intelligibility and conducting analysis to establish whether audio recordings are authentic and to determine if any manipulation has been performed on them. This review explores the methods and technologies used to verify voice authenticity, specifically in identifying AI-generated voices, discussing algorithms, techniques, and tools applied for voice authentication — including preprocessing, feature extraction (such as chromagram, spectrogram, Mel-spectrum, and MFCC), voiceprint databases, and matching logic using CNNs. It also examines the limitations and challenges of these systems, especially as AI continues to advance in generating highly realistic fake voices.

Audio Forensics Deepfake Audio Detection Voice Authentication MFCC CNN

Cite This Paper

S. A. Sohail, A. Nawaz, M. Zohaib, I. Rasheed, and U. Fayyaz, “Detecting AI-Generated Speech: Challenges and Advances in Audio Forensics,” Proc. Int. Conf. on AI, Cybersecurity, and Next-Gen Computing (ICACNC 2025), The Islamia University of Bahawalpur, Jun. 2025, doi: 10.67535/tsp.000001.028.