Deep fake Detection Techniques for Voice and Video-Based Social Engineering Attacks
Main Article Content
Abstract
The proliferation of generative artificial intelligence has enabled the creation of highly convincing synthetic audio and video content, commonly known as deep fakes. While these technologies have legitimate applications in entertainment, education, and accessibility, they have simultaneously become a potent weapon for social engineering attacks, including CEO fraud, voice phishing (vishing), identity impersonation, and disinformation campaigns. This paper presents a comprehensive review of detection techniques for voice and video-based deep fakes, with a specific focus on their application in identifying and mitigating social engineering threats. We examine signal-level, artifact-based, biological, and deep-learning-driven detection methods for both audio and visual modalities, along with emerging multimodal and behavioral approaches. The paper further discusses the datasets and evaluation benchmarks used in this domain, the adversarial arms race between generation and detection systems, and the organizational and human-factor countermeasures that complement technical detection. We conclude by identifying open challenges—including generalization across generation methods, real-time detection constraints, and the need for explainable and robust detection frameworks—and propose directions for future research aimed at securing communication channels against AI-enabled impersonation.
Downloads
Article Details
Section

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.