
Recent advances in deep learning have en-abled highly realistic synthetic speech, creating serious risks such as impersonation, fraud, and misuse of voice-based authentication systems. Detecting AI-generated speech is increasingly difficult because modern text-to-speech and voice conversion models can closely imitate human prosody and timbre across languages. This paper proposes VoiceGuard, a hybrid deep learning framework that combines complementary spectral and temporal rep-resentations for deepfake voice detection. A Convolutional Neural Network (CNN) branch learns frequency-domain artifacts from spectrograms, while a CNN-GRU branch models temporal inconsistencies from acoustic descriptors. An attention-based fusion mechanism adaptively weights branch outputs to improve discriminative power. The framework is evaluated on benchmark datasets and cross-lingual settings, and it improves performance compared to single-representation approaches while remaining compu-tationally practical for real-world deployment.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
