
The emergence of AI-generated voices has posed significant problems with the authenticity of media and their digital safety. False audio detection or fake audio has been critical in such areas as audio forensics and voice authentication. In this paper, a literature review of deep fake audio detection with deep learning is conducted. The system used currently works with Mel-frequency Cepstral Coefficients (MFCCs) as the input feature and a VGG16based Convolutional Neural Network (CNN) as transfer learning to classify the real and fake voices. VGG16 is an effective model that can capture spectral variations but it is not able to learn temporal dependencies. To overcome this hybrid CNN-LSTM models have been investigated, which combine both spatial and time based feature learning to make them more accurate and robust.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 1 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Top 10% | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
