
Soundscape, the sonic environment as perceived and understood by people, is a conglomerate of different sounds. It has been established that its appraisal by instantaneous annoyance is not solely determined by its calculated loudness, but also by recognised sounds. Hence, most previous research on annoyance has focused on single-source environments. Audio analytics aims at detecting and classifying sound sources, but does not explore human perception of these. This paper proposes a dual-input model to simultaneously perform sound source classification (SSC) and human annoyance rating prediction (ARP). The model takes mel features and root-mean-square value (rms) features as input, and uses convolutional blocks to extract high-level acoustic features. These are used to predict sound source classes and to estimate the human annoyance rating for the whole fragment. Experiments on the DeLTA dataset show that: 1) models using mel features and rms features outperform models using only one of them; 2) The proposed model achieves a SSC accuracy of 90.06%, and an ARP (scale 1 to 10) root mean square error of 1.05.
Technology and Engineering
Technology and Engineering
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
