Downloads provided by UsageCounts
The dataset contains 30 speakers (15 females and 15 males), with recorded utterances of single English digits from 0 to 9. The recordings are captured with the Kinect for Xbox One sensor from Microsoft. For each utterance, the video and depth streams with a frame rate of 30 fps are captured (cropped mouth region). Additionally, the audio data of the four-channel microphone array is stored.
speaker recognition, speech recognition, audio-visual
speaker recognition, speech recognition, audio-visual
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
| views | 16 | |
| downloads | 12 |

Views provided by UsageCounts
Downloads provided by UsageCounts