Downloads provided by UsageCounts
CLAP (Contrastive Language-Audio Pretraining) is a model that learns acoustic concepts from natural language supervision and enables “Zero-Shot” inference. The model has been extensively evaluated in 26 audio downstream tasks achieving SoTA in several of them including classification, retrieval, and captioning. Weights for the Microsoft CLAP model published in 2023 and 2022. clapcap is the audio captioning model that uses the 2023 encoders. Refer to the GitHub repository for the code. microsoft/CLAP: Learning audio concepts from natural language supervision (github.com)
CLAP, zero-shot, sound events, acoustic scenes
CLAP, zero-shot, sound events, acoustic scenes
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
| views | 923 | |
| downloads | 484 |

Views provided by UsageCounts
Downloads provided by UsageCounts