
IS3+ is an extended version of IS3 with clean audio/image pairs to ensure cross-modality consistency. The dataset has 4 GB of data. The dataset contains the following data: audio_wav: audio files (.wav) gt_segmentation: annotations of image bounding boxes and segmentation masks images: images (.jpg) IS3_annotation.json: file with image/audio/gt information for every dataset sample. This work was done as part of the paper Learning from Silence and Noise for Visual Sound Source Localization Models. Paper citation: @misc{juanola2025learningsilencenoisevisual, title={Learning from Silence and Noise for Visual Sound Source Localization}, author={Xavier Juanola and Giovana Morais and Magdalena Fuentes and Gloria Haro}, year={2025}, eprint={2508.21761}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2508.21761}, }
Visual Sound Source Localization, IS3+
Visual Sound Source Localization, IS3+
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
