Downloads provided by UsageCounts
Automated audio captioning (AAC) is the task of automatically creating textual descriptions (i.e. captions) for the contents of a general audio signal. Most AAC methods are using existing datasets to optimize and/or evaluate upon. Given the limited information held by the AAC datasets, it is very likely that AAC methods learn only the information contained in the utilized datasets. In this paper we present a first approach for continuously adapting an AAC method to new information, using a continual learning method. In our scenario, a pre-optimized AAC method is used for some unseen general audio signals and can update its parameters in order to adapt to the new information, given a new reference caption. We evaluate our method using a freely available, pre-optimized AAC method and two freely available AAC datasets. We compare our proposed method with three scenarios, two of training on one of the datasets and evaluating on the other and a third of training on one dataset and fine-tuning on the other. Obtained results show that our method achieves a good balance between distilling new knowledge and not forgetting the previous one.
The authors wish to acknowledge CSC-IT Center for Science, Finland, for computational resources. K. Drossos has received funding from the European Union's Horizon 2020 research and innovation programme under grant agreement No 957337, project MARVEL.
FOS: Computer and information sciences, Computer Science - Machine Learning, Sound (cs.SD), Clotho, Computer Science - Sound, 213, Machine Learning (cs.LG), learning without forgetting, Audio and Speech Processing (eess.AS), WaveTransformer, FOS: Electrical engineering, electronic engineering, information engineering, AudioCaps, automated audio captioning, continual learning, Electrical Engineering and Systems Science - Audio and Speech Processing
FOS: Computer and information sciences, Computer Science - Machine Learning, Sound (cs.SD), Clotho, Computer Science - Sound, 213, Machine Learning (cs.LG), learning without forgetting, Audio and Speech Processing (eess.AS), WaveTransformer, FOS: Electrical engineering, electronic engineering, information engineering, AudioCaps, automated audio captioning, continual learning, Electrical Engineering and Systems Science - Audio and Speech Processing
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
| views | 10 | |
| downloads | 11 |

Views provided by UsageCounts
Downloads provided by UsageCounts