WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal ResearchThanks: X. Mei, H. Liu, M. D. Plumbley, and W. Wang are with the Centre for Vision, Speech, and Signal Processing, University of Surrey, Guildford, GU2 7XH, U.K. (E-mail: [x.mei, haohe.liu, m.plumbley, w.wang]@surrey.ac.uk)Thanks: C. Meng is with Johns Hopkins University, U.S.A. (E-mail: cmeng9@jhu.edu)Thanks: Q. Kong is with The Chinese University of Hong Kong, Hong Kong, China. (E-mail: qqkong@ee.cuhk.edu.hk)Thanks: T. Ko, and C. Zhao are with ByteDance, China. (E-mail: [tom.ko, zhaochengqi.d]@bytedance.com)Thanks: Y. Zou is with the School of Electronic and Computer Engineering, Peking University, Shenzhen Graduate School, Shenzhen, 518055, China. (E-mail: zouyx@pku.edu.cn)
Xinhao Mei Chutong Meng Haohe Liu Qiuqiang Kong Affiliation: Tom Ko, Chengqi Zhao, Mark D. Plumbley, Yuexian Zou, Wenwu Wang
Abstract
The advancement of audio-language (AL) multimodal learning tasks has been significant in recent years, yet the limited size of existing audio-language datasets poses challenges for researchers due to the costly and time-consuming collection process. To address this data scarcity issue, we introduce WavCaps, the first large-scale weakly-labelled audio captioning dataset, comprising approximately 400k audio clips with paired captions. We sourced audio clips and their raw descriptions from web sources and a sound event detection dataset. However, the online-harvested raw descriptions are highly noisy and unsuitable for direct use in tasks such as automated audio captioning. To overcome this issue, we propose a three-stage processing pipeline for filtering noisy data and generating high-quality captions, where ChatGPT, a large language model, is leveraged to filter and transform raw descriptions automatically. We conduct a comprehensive analysis of the characteristics of WavCaps dataset and evaluate it on multiple downstream audio-language multimodal learning tasks. The systems trained on WavCaps outperform previous state-of-the-art (SOTA) models by a significant margin. Our aspiration
原文 arXiv:2303.17395;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2303.17395v2