WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research
Xinhao Mei, Chutong Meng, Haohe Liu, Qiuqiang Kong, Tom Ko, Chengqi Zhao, Mark D. Plumbley, Yuexian Zou, Wenwu Wang X. Mei, H. Liu, M. D. Plumbley, and W. Wang are with the Centre for Vision, Speech, and Signal Processing, University of Surrey, Guildford, GU2 7XH, U.K. (E-mail: [x.mei, haohe.liu, m.plumbley, Meng is with Johns Hopkins University, U.S.A. (E-mail: Kong is with The Chinese University of Hong Kong, Hong Kong, China. (E-mail: Ko, and C. Zhao are with ByteDance, China. (E-mail: [tom.ko, Zou is with the School of Electronic and Computer Engineering, Peking University, Shenzhen Graduate School, Shenzhen, 518055, China. (E-mail:
Abstract
The advancement of audio-language (AL) multimodal learning tasks has been significant in recent years, yet the limited size of existing audio-language datasets poses challenges for researchers due to the costly and time-consuming collection process. To address this data scarcity issue, we introduce WavCaps, the first large-scale weakly-labelled audio captioning dataset, comprising approximately 400k audio clips with paired captions. We sourced audio clips and their raw descriptions from web sources and a sound event detection dataset. However, the online-harvested raw descriptions are highly noisy and unsuitable for direct use in tasks such as automated audio captioning. To overcome this issue, we propose a three-stage processing pipeline for filtering noisy data and generating high-quality captions, where ChatGPT, a large language model, is leveraged to filter and transform raw descriptions automatically. We conduct a comprehensive analysis of the characteristics of WavCaps dataset and evaluate it on multiple downstream audio-language multimodal learning tasks. The systems trained on WavCaps outperform previous state-of-the-art (SOTA) models by a significant margin. Our aspiration
中文速览
音频与语言的联合理解研究长期受制于高质量配对数据的匮乏——现有最大音频描述数据集也只有约5万条。为此,研究者从FreeSound、BBC音效库、SoundBible及AudioSet等多个来源抓取了海量音频及其原始文字描述,并设计了一套三阶段自动化处理流程:先按文本频率做初步过滤,再借助大语言模型ChatGPT对嘈杂的原始描述进行内容筛查和改写,最后做后处理修正,最终构建出WavCaps——首个大规模弱标注音频描述数据集,包含约40万条音频及对应字幕。在文本-音频检索、自动音频描述等多项下游任务上,基于WavCaps训练的系统显著超越了此前的最优模型。这项工作不仅填补了音频语言多模态研究的数据缺口,还展示了用ChatGPT自动清洗和扩充学术数据集的可行路径,对整个音频AI领域具有重要的推动意义。
原文 arXiv:2303.17395;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2303.17395v2