Synthetic Data Augmentation for Zero-Shot Cross-Lingual Question Answering
Arij Riabi‡ Thomas Scialom⋆⋄∗ Rachel Keraron⋆ Benoît Sagot‡ Djamé Seddah‡ Jacopo Staiano⋆ ‡ Inria, Paris, France ⋄ Sorbonne Université, CNRS, LIP6, F-75005 Paris, France ⋆ reciTAL, Paris, France ∗: equal contribution. The work of Arij Riabi was partly carried out while she was working at reciTAL.
Abstract
Coupled with the availability of large scale datasets, deep learning architectures have enabled rapid progress on Question Answering tasks. However, most of those datasets are in English, and the performances of state-of-the-art multilingual models are significantly lower when evaluated on non-English data. Due to high data collection costs, it is not realistic to obtain annotated data for each language one desires to support.
原文 arXiv:2010.12643;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2010.12643v2