Demographic Dialectal Variation in Social Media: A Case Study of African-American English
Su Lin Blodgett† Lisa Green∗ Brendan O’Connor† †College of Information and Computer Sciences ∗Department of Linguistics University of Massachusetts Amherst
Abstract
Though dialectal language is increasingly abundant on social media, few resources exist for developing NLP tools to handle such language. We conduct a case study of dialectal language in online conversational text by investigating African-American English (AAE) on Twitter. We propose a distantly supervised model to identify AAE-like language from demographics associated with geo-located messages, and we verify that this language follows well-known AAE linguistic phenomena. In addition, we analyze the quality of existing language identification and dependency parsing tools on AAE-like text, demonstrating that they perform poorly on such text compared to text associated with white speakers. We also provide an ensemble classifier for language identification which eliminates this disparity and release a new corpus of tweets containing AAE-like language.
中文速览
非裔美国人英语(African-American English,AAE)在社交媒体上越来越活跃,但现有的自然语言处理工具几乎都是基于主流标准语言训练的,根本没法好好处理这类方言文本。研究者利用推特的地理定位数据和美国人口普查的街区级种族人口统计信息,设计了一个"远程监督"的混合成员概率模型,让机器自动从海量推文中识别出与非裔美国人社区相关的语言,无需人工逐条标注。验证结果表明,用这种方法提取出的文本确实符合已知的AAE语音和句法特征;进一步测试现有语言识别和依存句法分析工具后发现,这些工具在AAE文本上的表现明显比在白人用户文本上差,存在显著的种族性能差距,而研究者提出的集成分类器能有效弥补这一差距。这项研究不仅发布了一个包含83万条推文的AAE语料库,更揭示了NLP工具系统性忽视方言群体这一普遍问题,对推动更公平、更包容的语言技术开发具有重要意义。
原文 arXiv:1608.08868;中英对照 + 大白话阅读 https://aha.fim.ai/paper/1608.08868v1