Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection
Suchin Gururangan† Dallas Card♢ Sarah K. Dreier♡ Emily K. Gade♣ Leroy Z. Wang† Zeyu Wang† Luke Zettlemoyer† Noah A. Smith†♠ †University of Washington ♢ University of Michigan ♡University of New Mexico ♣Emory University ♠Allen Institute for AI {sg01, zwan4, lsz,
Abstract
Language models increasingly rely on massive web dumps for diverse text data. However, these sources are rife with undesirable content. As such, resources like Wikipedia, books, and newswire often serve as anchors for automatically selecting web text most suitable for language modeling, a process typically referred to as quality filtering. Using a new dataset of U.S. high school newspaper articles—written by students from across the country—we investigate whose language is preferred by the quality filter used for GPT-3. We find that newspapers from larger schools, located in wealthier, educated, and urban ZIP codes are more likely to be classified as high quality. We then demonstrate that the filter’s measurement of quality is unaligned with other sensible metrics, such as factuality or literary acclaim. We argue that privileging any corpus as high quality entails a language ideology, and more care is needed to construct training corpora for language models, with better transparency and justification for the inclusion or exclusion of various texts.
中文速览
训练大语言模型时普遍依赖的"质量过滤器"(quality filter)究竟在偏爱谁的语言?研究者以GPT-3所用的质量过滤器为研究对象,收集了全美近百万篇高中校报文章,并结合人口普查与教育统计数据,分析过滤器对不同来源文章的打分偏好。结果发现,来自更富裕、受教育程度更高、城市化程度更高地区以及规模更大学校的文章,更容易被判定为"高质量";而过滤器的这套"质量"标准,与事实准确性、文学奖项等更直观的质量指标并不吻合。这说明任何语料库的构建都隐含着一套"语言意识形态"(language ideology),当前主流训练数据系统性地偏向社会主流群体的语言风格,研究者呼吁业界在构建语言模型训练数据时提高透明度,明确说明文本纳入或排除的依据。
原文 arXiv:2201.10474;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2201.10474v2