Re-contextualizing Fairness in NLP: The Case of India
Shaily Bhatt Google Research \AndSunipa Dev Google Research \AndPartha Talukdar Google Research \ANDShachi Dave* Google Research \AndVinodkumar Prabhakaran* Google Research
Abstract
Recent research has revealed undesirable biases in NLP data and models. However, these efforts focus on social disparities in West, and are not directly portable to other geo-cultural contexts. In this paper, we focus on NLP fairness in the context of India. We start with a brief account of the prominent axes of social disparities in India. We build resources for fairness evaluation in the Indian context and use them to demonstrate prediction biases along some of the axes. We then delve deeper into social stereotypes for Region and Religion, demonstrating its prevalence in corpora and models. Finally, we outline a holistic research agenda to re-contextualize NLP fairness research for the Indian context, accounting for Indian societal context, bridging technological gaps in NLP capabilities and resources, and adapting to Indian cultural values. While we focus on India, this framework can be generalized to other geo-cultural contexts.
中文速览
现有NLP公平性研究几乎只关注西方社会的偏见问题,根本无法直接套用到印度这样拥有14亿人口、多宗教、多种姓、多民族的复杂社会。为此,作者系统梳理了印度社会中地区、种姓、性别、宗教等核心歧视维度,整理并创建了专用于印度语境的公平性评估资源(身份词汇表、印度人名列表等),并通过情感分析、语言模型等实验证明,主流NLP模型确实对达利特(Dalit)、穆斯林等边缘群体存在显著负面预测偏差,训练语料中也充斥着针对特定地区和宗教的刻板印象。在此基础上,文章提出了一套涵盖社会现实、技术能力与文化价值取向的整体研究议程,为印度乃至其他非西方地区的NLP公平性研究提供了可落地的路线图。这项工作的意义在于,它首次将NLP公平性问题从西方框架中真正"本土化"到印度语境,填补了全球最大民主国家在AI公平性研究上的空白。
原文 arXiv:2209.12226;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2209.12226v5