随着人工智能(AI) 侵入印度次大陆, 的不同地区,人们对研究人工智能如何影响这个文明的语言和文化基础产生了浓厚的兴趣。人工智能被视为''双刃剑'',一方面,它可以为大量人口提供访问和包容,,另一方面,它可以使世界观同质化并排除代表性不足的语言和世界观。在本文,中,我们试图通过解决印度语言学的广泛特征以及它们与文化实践和世界观紧密联系的方式来描述这个问题。然后,我们对自然语言处理 (NLP) 技术在这个领域的发展情况进行纵向调查,,追踪印度语 NLP, 的历史发展,涵盖关键里程碑, 方法转变, 和资源创建工作。此外, 该论文还研究了印度语言, 的结构和社会语言学特征,例如丰富的形态, 复杂的脚本和语法规则, 双语, 和较大的方言变化, 并解释了这些如何为构建人工智能基础模型带来独特的挑战。然后,我们讨论印度基础模型日益增长的作用,并分析这些模型如何解决这些长期存在的资源和代表性差距。最后,我们提出了一个名为'文化感知',的研究方向,它基于诠释学推理重新想象人工智能。文化感知旨在解决开放性问题,例如确保低资源语言的公平表现以及产生具有文化意义的产出。通过汇集过去的工作,当前技术,和新兴趋势,,本文概述了可以指导印度语 NLP 下一阶段并有助于开发更强大和更具包容性的印度语基础模型的研究方向。
As Artificial Intelligence (AI) makes inroads into different parts of the Indian subcontinent, there is significant interest in studying how AI impacts the linguistic and cultural foundations of this civilization. AI is seen as a ''double-edged sword where on the one hand, it can enable access and inclusion for a large population, on the other, it can homogenize worldviews and exclude underrepresented languages and worldviews. In this paper, we try to characterize this problem by addressing the extensive characteristic nature of Indian linguistics and the way they closely connect to cultural practices and worldview. We then perform a longitudinal survey of how Natural Language Processing (NLP) techniques have evolved in this space, tracing the historical development of Indic NLP, covering key milestones, methodological shifts, and resource creation efforts. In addition, the paper also examines the structural and sociolinguistic characteristics of Indian languages, such as rich morphology, complex scripts and grammar rules, diglossia, and large dialectal variation, and explains how these create unique challenges for building AI foundation models. We then discuss the growing role of Indic foundation models and analyze how these models address these long-standing resource and representation gaps. Finally, we propose a research direction called 'Culture Sensing', which re-imagines AI based on hermeneutic reasoning. Culture Sensing aims to address open problems such as ensuring equitable performance across low-resource languages and producing outputs that are culturally meaningful. By bringing together past work, current techniques, and emerging trends, this paper outlines research directions that can guide the next phase of Indic NLP and contribute to the development of more robust and inclusive Indic foundation models.
科目: 人工智能 (cs.AI); 计算和语言 (cs.CL)
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)