Xplore Publications
* Volume 3 of Transformations in Management is open for submissions until 30 August 2026. *

Chapter 61

VERNACULAR AI MANJILA MIRDAH

ISBN
978-81-992602-2-0
Published
21 July 2026
Accesses
2 views · 0 downloads
Reading time
~2 min

Full text

Vernacular AI Commons: Reimagining Data Sovereignty as Collective Linguistic Intellectual Property in Postcolonial India

Manjila Mirdah Khatun

Research Scholar of Adamas University

Assistant Professor of Law, L.J.D. Law College, Tollygunge Campus, affiliated to the University of Calcutta.

ORCID iD: 0009-0000-3288-4160

Email id- manjila0204@gmail.com,

ABSTRACT

India, with 22 languages under eight schedules of Constitution of India and 121 widely spoken languages, which comprise codified knowledge, oral traditions, and cultural expression, is threatened structurally by the commercial use of large language models that were primarily trained on English-language data. The existing intellectual property law is inadequate to compensate for damages arising from large scale consumption of vernacular language data by commercial Ai system. It does not acknowledge collective linguistic communities as subjects with rights, nor does it offer protection to oral and unwritten knowledge that copyright law inherently leaves out. For this purpose, there is a need to draw a comparative analysis of international laws including the Nagoya Protocol, the UNESCO Convention on Cultural Diversity, and also the Digital Personal Data Protection Act, 2023.Utilizing the ABS model created for biological genetic resources, the paper contends by analogy that vernacular linguistic corpora represent sovereign community assets, and that extensive commercial use by AI developers without prior informed consent or obligations for benefit-sharing equates to a type of digital linguistic biopiracy. This article proposes to consider the right to equitable access to the intellectual property as a sui generis right distinct from copyright, patent, and trademark, and advocates for a statutory National Vernacular Data Commons governed by community-representative approval protocols with mandatory benefit-sharing obligations binding on commercial AI developers. Articles 29 and 30 of the Constitution of India, read with the Preamble's guarantee of social, economic, and cultural justice, impose a positive obligation on the state to ensure this right to protect linguistic minorities. Without a robust legal architecture, the AI revolution is becoming a threat to India's long history of traditional knowledge in the post-colonial era.

Keywords: linguistic communities, large language models, sui generis.

Get an email when we publish new research and open calls for chapters.

Create a free account