Will AI Preserve the Georgian Language?

In the age of artificial intelligence, a language no longer lives only in human speech. It must also exist in search engines, digital archives, voice recordings, software interfaces and the datasets from which algorithms learn. Georgian remains strong in daily life, yet its position in AI is much more fragile. In August 2026, W3Techs estimated that Georgian was used by less than 0.1% of websites whose content language could be identified.

The contrast is striking because Georgia itself is highly connected. Geostat’s 2025 household ICT survey indicates that roughly 88% of people aged six and above had used the internet during the previous three months. Georgian citizens are online, but a substantial share of Georgian knowledge remains in closed repositories, scanned PDFs, unstructured websites and archives with unclear reuse rights. Content that a person can see online is not automatically usable as responsible AI training data.

Whether AI will preserve Georgian is therefore not a matter of technological prediction. It depends on whether Georgia builds a sufficiently large, diverse, high-quality and legally reusable digital language resource. AI cannot strengthen a language that it is not given an opportunity to learn; properly organised Georgian knowledge, however, can turn AI into one of the language’s most scalable distribution tools.

High digital use, limited global representation

Indicator Period Value Meaning Source/status
Internet users, age 6+ 2025 about 88% The Georgian audience is already digital Geostat; official estimate
Georgian among identifiable website languages Aug. 2026 below 0.1% Georgian is a small language on the global web W3Techs; current estimate
Active spoken or signed languages worldwide 2026 7,000+ Global language diversity is vast UNESCO
Languages well supported by AI 2026 a small fraction Leading systems concentrate on data-rich languages UNESCO; qualitative
BTU open Georgian language resource 2026 about 12m tokens Georgian data are openly available to the international AI ecosystem BTU; open GitHub dataset

 

A share below 0.1% does not mean Georgian is disappearing. It means Georgian is a very small supplier in the global data market. Multilingual transfer allows large models to produce Georgian text, but limited local material raises the risk of awkward terminology, imported sentence structures, shallow knowledge and weak interpretation of Georgian realities.

For AI, both volume and quality matter. A million near-duplicate news items cannot replace a balanced collection covering history, science, business, law, medicine, literature, dialects, public data and natural speech. Georgia needs more text, but also better described and more diverse text.

BTU’s project: 12 million Georgian tokens in open access

BTU has already responded with a concrete resource. With Palitra Media’s support, the university created a large Georgian-language digital database. Palitra Media supplied millions of media items accumulated over many years. The source archive and the current open release should be distinguished: the archive provides a foundation for expansion, while the initial dataset published on GitHub contains approximately 12 million tokens.

The resource includes words, sentences, verb forms, declensions, idioms and terminology, as well as the language of media, business and economics and official data. Such diversity helps AI systems process Georgian not only formally but contextually and semantically.

Publishing the dataset on GitHub makes it available to international researchers, developers and the broader AI ecosystem. It can support model adaptation, classification, search, translation and educational applications. It also creates an opportunity for systems such as ChatGPT, Google’s AI models, Claude and others to use Georgian material. Availability, however, does not prove that any of those companies has already incorporated the dataset into model training.

The project forms part of BTU’s Georgian Language Digital Sovereignty initiative. Here sovereignty does not mean isolating Georgian data. It means enabling Georgia to create, document and lawfully open a high-quality language foundation so Georgian enters AI with its own structure, meaning and cultural context. Palitra Media contributes scale and linguistic variety; BTU converts that material into a resource suitable for research and technological use.

How much material is needed?

There is no token count after which AI suddenly masters a language. Outcomes depend on data quality, the model, intended use and domain balance. BTU’s 12-million-token release nevertheless provides a measurable baseline. The 50-million and 200-million-token levels below are indicative expansion scenarios only.

Scenario Token volume Share represented by current 12m Additional volume Status
BTU initial open resource 12 million 100% Published
Expanded multi-domain resource 50 million 24% 38 million tokens Indicative scenario
National-scale resource 200 million 6% 188 million tokens Indicative scenario

 

Georgia does not need to write all of this from scratch. It already has books, studies, legislation, court decisions, statistics, media archives and audiovisual materials. The work is to inventory them, clarify rights, digitise content, remove duplication, protect personal data and publish eligible material under common technical standards.

Open data is not synonymous with unconditionally free data. Copyright, privacy, commercial confidentiality and culturally sensitive content require clear rules. A national language resource should be built around consent, licensing, provenance and meaningful opt-out mechanisms.

The next steps

First, Georgia needs a national inventory of text, speech, video, terminology and parallel translations: what exists, who owns it and how it may be used. Second, it needs a minimum publication standard covering machine-readable text, metadata, licensing, persistent links and quality labels. Third, Georgia needs a recurring AI benchmark that measures grammar, factual accuracy, local context, dialects, professional terminology and safety.

BTU’s publication of approximately 12 million tokens on GitHub is already a measurable starting point. Its broader effect will emerge if the resource grows continuously, quality and licensing remain controlled and other universities, publishers, media organisations and public bodies join it. The objective should not be a larger repository for one institution, but a shared, lawful and high-quality digital environment for Georgian knowledge.

The future of Georgian in AI is not solely a decision for technology companies. It depends on how much knowledge Georgians create, how much they make accessible, how carefully they describe it and how responsibly they make it usable. If Georgian knowledge remains invisible, AI may speak Georgian while knowing little about Georgia. If that knowledge becomes open, structured and reliable, AI can become not a replacement for the language, but a new channel for its expansion.

Recent Posts