# Czy nowe technologie językowe są inkluzywne? | Joanna Dolińska | TEDxUniversityofWarsaw

Source: https://www.youtube.com/watch?v=0ISUObr8H7A
Recap page: https://rapidrecap.app/video/0ISUObr8H7A
Generated: 2025-12-03T18:37:42.78+00:00

---
## Quick Overview

New language technologies are not inherently inclusive, as demonstrated by the speaker's research showing that minority and under-resourced languages like Dagur and Basque lack sufficient digital resources, creating an exclusion gap that requires careful consideration of data sourcing and community engagement during development to ensure equitable technological advancement.

**Key Points:**
- The speaker questions the inclusivity of new language technologies, noting that minority and under-resourced languages like Khakas, Basque, and Dagur are often excluded due to lack of sufficient digital resources.
- The speaker's research involved working with Mongolian and Thai minority languages, specifically creating an experimental Dagur language corpus in collaboration with researchers from the University of Strasbourg.
- The research revealed that languages like Basque and Dagur share morphological features, such as agglutinative structure and the lack of grammatical gender, but often lack unified, standardized data for technology development.
- The speaker cites the Horizon Europe project, 'Fostering Language Richness in the European Union,' as a guiding principle for promoting digital language equality.
- A key takeaway is the necessity of including language communities in the development process, ensuring they understand how data is used and how they benefit from the technology.
- The speaker points to Google's plan to add around 100 languages to Google Translate by 2024 as a positive example of expanding access, but stresses that data quality and community involvement are paramount.
- The presentation emphasizes the need to address 'heritage data'—historical data collected by researchers or missionaries—cautiously, ensuring it aligns with current community standards.

![Screenshot at 04:47: The slide poses the central question, "What shall we do when we want to carry out experiments on a low-resource language corpus?" and immediately suggests the solution: using transfer learning from a multilingual language model, highlighting the core technical challenge and proposed methodology.](https://ss.rapidrecap.app/screens/0ISUObr8H7A/00-04-47.png)

**Context:** Joanna Dolińska, affiliated with the Faculty of Artes Liberales at the University of Warsaw, presents her work at a TEDxUniversityofWarsaw event focusing on the inclusivity of modern language technologies. The core concern is the digital divide affecting under-resourced languages, such as Mongolian minority languages (like Dagur) and Basque, which are often marginalized by AI and language processing tools developed primarily for high-resource languages. Her research focuses on methods to build computational resources for these endangered languages.

## Detailed Analysis

The talk addresses whether new language technologies are inclusive, concluding that they are not, particularly for under-resourced languages. Dolińska argues that the lack of sufficient digital resources for languages like Khakas, Basque, and Dagur excludes speakers from the benefits of modern AI tools like machine translation and voice assistants. She details her research, which focused on Mongolian minority languages, including creating a Dagur language corpus in collaboration with the University of Strasbourg, addressing linguistic features like agglutination and the absence of grammatical gender. This work was guided by the Horizon Europe project promoting digital language equality. A major challenge identified is the reliance on 'heritage data'—historical texts collected by missionaries or early researchers—which may not align with the community's current orthography or needs. The speaker stresses that any development must involve the language community to explain data usage and ensure benefits, contrasting this ethical approach with the sheer scale of major tech company rollouts, such as Google's plan to add 100 languages to Translate by 2024. The overarching message is that inclusivity requires careful consideration of data sources, community involvement, and the creation of standardized, ethical resources for all languages.

### Introduction to Language Technologies & Exclusion

- High-resource vs. minority languages (Khakas, Basque, Dagur)
- Lack of digital resources hinders minority language use in AI tools like Google Translate and GPT
- Speaker's background in computational linguistics and work with Mongolian and Thai minorities.

### Case Study

- Dagur Corpus Development: Research involved creating an experimental corpus for Dagur, a critically endangered Mongolian language spoken in China, in collaboration with the University of Strasbourg.

### Linguistic Features of Under-Resourced Languages

- Dagur and Basque share morphological features (agglutinative, no grammatical gender, Subject-Object-Verb sentence structure) but lack standardized data/orthographies.

### Ethical Considerations for Low-Resource Language Development

- Key principles include explaining data usage, discussing community benefits, addressing competing scripts/orthographies, and carefully evaluating 'heritage data' collected historically (e.g., by missionaries).

### Conclusion and Future Goals

- The goal is promoting digital language equality across the EU by inviting language communities into the development process, ensuring tools offer equitable access to services.

![Screenshot at 00:03: Title slide for the TEDx University of Warsaw event: "YOUNG SCIENCE: UNLEASH YOUR CURIOSITY" featuring the speaker's name and topic: "Are new language technologies inclusive?"](https://ss.rapidrecap.app/screens/0ISUObr8H7A/00-00-03.png)
![Screenshot at 01:00: Diagram illustrating various language technologies \(Machine translation, Voice assistants, Grammar/spell check\) feeding into speech/text tools, emphasizing the interconnected nature of NLP.](https://ss.rapidrecap.app/screens/0ISUObr8H7A/00-01-00.png)
![Screenshot at 00:46: Slide explicitly defining the challenge: "Under-resourced languages" lacking sufficient digital resources for technology development.](https://ss.rapidrecap.app/screens/0ISUObr8H7A/00-00-46.png)
![Screenshot at 03:37: Map showing the current distribution of Mongolian languages across Eurasia, highlighting locations like Buryat, Khalkha, and Dagur.](https://ss.rapidrecap.app/screens/0ISUObr8H7A/00-03-37.png)
![Screenshot at 04:47: Slide outlining the key questions and answers for developing technology for minoritized languages: focusing on data usage, community benefit, script discussion, and heritage data identification.](https://ss.rapidrecap.app/screens/0ISUObr8H7A/00-04-47.png)
