Czy nowe technologie językowe są inkluzywne? | Joanna Dolińska | TEDxUniversityofWarsaw

Quick Overview

New language technologies are not inherently inclusive, as demonstrated by the speaker's research showing that minority and under-resourced languages like Dagur and Basque lack sufficient digital resources, creating an exclusion gap that requires careful consideration of data sourcing and community engagement during development to ensure equitable technological advancement.

Key Points: The speaker questions the inclusivity of new language technologies, noting that minority and under-resourced languages like Khakas, Basque, and Dagur are often excluded due to lack of sufficient digital resources. The speaker's research involved working with Mongolian and Thai minority languages, specifically creating an experimental Dagur language corpus in collaboration with researchers from the University of Strasbourg. The research revealed that languages like Basque and Dagur share morphological features, such as agglutinative structure and the lack of grammatical gender, but often lack unified, standardized data for technology development. The speaker cites the Horizon Europe project, 'Fostering Language Richness in the European Union,' as a guiding principle for promoting digital language equality. A key takeaway is the necessity of including language communities in the development process, ensuring they understand how data is used and how they benefit from the technology. The speaker points to Google's plan to add around 100 languages to Google Translate by 2024 as a positive example of expanding access, but stresses that data quality and community involvement are paramount. The presentation emphasizes the need to address 'heritage data'—historical data collected by researchers or missionaries—cautiously, ensuring it aligns with current community standards.

Context: Joanna Dolińska, affiliated with the Faculty of Artes Liberales at the University of Warsaw, presents her work at a TEDxUniversityofWarsaw event focusing on the inclusivity of modern language technologies. The core concern is the digital divide affecting under-resourced languages, such as Mongolian minority languages (like Dagur) and Basque, which are often marginalized by AI and language processing tools developed primarily for high-resource languages. Her research focuses on methods to build computational resources for these endangered languages.

Raw markdown version of this recap