# Tiny Aya - Cohere's Mini Multilingual Models

Source: https://www.youtube.com/watch?v=8i0zxyHKbfk
Recap page: https://rapidrecap.app/video/8i0zxyHKbfk
Generated: 2026-02-23T11:30:58.774+00:00

---
## Quick Overview

Cohere Labs launched the Tiny Aya family of specialized, small-sized (3.35B parameter) multilingual models, including TinyAya-Base and TinyAya-Global, designed to offer high performance across 67 languages, significantly outperforming larger models on low-resource languages and tokenization efficiency compared to models like Gemma 3 and Qwen 3.

**Key Points:**
- Cohere Labs released the Tiny Aya family of models, including the 3.35B parameter TinyAya-Base, which covers over 70 languages, including many lower-resourced ones.
- TinyAya-Global is a powerful instruction-tuned multilingual model built on TinyAya-Base, achieving strong, balanced performance across 67 supported languages.
- The development process involved regional specialization: models were trained on region-specific data (Europe, West Asia, Asia Pacific, Africa, South Asia) and then merged, resulting in specialized models like TinyAya-Water, Earth, and Fire.
- TinyAya models demonstrated superior tokenization efficiency compared to larger models like Gemma 3, Qwen 3, and SimoLLM 3 across various scripts, especially for non-Latin scripts like Greek and Indic languages.
- In generation quality benchmarks across five regions, TinyAya instruction-tuned models competed effectively with existing massively multilingual models, notably showing superior performance in Africa compared to competitors.
- The overall goal of Tiny Aya is to bridge the gap between model performance, coverage, and efficiency, enabling powerful, adaptable AI to run locally on devices like phones.

![Screenshot at 00:54: The performance graph compares the average multilingual generation quality \(y-axis\) against model size in billions of parameters \(x-axis\), clearly positioning TinyAya-Global above the performance/size frontier occupied by models like Gemma 3-4B, Qwen3-4b, and Mistral 3-3b.](https://ss.rapidrecap.app/screens/8i0zxyHKbfk/00-00-54.jpg)

**Context:** The video discusses the release of Cohere Labs' Tiny Aya suite of multilingual AI models, focusing on how these smaller, efficient models address the performance gap often experienced by lower-resource languages in larger, more common LLMs. The presentation contrasts the Tiny Aya approach, which emphasizes regional specialization and efficient tokenization, against established models like those from the Gemma and Qwen families, using performance charts and diagrams to illustrate their findings.

## Detailed Analysis

The video introduces Cohere's Tiny Aya family of small, multilingual AI models, emphasizing their efficiency and deep language coverage. The core model, TinyAya-Base, is a 3.35B-parameter model trained on over 70 languages, many of which are lower-resourced. Building on this is TinyAya-Global, an instruction-tuned version that performs strongly across 67 supported languages. The training methodology is complex, involving initial region-specific SFT models (Europe, West Asia, Asia Pacific, Africa, South Asia) which are then grouped and optimized regionally (e.g., Europe + West Asia + Asia Pacific) before merging into final models like TinyAya-Water, Earth, and Fire. A key finding highlighted is tokenization efficiency: a bar chart comparing Tiny Aya against Gemma 3, Qwen 3, and SimoLLM 3 shows Tiny Aya achieving significantly lower tokens per character/word for many non-European languages, like Gujarati and Burmese, demonstrating better representation for these scripts. Furthermore, a generation quality chart across five regions shows TinyAya instruction-tuned models competitively matching or exceeding existing massively multilingual models, particularly excelling in regions like Africa where other models struggle. The ultimate aim is providing powerful, adaptable AI that can run locally on devices like phones, focusing on performance, coverage, and efficiency.

### Audience Questions & Context

- Viewers requested Turkish, Indonesian, Polish, and Arabic support, highlighting demand for non-Western European languages
- The video addresses this by showing language coverage maps and subsequent model performance.

### LLaMA 2 Tokenizer Analysis

- Demonstrates LLaMA 2 tokenization inefficiency, showing English and French are compact (4 tokens for 'Me llamo Sam'), while Thai requires 14 tokens and Greek requires 23 tokens, illustrating the challenge for non-Latin scripts.

### Model Landscape Overview

- Compares model evolution, showing the shift away from 32K tokenizers (like LLaMA 2) toward newer models like Llama 4, Qwen3, Gemini 3.1 Pro, and Claude Opus 4.6, all of which still lag in low-resource language support.

### Tiny Aya Model Architecture & Data

- Tiny Aya prioritizes 67 languages across 5 regions (Europe, West Asia, South Asia, Asia Pacific, Africa) during post-training; specialized models are created by merging region-specific SFT models, while TinyAya-Global is created from an 'All region SFT Model'.

### Performance Benchmarks (Model Size vs. Quality)

- A scatter plot shows TinyAya-Global (3.35B parameters) achieving higher generation quality (approx. 0.87) than larger models like Gemma 3-4B (0.77), Ministral-14b (0.75), and others, fitting into the 'faster/cheaper' efficiency zone.

### Tokenization Efficiency by Script

- A bar chart explicitly demonstrates Tiny Aya's superior tokenization efficiency (fewer tokens per character) over Gemma 3, Qwen 3, and SimoLLM 3 for scripts like Burmese, Gujarati, and Greek.

![Screenshot at 00:00: A user comment asks for Turkish language support, setting the context for the video's focus on multilingual capabilities.](https://ss.rapidrecap.app/screens/8i0zxyHKbfk/00-00-00.jpg)
![Screenshot at 00:14: A bar chart titled 'Artificial Analysis Multilingual Index' compares the average multilingual generation quality of several models across seven languages, showing performance discrepancies.](https://ss.rapidrecap.app/screens/8i0zxyHKbfk/00-00-14.jpg)
![Screenshot at 00:54: A performance chart comparing four models \(Tiny Aya, Ministral, Gemma, Qwen\) across five geographical regions, highlighting Tiny Aya's competitive performance, especially in Africa.](https://ss.rapidrecap.app/screens/8i0zxyHKbfk/00-00-54.jpg)
![Screenshot at 01:31: Code output comparing LLaMA 2 tokenization: English \('My name is Sam'\) uses 4 tokens, while Thai uses 14 tokens and Greek uses 23 tokens, illustrating tokenization inefficiency for non-Latin scripts.](https://ss.rapidrecap.app/screens/8i0zxyHKbfk/00-01-31.jpg)
![Screenshot at 03:51: A graphic from the Cohere blog detailing the Tiny Aya family, which prioritizes 67 languages across 5 regions \(Europe, West Asia, South Asia, Asia Pacific, Africa\).](https://ss.rapidrecap.app/screens/8i0zxyHKbfk/00-03-51.jpg)
