# Typhoon OCR: Open Vision–Language Model For Thai Document Extraction

Source: https://www.youtube.com/watch?v=xwctfmMZemU
Recap page: https://rapidrecap.app/video/xwctfmMZemU
Generated: 2026-01-23T19:06:14.46+00:00

---
## Quick Overview

The Typhoon OCR model, developed by a team from SCB 10X in Thailand, significantly outperforms large models like GPT-5 and Gemini 2.5 Pro in extracting data from Thai documents, achieving a RoJie score of 0.97 against GPT-5's 0.96, demonstrating superior performance by leveraging specialized architecture and data handling techniques.

**Key Points:**
- Typhoon OCR achieved a RoJie score of 0.97 on Thai document extraction, outperforming GPT-5's score of 0.96.
- The model was developed by a team from SCB 10X in Thailand to address the poor performance of general models on Thai documents.
- Key to Typhoon's success is its specialized architecture, which handles Thai characters (like vowels floating above/below consonants) and complex layouts better than general models.
- The training process involved a four-stage process: gathering real data, extracting text via traditional OCR, smart tokenization/formatting, and finally, human review/agentic quality control.
- Typhoon successfully extracts data, including LaTeX equations mixed with Thai text, which general models often fail to process correctly.
- The model's approach focuses on specialized training data (Thai documents, financial reports) rather than relying on massive, general internet data like larger models.

![Screenshot at 00:36: The presentation of the research paper's model name, Typhoon OCR, which is the core subject of the discussion regarding specialized document extraction.](https://ss.rapidrecap.app/screens/xwctfmMZemU/00-00-36.jpg)

**Context:** The discussion centers on a specialized Optical Character Recognition (OCR) model named Typhoon OCR, created by researchers at SCB 10X in Thailand. This research addresses the known difficulty that large, general-purpose Language Models (LLMs) like GPT-5 and Gemini 2.5 Pro have in accurately processing documents written in low-resource languages, particularly Thai, which has complex orthography involving diacritics that float above and below consonants.

## Detailed Analysis

The speakers discuss the Typhoon OCR model, a specialized solution for Thai document extraction developed by an SCB 10X team from Thailand, which significantly outperforms general-purpose large language models (LLMs) like GPT-5 and Gemini 2.5 Pro on this task. The primary metric discussed is the RoJie score, where Typhoon achieved 0.97, slightly better than GPT-5's 0.96. The core advantage of Typhoon lies in its specialized handling of the Thai language's complex structure, where characters are annotated with vowels floating above and below the main character, a feature general models struggle with, often leading to gibberish or incorrect structure (like reading across columns instead of down). The paper outlines a four-stage process: gathering real data (Thai documents, financial reports from the Thai stock exchange), using traditional OCR for initial extraction, applying smart tokenization (using spaces as delimiters, which is natural for Thai), and finally, implementing agentic quality control involving human review. The paper explicitly mentions that Typhoon excels at handling complex inputs like LaTeX equations mixed with Thai text, whereas larger models often fail or produce inaccurate results when processing such documents. The success of Typhoon suggests that specialized models trained on domain-specific, high-quality data can outperform massive general models on niche tasks, challenging the prevailing dogma that bigger is always better.

### Typhoon OCR Introduction

- Developed by SCB 10X team in Thailand
- Focuses on Thai document extraction
- Outperforms GPT-5 (0.97 RoJie vs 0.96)

### Key Performance Metrics

- Typhoon 0.97 vs GPT-5 0.96 on RoJie score
- Outperforms Gemini 2.5 Pro
- Superiority shown on complex tasks like Thai financial reports

### Typhoon's Technical Advantage

- Specialized architecture handles Thai orthography (vowels above/below characters)
- Uses Thai script structure (no spaces)
- Superior handling of LaTeX equations mixed with Thai text

### Data Curation Process

- Four stages: 1. Gather real data (Thai docs, financial reports)
- 2. Traditional OCR extraction
- 3. Smart tokenization/formatting
- 4. Agentic QC/Human review

### Comparison to General Models

- General models (like GPT-5) struggle with Thai structure and fail on complex formats
- Typhoon avoids forcing English layout conventions onto Thai data

![Screenshot at 00:00: The video opens with a promotional graphic inviting viewers to "Become a Member Today!" over an oscilloscope-like background, indicating a podcast or discussion format.](https://ss.rapidrecap.app/screens/xwctfmMZemU/00-00-00.jpg)
![Screenshot at 00:16: A speaker introduces the topic: a paper that is going to annoy Silicon Valley, setting up a comparison between models.](https://ss.rapidrecap.app/screens/xwctfmMZemU/00-00-16.jpg)
![Screenshot at 00:36: The slide or graphic explicitly names the model being discussed: "Typhoon OCR version 1.5."](https://ss.rapidrecap.app/screens/xwctfmMZemU/00-00-36.jpg)
![Screenshot at 02:26: Visual representation of the vertical nature of Thai script issues, contrasting with the horizontal layout favored by general models.](https://ss.rapidrecap.app/screens/xwctfmMZemU/00-02-26.jpg)
![Screenshot at 09:08: The speaker contrasts Typhoon's performance against GPT-5 and Gemini 2.5 Pro scores \(0.96 vs 0.97\), highlighting the performance gap.](https://ss.rapidrecap.app/screens/xwctfmMZemU/00-09-08.jpg)
