# GPT 5.3 Codex Spark: Its Crazy Fast

Source: https://www.youtube.com/watch?v=gP-3S7ITWzY
Recap page: https://rapidrecap.app/video/gP-3S7ITWzY
Generated: 2026-02-12T19:35:22.241+00:00

---
## Quick Overview

GPT-5.3-Codex-Spark achieves drastically faster, near-instantaneous coding completion speeds, delivering over 1000 tokens per second on ultra-low latency hardware while maintaining strong reasoning capabilities, significantly outperforming the standard GPT-5.3-Codex in real-time coding tasks.

**Key Points:**
- GPT-5.3-Codex-Spark is an ultra-fast model for real-time coding in Codex, optimized for near-instant response times.
- Codex-Spark delivers over 1000 tokens per second when served on specialized ultra-low latency hardware built by Cerebras.
- The model is being shared as a research preview with ChatGPT Pro users, alongside a partnership with Cerebras to ramp up datacenter capacity.
- Codex-Spark has a smaller 128k context window and is text-only during the research preview, operating under its own rate limits.
- Benchmarks show Codex-Spark achieving higher accuracy than GPT-5.3-Codex at shorter task durations (e.g., around 2 minutes) in SWE-Bench Pro, although Codex eventually pulls ahead in overall accuracy at longer durations.
- The performance difference highlights a trade-off: Codex-Spark prioritizes speed for interactive work, while Codex maintains higher performance for longer, more complex reasoning tasks.
- The partnership with Cerebras leverages wafer-scale custom hardware to achieve this significant speed improvement for real-time collaboration.

![Screenshot at 00:09: Comparison screen splitting between GPT-5.3-Codex \(left\) and GPT-5.3-Codex-Spark \(right\) showing Spark completing a complex software planning task significantly faster than the standard Codex model.](https://ss.rapidrecap.app/screens/gP-3S7ITWzY/00-00-09.jpg)

**Context:** OpenAI announced GPT-5.3-Codex-Spark on February 12, 2026, describing it as a smaller, faster version of GPT-5.3-Codex specifically designed for real-time coding experiences within the Codex environment. This development marks a milestone achieved in partnership with Cerebras, utilizing their custom wafer-scale hardware to drastically reduce latency for coding assistance.

## Detailed Analysis

OpenAI introduced GPT-5.3-Codex-Spark, a specialized, ultra-fast model for real-time coding in Codex, developed in partnership with Cerebras utilizing specialized low-latency hardware. Codex-Spark is optimized to feel 'near-instant' by delivering over 1000 tokens per second, a significant speed increase compared to previous models. While it retains high capability for real-world coding tasks, its primary focus is speed, meaning it keeps its default working style lightweight, making minimal, targeted edits without automatically running tests unless explicitly asked. The model is initially available as a research preview to ChatGPT Pro users, with usage governed by separate rate limits during this research phase. Benchmarks, like the SWE-Bench Pro graph (00:51), illustrate the performance trade-off: Codex-Spark achieves higher accuracy faster (e.g., 51.8% accuracy at 2.29 minutes), but the larger GPT-5.3-Codex model eventually surpasses it in overall accuracy at longer task durations (e.g., around 16 minutes). The presentation also contrasted this speed-focused approach with announcements from other frontier labs, such as Gemini 3 Deep Think's focus on reasoning and complex scientific domains, and MiniMax M2.5's performance on coding and agentic tasks, emphasizing that different models cater to different needs, with Codex-Spark targeting low-latency, interactive coding.

### Introduction of Codex-Spark

- Releasing a research preview of GPT-5.3-Codex-Spark, a smaller, ultra-fast model for real-time coding in Codex
- Marks the first milestone in the partnership with Cerebras announced in January
- Optimized to feel near-instant when served on ultra-low latency hardware, delivering over 1000 tokens per second.

### Availability and Details

- Rolling out today as a research preview for ChatGPT Pro users in Codex app, CLI, and VS Code extensions
- Runs on specialized low-latency hardware, usage governed by a separate rate limit adjustable based on demand
- Currently text-only with a 128k context window.

### Speed vs. Intelligence Trade-off

- Codex-Spark is tuned for speed, keeping its default style lightweight; it makes minimal edits and doesn't run tests automatically
- Performance comparison shows Codex-Spark faster at short durations but GPT-5.3-Codex achieving higher peak accuracy.

### Competitive Landscape Context

- Referenced performance of Gemini 3 Deep Think in reasoning tasks (e.g., 84.6% on ARC-AGI-2) and MiniMax M2.5 achieving SOTA in coding benchmarks (e.g., 80.2% in SWE-Bench Verified).

### Latency Improvements

- End-to-end latency reduction implemented via streamlined response streaming, optimized inference stack, and persistent WebSocket connection, resulting in 80% reduction in client/server roundtrip time.

![Screenshot at 00:00: Title slide introducing GPT-5.3-Codex-Spark as an ultra-fast model for real-time coding in Codex.](https://ss.rapidrecap.app/screens/gP-3S7ITWzY/00-00-00.jpg)
![Screenshot at 00:09: Side-by-side comparison demonstrating Codex-Spark \(right\) completing a complex project plan significantly faster than GPT-5.3-Codex \(left\).](https://ss.rapidrecap.app/screens/gP-3S7ITWzY/00-00-09.jpg)
![Screenshot at 00:51: SWE-Bench Pro graph comparing Accuracy vs. Task duration for Codex-Spark \(fastest at low duration\) against Codex and Codex-mini.](https://ss.rapidrecap.app/screens/gP-3S7ITWzY/00-00-51.jpg)
![Screenshot at 01:16: Section showing Gemini 3 Deep Think benchmark results, achieving 84.6% on ARC-AGI-2, highlighting external model performance context.](https://ss.rapidrecap.app/screens/gP-3S7ITWzY/00-01-16.jpg)
![Screenshot at 08:29: Graph showing SWE-bench Verified Score Evolution over time for various models, illustrating the rapid progress of the MiniMax M Series.](https://ss.rapidrecap.app/screens/gP-3S7ITWzY/00-08-29.jpg)
