GPT 5.3 Codex Spark: Its Crazy Fast
Quick Overview
GPT-5.3-Codex-Spark achieves drastically faster, near-instantaneous coding completion speeds, delivering over 1000 tokens per second on ultra-low latency hardware while maintaining strong reasoning capabilities, significantly outperforming the standard GPT-5.3-Codex in real-time coding tasks.
Key Points: GPT-5.3-Codex-Spark is an ultra-fast model for real-time coding in Codex, optimized for near-instant response times. Codex-Spark delivers over 1000 tokens per second when served on specialized ultra-low latency hardware built by Cerebras. The model is being shared as a research preview with ChatGPT Pro users, alongside a partnership with Cerebras to ramp up datacenter capacity. Codex-Spark has a smaller 128k context window and is text-only during the research preview, operating under its own rate limits. Benchmarks show Codex-Spark achieving higher accuracy than GPT-5.3-Codex at shorter task durations (e.g., around 2 minutes) in SWE-Bench Pro, although Codex eventually pulls ahead in overall accuracy at longer durations. The performance difference highlights a trade-off: Codex-Spark prioritizes speed for interactive work, while Codex maintains higher performance for longer, more complex reasoning tasks. The partnership with Cerebras leverages wafer-scale custom hardware to achieve this significant speed improvement for real-time collaboration.
Context: OpenAI announced GPT-5.3-Codex-Spark on February 12, 2026, describing it as a smaller, faster version of GPT-5.3-Codex specifically designed for real-time coding experiences within the Codex environment. This development marks a milestone achieved in partnership with Cerebras, utilizing their custom wafer-scale hardware to drastically reduce latency for coding assistance.
Detailed Analysis
OpenAI introduced GPT-5.3-Codex-Spark, a specialized, ultra-fast model for real-time coding in Codex, developed in partnership with Cerebras utilizing specialized low-latency hardware. Codex-Spark is optimized to feel 'near-instant' by delivering over 1000 tokens per second, a significant speed increase compared to previous models. While it retains high capability for real-world coding tasks, its primary focus is speed, meaning it keeps its default working style lightweight, making minimal, targeted edits without automatically running tests unless explicitly asked. The model is initially available as a research preview to ChatGPT Pro users, with usage governed by separate rate limits during this research phase. Benchmarks, like the SWE-Bench Pro graph (00:51), illustrate the performance trade-off: Codex-Spark achieves higher accuracy faster (e.g., 51.8% accuracy at 2.29 minutes), but the larger GPT-5.3-Codex model eventually surpasses it in overall accuracy at longer task durations (e.g., around 16 minutes). The presentation also contrasted this speed-focused approach with announcements from other frontier labs, such as Gemini 3 Deep Think's focus on reasoning and complex scientific domains, and MiniMax M2.5's performance on coding and agentic tasks, emphasizing that different models cater to different needs, with Codex-Spark targeting low-latency, interactive coding.