Context window comparison, 2026

A context window caps how much input a model can attend to in a single request. Max output caps how long a single response can be. The two limits are separate, and providers publish them separately. This page collects both numbers for the current models, verified against provider documentation.

Window size also sets the ceiling on cost per request. Input billing scales with the tokens you send, so a model that accepts a million tokens can also bill you for a million tokens in one call. A bigger window is a bigger budget ceiling, not a discount. See the pricing pages for per-token rates.

Every current model, side by side

Verified 2026-08-10.

ModelContext window (tokens)Max output (tokens)
GPT-5.6 Sol1,050,000128,000
Gemini 2.5 Pro1,048,57665,536
Gemini 2.5 Flash1,048,576
GPT-5.51,000,000
Claude Fable 51,000,000128,000
Claude Opus 51,000,000128,000
Claude Opus 4.81,000,000128,000
Claude Sonnet 51,000,000128,000
Claude Sonnet 4.61,000,000128,000
Kimi K31,048,576
DeepSeek V4 Flash1,000,000384,000
DeepSeek V4 Pro1,000,000384,000
Grok 4.31,000,000
Grok 4.5500,000
GPT-5400,000128,000
Kimi K2.6262,144
Grok 4256,000
Claude Haiku 4.5200,00064,000
GPT-4o128,00016,384

A dash means the provider does not publish a verified max output figure that we track for that model.

Will your input fit?

Paste text or enter a token count to check it against any window in the table.

Are there models with a 10 million token context window in 2026?

Not among the models verified here. As of 2026-08-10, no model in the table above ships a 10,000,000-token window. The largest verified windows are around one million tokens: GPT-5.6 Sol at 1,050,000, Gemini 2.5 Pro, Gemini 2.5 Flash, and Kimi K3 at 1,048,576, and GPT-5.5, the current Claude models, DeepSeek V4, and Grok 4.3 at 1,000,000. If you see a 10M claim, treat it as unverified marketing until the provider documents the number in its API reference. We will update this table when a provider does.

Two cautions before you fill a big window

First, cost. A full one-million-token prompt is billed as one million input tokens. On Gemini 2.5 Pro the per-token price itself rises once a prompt passes 200,000 input tokens, which is easy to miss when budgeting long-context workloads. Details are on the Gemini 2.5 Pro page.

Second, quality. Long prompts can degrade retrieval of facts placed in the middle of the window, an effect well documented in long-context evaluations. A window that accepts your document does not guarantee the model uses all of it equally well. When accuracy matters, keep the key material near the start or end of the prompt and test at your real prompt length.

Model pages

For per-token rates on these models, see API pricing. To count tokens in your own text, start at the token calculator.