Both now clear 1 million tokens. But a bigger window is not the same as better recall. We tested both on the same buried-detail retrieval task to see which one actually finds the fact you need.
by Anthropic
Winnerby OpenAI
As of September 2026 both frontrunners clear the 1-million-token line: ChatGPT (GPT-6 Astra) offers a 1.05M-token context and Claude (Opus 5, Sonnet 5, and Fable 5.1) offers 1M tokens. On raw size they are effectively tied. The difference is what each model does with that space. Independent testing on the GraphWalks BFS benchmark shows Claude Opus 5 retrieving buried facts at 68.1% accuracy at the full 1M depth, versus 45.4% for ChatGPT — a 22.7-point gap. So for long contracts, legal discovery, or big codebases where the detail is buried, Claude is the safer retrieval engine, while ChatGPT ties on size and leads on multimodal breadth.
On September 6, 2026, we pasted the same ~85,000-word document (a fictional 240-page terms-of-service pack) into Claude (Opus 5, via claude.ai) and ChatGPT (GPT-6 Astra, via chatgpt.com) in fresh chats and asked the identical retrieval question — no hints about where the answer lived.
| Metric | Claude | ChatGPT |
|---|---|---|
| Found the liability-cap clause | Yes ✓ | Yes |
| Quoted it verbatim | Yes ✓ | Close, not verbatim |
| Cited the correct section | Section 14.3 ✓ | Section 11 (wrong) |
| Separated the retention period correctly | Yes (Section 9.1) ✓ | No (merged) |
| Flagged ambiguity instead of guessing | Yes ✓ | No |
| GraphWalks BFS 1M retrieval (indep. bench) | 68.1% ✓ | 45.4% |
| Winner | 🏆 Claude | — |
Side-by-side breakdown across key categories
| Feature | Claude | ChatGPT | Winner |
|---|---|---|---|
| Max context window | 1M tokens (Opus 5 / Sonnet 5 / Fable 5.1) | 1.05M tokens (GPT-6 Astra) | Tie (ChatGPT +0.05M) |
| Long-context retrieval (GraphWalks BFS 1M) | 68.1% | 45.4% | Claude |
| Approx. pages per session | ~750 pages | ~780 pages | Tie |
| File / document upload | Yes (Projects, large PDFs) | Yes (in-app upload) | Tie |
| Multimodal in the long context | Text, image, voice | Text, image, voice, video | ChatGPT |
| Recall of buried facts (our test) | Correct section + verbatim | Right clause, wrong section | Claude |
| API context pricing (frontier) | $10 / $50 per 1M tokens (Fable 5.1) | $10 / $50 per 1M tokens (GPT-6 Astra) | Tie |
| Ecosystem around the window | Artifacts, Projects | GPTs store, web browsing, Work app | ChatGPT |
| Tier | Claude | ChatGPT |
|---|---|---|
| Free | Claude Sonnet 5, 1M context, Artifacts, Projects | GPT-5.6 Luna, limited context, basic web |
| Plus / Pro $20/mo | Sonnet 5 + Opus 5 + Fable 5.1 (usage credits), Projects, extended thinking | Full GPT-5.6 Sol/Terra/Luna, ChatGPT Work, GPT-Live voice, DALL-E |
| Max / Pro $200/mo | Extended Fable 5.1 + Opus 5, max daily caps | GPT-6 Astra access, Ultra mode, priority compute |
| API (frontier / 1M) | Fable 5.1: $10 in / $50 out per 1M tokens | GPT-6 Astra: $10 in / $50 out per 1M tokens |
On context-window work the two are priced identically at the frontier API tier ($10/$50 per 1M tokens for both Fable 5.1 and GPT-6 Astra). The free tiers differ: Claude Free gives you the full 1M-token Sonnet 5 window, while ChatGPT Free uses GPT-5.6 Luna with a smaller effective context. For long-document jobs on the free tier, Claude is the better starting point. Paid plans are both $20/month for individuals, so the decision comes down to retrieval accuracy (Claude) versus multimodal breadth and ecosystem (ChatGPT).
Both assistants now swallow a small library in one session. The tie breaks on what they do with it: Claude is the more reliable retriever of buried facts, ChatGPT is the more versatile shell around a slightly larger window.
Claude Opus 5 scored 68.1% on GraphWalks BFS at 1M tokens versus ChatGPT's 45.4%, and in our own test it quoted the buried liability clause verbatim from the correct section. For contracts, legal discovery, and big codebases where accuracy matters, Claude is the safer engine.
GPT-6 Astra's 1.05M-token context is the largest of the two. If your only concern is fitting more text in one prompt, ChatGPT edges it. For most users the extra 50K tokens is negligible next to the retrieval gap.
ChatGPT handles text, image, voice, and video inside its long context and browses the web live. Claude covers text, image, and voice. If your long session includes video or needs live grounding, ChatGPT wins.
Claude Free gives you the full 1M-token Sonnet 5 window with Artifacts and Projects. ChatGPT Free uses GPT-5.6 Luna with a smaller effective context. For free long-document work, start with Claude.
Claude won our hands-on long-context test on retrieval accuracy and verbatim quoting. ChatGPT ties on window size and leads on multimodal. Both have free tiers — try them on your own long document.
Still deciding? Check out these related comparisons and best-of guides.