In short
- Claude or ChatGPT: neither wins everywhere. After five new models in September, OpenAI and Anthropic each claim the lead, but on different tasks.
- So choose per task, and test on your own work rather than on a comparison table.
- What lasts is your own data, processes and instructions. Keep those under your own control, and switching is not a project.
Which is better, ChatGPT or Claude? The short answer: it depends on the task, and it shifts with every new release. In September 2026, OpenAI and Anthropic released five new models between them, and both publish tests on which their own model comes out on top. Those tests do not contradict each other; they measure different things.
For an organisation, the more useful question is therefore: can we switch models tomorrow without starting over? You can, as long as your data, processes and instructions are not locked into one tool. Below you can read what was released, where the differences lie and how to choose for yourself.
What was released in September
- 1 September. Anthropic releases Claude Fable 5.1, its most capable model for complex, long-running work.
- 3 September. OpenAI launches GPT-6 Astra, the first model of a new generation, with a focus on operating a computer on its own, software and professional work.
- 22 September. OpenAI adds the faster, cheaper GPT-6 Sol and GPT-6 Luna. The same day, Anthropic releases Claude Opus 5.5, which according to Anthropic performs at the level of Fable 5.1 on most work and costs 40% less than its predecessor Opus 5 on typical workloads.
Also on 22 September, Microsoft brought both Claude Opus 5.5 and GPT-6 Sol to Copilot, including in Word, Excel and PowerPoint. Microsoft noted that no single model is ideal for every task (Microsoft, 2026).
Claude vs ChatGPT: where is the difference?
The comparison below is based on what the makers publish themselves. These are vendor claims, not independent measurements.
| Claude (Anthropic) | ChatGPT (OpenAI) | |
|---|---|---|
| Latest top model | Opus 5.5 | GPT-6 Astra |
| Strong according to the maker | Agentic coding, knowledge work, long and complex tasks | Operating a computer on its own, professional work, documents and presentations that follow your own templates |
| API price of top model, per million tokens | $4 input, $20 output | $10 input, $50 output |
| Cheaper variants | Sonnet and Haiku 5.5 to follow in the coming weeks, according to Anthropic | GPT-6 Sol ($2 / $10) and Luna ($0.10 / $0.50) |
| In Microsoft Copilot | Opus 5.5 | GPT-6 Sol |
How close they are shows in the figures Anthropic itself publishes. On a test of coding tasks in the terminal, Opus 5.5 scores 66.4% against 57.9% for GPT-6 Astra. On a test of business workflows across connected apps, Astra scores 41.4% against 40.0% for Opus 5.5 (Anthropic, 2026). Which model "wins" depends on which task you measure.
Anthropic itself adds that at this level, benchmark margins say less and less about differences in practice. That is a useful warning for anyone who wants to base a choice on a comparison table, including this one.
Why the choice of one model matters less than you think
Whoever picks a model today will make that choice again in a few months. That is not a problem. It only becomes one when switching means starting over. And in many teams, more is tied to a single tool than they realise:
- Knowledge in scattered chats. Someone has worked out a good prompt for quotes, but it lives in a personal account. If the team switches, that knowledge is gone.
- Data with the vendor. Documents, brand guidelines and examples have been uploaded to a tool you cannot easily get them out of.
- Ways of working built around one screen. The team knows the buttons of one product, but the process behind it is not written down anywhere.
Microsoft makes the same point in its announcement: the choice of model is only part of a good answer. Just as important is what the model knows about your work.
What you do keep in your own hands
- Your data. Customer data, product information, brand guidelines and past campaigns. Keep them in your own environment, structured and usable by any model.
- Your processes. Which steps does your team take, where are the checks, what counts as a good result? Write it down, independent of the tool.
- Your shared instructions. Prompts, examples and writing rules that work, in a place the whole team can reach, not in someone's chat history.
Get those three in order, and a new model is good news instead of a migration project. You try it on your own tasks and keep it if it is better.
How to choose between Claude and ChatGPT yourself
A comparison on your own work tells you more than any benchmark:
- Pick three recurring tasks that matter to you, such as drafting a quote, summarising a report or writing product copy.
- Give both models the same input, with the same instructions and examples.
- Judge on fixed criteria: is it correct, is it usable straight away, how much correction does it need, and what does it cost at your volume?
- Repeat this with every new release. With your tasks and criteria written down, it takes an hour, not a project.
Cheaper is news too
September's releases are not only about capability. Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, against $5 and $25 for Opus 5. OpenAI cut the API prices of Sol and Luna by 50% compared with their predecessors (OpenAI, 2026).
A caveat when comparing: a price per token is not the same as the cost per task. One model uses more tokens for the same task than another, and both makers publish cost-per-task comparisons that favour their own model. For a single question the difference is negligible. It becomes interesting at volume: thousands of product descriptions, a daily analysis of your data, an assistant the whole team uses. Applications that were too expensive to automate last year may not be any more.
When the choice does matter
Sometimes the choice is already largely made. If your organisation works in Microsoft 365 and has Copilot, both models are available there, depending on licence and region. If your organisation has rules about where data may be stored, those help decide which tool qualifies. And for a specific, demanding task, such as large-scale coding work, one model can clearly fit better. Test that specifically.
How we approach it
In an AI workplace for a client, we use the tool that works best for each task, and data and instructions live in the client's own environment. That makes a new model a matter of testing and switching, not starting over.
Not sure which model suits your work? Book a free conversation and we will test it on your own tasks rather than on a benchmark.
Also read Using AI agents safely: five questions before you start on letting these models work on their own, and EU AI Act 2026: a delay for high risk, not for transparency on the European rules.
Sources
- Anthropic (22 September 2026). Introducing Claude Opus 5.5.
- Anthropic (1 September 2026). Introducing Claude Fable 5.1 and Claude Mythos 5.1.
- OpenAI (3 September 2026). GPT-6 Astra: A new generation of intelligence.
- OpenAI (22 September 2026). Introducing GPT-6 Sol and Luna.
- Microsoft (22 September 2026). More Models, One Copilot.
