A new best AI model every fortnight. Should you switch?
Three 'new best model' launches in three weeks. Why we stay with one maker, switch only within its family, and why how you use AI beats which one you have.

Three times in three weeks the news has said the best AI model in the world just changed. OpenAI launched GPT-6 Astra on 3 September. A start-up called TypeSafe released Jev on 15 September, claiming to be hundreds of times faster and cheaper than anything before it. And yesterday Anthropic shipped Claude Opus 5.5. Each launch came with a wave of videos and a confident person explaining why everything before it is now obsolete.
If you run a business and pay for one of these tools, the obvious question is whether you are on the wrong one. Our answer, having built our whole way of working on one of them: almost certainly not. Here is why we don't switch, what we do instead, and what it means for how you use AI in your own firm.
What actually launched
Three different things, which the headlines flatten into one race. Anthropic's own announcement says Opus 5.5 "performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5". The maker of the model we use every day says its cheaper sibling is about as good. Al Jazeera's report on Astra quotes OpenAI calling it "the most intelligent and aligned model in the world", which is what every launch says. And Jev, according to TechCrunch, is not a chatbot at all. It does not write text. It returns a decision from a list you defined in advance, with a confidence score, which is why it is so fast and why it cannot make things up. A useful tool, and a different one.
We have not tested Astra or Jev, so we cannot tell you whether the videos are right. We can tell you that a fortnight of "new best model" coverage is now normal, and it will not stop.
Why we don't switch providers
Because we tried once, and it hurt. Early in building the first version of our platform we moved from Claude to Gemini for a stretch. Same project, same brief. Gemini had its own ideas about how the thing should be structured. It reorganised decisions we had already settled, produced second copies of files it could not see, and by the time we noticed how far it had drifted we think it had cost us two months. Not because Gemini was bad. Because a tool you have worked with for months has absorbed your judgement about the project, and the new one starts from zero.
That was the era people now call vibe coding, and we don't build that way any more. The lesson held anyway. The value is not in the model. It is in everything around it: the tools you have configured, the habits you have built, the failure modes you have learned to spot. Every switch throws that away and restarts the clock.
So we stay with Anthropic. We use Claude Desktop and Claude Code every working day, and we built SDUK Studio, our own platform, around them. We cannot think of a reason we would move. Someone who started on OpenAI or Google has every reason to make the same decision about theirs.
Move within the family, not between families
Where we do change is between models from the same maker, and a solicitors' practice is the closest everyday picture of how. A partner costs far more per hour than a junior solicitor, so the partner does not do all the work. They set the approach and delegate what is safe to delegate, with instructions. The junior may be a little slower and costs a great deal less. That is how we run it. We plan with the more expensive model, a cost we are happy to pay for judgement, and when the plan is safe to hand over we build with the cheaper one. Opus 5.5 arrived yesterday and this morning it is reviewing a client project in a separate session. If it earns the job, it keeps it.
Now and then even the partner wants a barrister's opinion, at five or ten times the cost. Sensible for the case that matters, not for most Tuesdays. Our version is the blind test. For design work we give the same written brief to two models, blind, and choose the result without knowing which produced which. Over eight rounds this summer one model won most of the expansive first drafts and the other won the judgement-heavy work, and we only found out which was which after scoring. On client projects we have run the same blind test across three models. We cannot publish client material, so we built a public version of the same idea: Battle of the Bots, nine models from three makers predicting the same football matches from the same brief, marked every Tuesday. Test the model on your own work and let the score decide.
What you have today is enough
Whatever provider you use, the tool you have now is better than the one you had in June, and it will be better again by Christmas. It is also already very good.
The gap between two businesses using AI is far wider than the gap between the two models they happen to use. How well you use it beats which one you have got.
So what should a business use it for?
One test covers most cases: use AI where you can judge the output. Where you can't, be careful.
Contracts are the example we get asked about most. If you have a standard agreement you have used for years and need it adjusted for a new client, an AI will do the edit well and you will know whether the result is right, because you know the document. Ask it to summarise a contract someone has sent you or to explain a clause in plain English, and the same holds. You are the judge and it is the clerk.
Drafting a new agreement from scratch and signing it unreviewed is the other side of the line. The draft will read fluently whether or not it is right, and if you could not have written it you cannot check it. That is the same failure as the impressive prototype that falls over in its first month. It just happens in Word instead.
Meeting notes crossed this line a few years ago. An AI note-taker was a novelty then; now most calls we join have one, and they are more accurate and do more than they used to. From next week we will publish a short post every Thursday on one thing a firm can do with AI now, starting with contracts. Wednesdays stay for the main post.
How we build with this
A client's software built on Studio inherits our one-family decision without inheriting the churn. When a client's system uses AI for a task, the model it calls is a setting, not a rebuild. Every prompt is tested against real material and changed only when the evidence says so. When a better model arrives, we change the setting and rerun the tests. The client's product does not notice.
The best model in the world will change again in a fortnight. The best set-up in your business changes when you decide it should.