AI & Automation · 3 October 2026 · 6 min read
The model is the most replaceable part of your AI project
On broad leaderboards, the leading AI models are now bunched closely together. The research points elsewhere for what separates businesses getting value from AI: a measured goal, redesigned work and a team that uses it. Here’s how to stop chasing and start measuring.

Every week brings a new AI model, a new leaderboard crowning the “best” one and a fresh worry that you’re falling behind. So does the newest model decide whether AI pays off for your business? On its own, rarely. On broad leaderboards the top models are now bunched closely together1, and the companies getting real value from AI stand out for how they measure results and change their work, not for running the newest model2.
We’ll admit it: we chased the news cycle too. Every launch felt like a reason to start again. It felt like progress. Mostly, it wasn’t.
What is the best AI model for business?
There’s no clear winner: the top models are closely bunched, although their strengths differ from task to task. Stanford’s AI Index, one of the most widely cited independent reviews of AI, found that top model performance is converging, with four companies clustered within 25 Elo points of each other on Arena, a leaderboard where people vote between two models’ answers1. (Elo is a rating system borrowed from chess. By the standard Elo formula, a 25-point gap means the higher-rated model is preferred only slightly more than half the time.) The same report says competition is shifting towards cost, reliability and domain-specific performance1.
The leaderboards themselves are shaky. Tests designed to stay hard for years are now beaten within months, and audits have found invalid-question rates as high as 42% on some widely used benchmarks1. There’s a running joke online (opens in a new tab) that every AI benchmark chart presents a lead of a fraction of a point as a landslide. It’s funny because it’s true.
And models are uneven. AI models can now win a gold medal at the International Mathematical Olympiad, yet even the top model reads an analogue clock correctly only about half the time3. A leaderboard can’t tell you how a model handles your quotes, invoices or customer emails. Only testing on your own work can.
Why do so many AI projects fail to pay off?
The research points to execution more than the model: clear goals, measurement and changing how the work gets done.
AI use is already widespread: Stanford’s AI Index reports that 88% of surveyed organisations used AI in 20254. Yet in McKinsey’s 2026 survey of 1,719 respondents, only 37% of organisations reported any positive contribution to profit (EBIT) from AI, essentially flat on the year before2. Meanwhile, 80% said AI had improved their own productivity2. People are reporting productivity gains much faster than businesses are reporting bottom-line gains.
Boston Consulting Group put it bluntly after surveying 152 chief executives of companies with revenues of at least $500 million: “The core problem is not technology. It is execution.”5 Only 14% of those CEOs had clearly defined the P&L impact for all of their AI initiatives5. These are big companies, but we think the lesson applies at any size. And Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027, because of escalating costs, unclear business value or inadequate risk controls. Its analyst described most of today’s agentic AI projects as early experiments that are “mostly driven by hype and are often misapplied”6.

What do businesses getting real value from AI do differently?
They set clear goals, measure the results and change the work around AI.
McKinsey’s “high performers” (the roughly 6% of respondents who attribute at least 5% of EBIT to AI and call its impact significant)2 don’t just add AI to what they already do. Nearly three-quarters of them have fundamentally redesigned workflows because of AI, compared with just one-quarter of everyone else, and they are twice as likely to have defined processes for measuring the impact of their AI work2. BCG found the same pattern: high performers are roughly seven times more likely to redesign workflows5.
Think of how businesses buy computers. Chipmakers announce faster processors every few months, yet no sensible IT manager replaces the company’s laptops each time one launches. They buy machines that do the job, run them for years and upgrade on their own schedule. The value comes from the work done on them, not the spec sheet. AI models deserve the same treatment: land one solution that moves a real number, enjoy the gains, and upgrade when your results call for it, not when the launch calendar does.

Will my AI project be out of date when a better model comes out?
Not if it’s built so the model can be swapped. But don’t expect to set it and forget it.
Model makers retire older models on a schedule. Anthropic, which makes Claude, says that as more capable models launch it “regularly retires older ones” and that applications “may need occasional updates”7. OpenAI gives every deprecated model a shut-down date8. So any AI solution needs light maintenance, whichever model it runs on.
That’s fine if you plan for it. Keep the parts that carry the value (your process, your data, your connections to other systems, your team’s habits) separate from the model itself. Then a new model is more likely to be a component you test and swap in than a full rebuild. Anthropic’s own advice is to carry out “thorough testing of your applications with the new models well before the retirement date”7. The simplest way to do that is to keep a small set of real examples from your business, with the answers you’d expect, and run every candidate model against them.
When is it actually worth switching AI models?
When your own results say so, not when a leaderboard does. Switching is worth a look when:
- a new capability makes a task possible that wasn’t before;
- you can get the same quality for clearly less money;
- your current model is being retired;
- your own tests, on your own examples, show a clear gain on the number you care about.
It isn’t worth it because of a launch-day headline, a benchmark chart or a demo video. Every switch costs testing time and attention, and the gain is often marginal. Meanwhile, every month spent waiting for the “right” model is a month without the time or money a working solution could already be saving you.

Should a small business commit to one AI platform, like ChatGPT or Claude?
For many small teams, we think so: depth reduces complexity, as long as switching stays possible.
In the interest of transparency: NightForward builds mainly on Anthropic’s Claude. That’s not because it tops every benchmark. The top spot has changed hands multiple times since early 20251, which is rather the point of this article. It’s because knowing one platform deeply lets us deliver working solutions faster. Every platform creates some switching costs, so we design the workflow around it so that the model can be changed where practical. If a different model would clearly serve a client’s goal better, we’ll say so. Either way, we test on the client’s own work before we recommend anything.
Researched and drafted with AI tools, fact-checked against the sources below, and edited by Kringle L.
Sources
- Stanford HAI, Technical Performance, The 2026 AI Index Report (opens in a new tab) (13 April 2026)
- McKinsey & Company, The state of AI in 2026: On the road to ROI (opens in a new tab) (25 August 2026)
- Stanford HAI, The 2026 AI Index Report (opens in a new tab) (13 April 2026)
- Stanford HAI, Economy, The 2026 AI Index Report (opens in a new tab) (13 April 2026)
- Boston Consulting Group, Nearly Nine in Ten CEOs See Some Cost or Revenue Benefits from AI in Targeted Areas, But Most Are Struggling to Scale It (opens in a new tab) (22 July 2026)
- Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (opens in a new tab) (25 June 2025)
- Anthropic, Model deprecations (Claude Platform Docs, accessed October 2026) (opens in a new tab)
- OpenAI, Deprecations (API docs, accessed October 2026) (opens in a new tab)