AI & Building
AI Loyalty
Why constantly switching models may be costing you more than you think.
Every few weeks there is a new leaderboard. A faster model, a smarter model, a cheaper model. The industry rewards novelty, and it rewards it loudly, because novelty is the easiest thing to announce. Businesses watching from the outside often mistake that novelty for progress.
The result is predictable. Teams jump from one model to another, rebuilding prompts, retraining staff, rewriting workflows, and rediscovering the same lessons they had already learned six months earlier. They are optimizing benchmarks. They are not compounding knowledge.
Imagine hiring a new employee every month. Each one slightly more talented than the last. None of them there long enough to understand your customers.
The continuity problem
Picture that hiring pattern for a moment. Every thirty days, someone new. On paper each hire is an upgrade — better credentials, faster on the keyboard, higher score on whatever test you administer. And yet the work gets worse. Nobody knows which client hates phone calls. Nobody remembers why the invoice template has that odd second page. Nobody has been present for enough mistakes to develop judgment about your particular business.
Eventually you realize the problem was never talent. It was continuity. The value of an employee is not their raw capability; it is their capability multiplied by their accumulated understanding of your world. The second term takes time to grow and resets to zero every time you replace them.
AI relationships work similarly, and I do not mean that as a metaphor stretched for effect. As your team learns a model's strengths, its limitations, its reasoning patterns, the phrasings it responds to and the ones it ignores, your organization develops something real. Institutional knowledge. It lives in your prompt library, your review habits, your sense of which tasks to hand over and which to keep.
What compounds
That knowledge compounds quietly and in several directions at once. The prompts get better, because each one is written by people who have seen a hundred outputs. The reviews get faster, because your editors know where to look first. The edge cases become understood rather than merely encountered. And trust increases — the human kind, the sort where a team stops second-guessing every output and starts using the tool with confidence, which is the moment any tool actually begins paying for itself.
None of that transfers cleanly to a new system. Some of it does. The habits of mind survive. But the specifics — the prompt that finally worked, the failure mode you learned to watch for, the calibration of how much to trust which kind of answer — all of that has to be rebuilt, and rebuilding it costs a quarter of quiet, unmeasured productivity while everyone pretends the migration was seamless.
Technology catches up
Here is the part that makes the constant switching especially poor economics. Technology catches up remarkably quickly in this field. Today's leader is often tomorrow's commodity. The gap between the best available model and the second or third best has, over the last several years, narrowed almost as fast as it has opened, again and again.
Which means the advantage you chase by switching is temporary and shrinking, while the advantage you abandon by switching is durable and growing. The organizations that consistently outperform are not, in my experience, the ones using the newest model. They are the ones running mature systems built on deep understanding of a model they chose deliberately and stayed with long enough to learn.
Don't confuse changing tools with improving capability.
This is not an argument against innovation
It would be foolish to read this as a case for ignoring what is happening. Evaluate the new thing. Experiment with it. Benchmark it honestly against your actual work rather than against a public test set that resembles nothing you do.
But hold the evaluation to a real standard. The question is not whether the new model is better in the abstract. It is whether it is better enough to justify resetting your institutional knowledge to a fraction of its current value. That is a high bar, and it is met occasionally — genuine step changes do happen. It is not met by a two-point improvement on a reasoning benchmark, and it is certainly not met by a headline.
A useful practice: keep a small, standing evaluation drawn from your own hardest real tasks. Run new models against it when they appear. Most of the time you will find the difference is smaller than the announcement suggested. Occasionally you will find it is not, and then you will switch with a clear head rather than out of anxiety.
Depth as strategy
The greatest competitive advantage available in this moment may not be finding the next model. It may be becoming exceptionally good at using the one you have already chosen — knowing its temperament the way a carpenter knows a particular saw, working with the grain instead of fighting it, and building processes around what it reliably does rather than what it might do someday.
That kind of depth is unglamorous. It does not make announcements. It accrues in small increments that are invisible from outside and obvious from inside, and one day you notice your team is producing work at a level that competitors with better models cannot match, and there is no single decision you can point to that explains it.
Technology evolves. Wisdom compounds. The businesses that understand the difference will usually move faster over the long run, and they will do it while spending less energy on migrations that nobody outside the company ever noticed.
If your team is midway through its third migration this year, it may be worth asking what is actually being pursued. That is a conversation I am always happy to have.