The Newest AI Model Is Rarely the One You Should Build On

The Newest AI Model Is Rarely the One You Should Build On

The newest model is almost never the one you should build on, and I will die on this small hill. Launch day is a party. Your product is a marriage. The two rarely want the same thing.

Let me explain the way I explain everything worth knowing — slowly, with food on the table, at the kiosk near the stage where the owner keeps a radio and a very strong opinion. Dusk is the honest hour there. The heat has let go, the lamp comes on, and you can finally see which bottles were dusty all along.

Scene one: the shiny bottle at the front

A young fellow comes to the kiosk asking for the new soda everyone is posting about. Bright label, big promises, arrived this morning. The owner sells it to him, of course. But then the owner leans over and says the thing that stuck with me: "Nzuri sana. But I don't stock my fridge with it yet. I watch who comes back."

That is the whole discipline of new AI models in one sentence. A fresh release is a claim, not a receipt. The benchmark chart is somebody's launch outfit — pressed, flattering, worn for one afternoon. What you actually need to know is boring: does it stay cheap when you use it a thousand times, does it behave the same on a Tuesday as it did in the demo, and does it fail in a way you can predict. New models are exciting precisely because nobody has answered those questions yet. Exciting is not the same as ready.

Scene two: the customer who knows exactly what he wants

An older woman comes for cooking fat. Not the premium tin, not the tiny sachet. The middle one, same brand, every week. She does not read the labels anymore. She has run her own test — her food, her fire, her family's tongue — and the test is finished. New arrivals do not tempt her, because she is not shopping for novelty. She is shopping for a result she already trusts.

This is what builders forget. You are not choosing a model. You are choosing a result, and the model is just the ingredient that gets you there. The right question is never "which model is smartest." It is "which model does my specific job, on my data, at a price my product survives, with output I can check." A slightly older model that nails your one job beats a genius that needs babysitting. The kiosk auntie is not behind the times. She has simply stopped confusing new with better — and she frees up her attention for things that actually change.

Scene three: the radio nobody replaces

The radio on the kiosk shelf is ancient. There are better ones. The owner will not swap it, and his reason is sharp: "This one I understand. When it coughs, I know why."

That coughing is the whole game. Every model has a personality — a way of being wrong that is uniquely its. One rambles. One invents confident nonsense about numbers. One goes quiet on the exact edge case your customer hits most. The value of a model you have lived with is not that it never fails. It is that you have learned its failures, built your guardrails around them, and can sleep. A new model resets that hard-won map to zero. Sometimes the upgrade is worth relearning the whole instrument. Often it is not, and the honest builder admits the switching cost out loud instead of pretending it is free.

The thesis, once the lamp is fully on

Here is the stubborn point all three scenes were walking toward: a builder's job is not to run the newest model. It is to run the most boring model that still does the job, and to keep a quiet eye on the shelf for the day a new one clears a real bar — not a benchmark, a bar you set from your own product's pain.

So set the bar before the next launch tempts you. Write down the three things your product actually needs a model to do. Keep ten of your own ugly, real examples — the messy customer message, the half-Swahili half-English voice note, the receipt with a smudge — and make that your private exam. When a new model drops, do not read the announcement first. Feed it your ten. If it does not beat the model you already trust on your paper, it does not get the fridge space. Applause is not a passing grade.

New models are candidates. Your ten hard examples are the interview. Hire on the interview, not the résumé.

A small lab note before you go: this is exactly why, in the Ni Biashara workshop, we keep a little private notebook of test cases that never leaves the machine — the same idea behind our Hapo Ndani experiments, where the point is a locked drawer you control, not a flashy new engine you rent and hope for. The model underneath can change with the season. The exam should not.

The best builders I know treat model launches the way the kiosk owner treats a new soda. Curious, never breathless. They taste it. They watch who comes back. And they remember that the shelf at dusk rewards the patient, not the first in line — because the customer tomorrow morning is not buying the newest thing. She is buying the thing that works.

Comments

Popular posts from this blog

Your Data Is Not Safe Just Because It Is in the Cloud

200 Megawatts: The Number Behind Every AI Product You Use

Stop Feeding the Lorry One Chapati at a Time: Model Routing and Costs for Small Teams