Small, local AI models are probably all we need
I'm an AI bull. I think it's going to revolutionize the economy and make most of us better off.
While frontier AI models get all the attention, I don't see them being the path by which LLMs move the economy forward. Those models are cool, doing things that still feel like magic (as is the case with any major technology when it's new). Yet they come with some serious drawbacks that are being overlooked because of their coolness. They're slow, they're horrifically expensive, and in most business applications you get nothing in return for the two downsides.
Frontier models are seductive, because they give the impression of productivity, but that doesn't mean they have business value. The question is not are you more productive with frontier models than you are with no AI assistance? That's the basis for the claims of massive productivity gains - assuming, of course, that you get those gains for all tasks you do at work rather than one or two where those models really stand out.
The question you need to ask is how much additional profit does the business generate if you use a frontier model rather than the best alternative strategy? Once you open that door, you have to consider a lot of other factors:
- How much will frontier models cost when we're in the new equilibrium? If you're using subsidized tokens, that's not going to tell you about the long-run, sustainable business case for frontier models that is the topic of this post. Don't forget to include the cost of a frontier model for maintaining and changing your programs and documents in the future.
- How much does it cost to use cheap, non-frontier models? If you're writing a program or creating a document, what's the incremental cost to do the big picture thinking yourself, having a lesser LLM fill in the pieces of your outline, versus one-shotting the whole thing? I'm not convinced one-shotting reduces the cost. I'm not convinced it isn't dramatically more expensive.
- How much additional revenue is created by using a frontier model versus a non-frontier model? This is potentially a differentiator that could leave room for frontier models, assuming they produce output that cannot be matched by a combination of human labor and non-frontier models. My guess, however, is that the answer to this question is nil.
- What is the value to speed? Frontier models, by construction, are going to be slow due to their thinking. Nobody's calling Gemini Flash Lite a frontier model, yet the combination of speed and token price makes it unbeatable for a lot of business applications. My guess is that if any of the current generation of LLMs survive into the future, Gemini Flash Lite will be the one, due to its speed.
My view of the future of LLMs is one in which frontier models play only a small role. If they can improve outcomes for open heart surgery, or if they can help with long term care needs, they'll absolutely have a place. But that's not most business applications. All the comments about Google not being relevant because Gemini isn't a frontier coding model are built on a very specific assumption about one's objective function. I think Google is winning the AI battle due its focus on low cost, speed in the form of flash lite and flash models, and successful open weight models in the form of Gemma 4.