Why open weights do not mean free to run, and the API first strategy that actually keeps your AI costs low without owning any infrastructure.
The Hook and the Direct Answer
Every few weeks, a powerful new AI model gets released as open, and a wave of business owners have the same thought: this is my chance to own my AI, cut out the subscriptions, and finally be independent. It is a seductive idea. It is also, for almost every small business, a trap.
Here is the direct answer to the question in the headline. No, your small business almost certainly should not self host an open AI model to save money. Open weights are a licensing story, not a cost story. The fact that you can download a model does not mean you can afford to run it. Today’s frontier open models are so large that running one requires serious hardware, real engineering talent, and ongoing maintenance that will cost you far more than simply paying for access through an API. The smarter strategy is API first: consume the best models on demand, keep your setup flexible, and let the fierce price competition between providers work in your favor.
This week made the point almost comic. Moonshot released Kimi K3, a genuinely open, frontier class model with 2.8 trillion parameters. The developer community’s reaction was not celebration about running it at home. It was a collective laugh about how impossible that would be. You could download the weights for free and still have no realistic way to run them without a data center. That gap between open and runnable is the entire lesson, and understanding it will save you from an expensive mistake dressed up as a smart one.
Key Takeaways
- Open weights mean a model is freely licensed, not that it is free or practical for a small business to actually run.
- Today’s frontier open models are enormous, requiring hardware, engineering, and maintenance costs that usually exceed simply paying for API access.
- An API first strategy lets you consume the best available model on demand without owning or maintaining any infrastructure.
- Keeping your workflows model neutral lets you switch providers quickly and capture price drops as they happen.
- The real benefit of open models to a small business is competition, which drives down the price of the APIs you actually use.
The Problem
I understand the appeal of self hosting deeply, because I have felt it myself. There is something powerful about the idea of owning your tools instead of renting them. Subscriptions feel like a leak in the boat, a monthly reminder that you depend on someone else. So when a capable model becomes available for free download, the instinct is immediate: bring it in house, run it yourself, and stop paying the toll. Independence. Control. No more vendor holding the keys.
The problem is that this instinct is built on a misunderstanding of what open actually means. Open weights means the model’s parameters are released under a license that lets you use them. It does not mean the model is small, cheap, or practical to operate. And the frontier open models arriving now are staggeringly large. We are talking about models with trillions of parameters, systems designed to run on clusters of specialized, expensive hardware in data center conditions. The idea of running one on the kind of equipment a normal business owns is not just difficult. It is a punchline.
And the hardware is only the beginning. Even if you acquired the machines, you would need engineering talent to deploy and maintain the model, to keep it running, to handle updates and failures. That is a salary, or several. You would be responsible for uptime, security, and performance. You would be running a small AI infrastructure operation, which is a business in itself, and probably not the business you actually want to run. All of this to avoid a subscription that, in most cases, costs a tiny fraction of what self hosting would.
There is a final twist that makes self hosting even worse. These models are being replaced every few weeks. So even if you did stand up the infrastructure to run today’s open frontier model, it would be behind the curve almost immediately, and you would be stuck maintaining an aging system while the world moved on. But what if you never had to own any of it, and could always run on a current model instead?
The Evidence
The evidence from this week is unusually clear, because the numbers speak for themselves.
Start with Kimi K3. Moonshot released it as an open, frontier class model with 2.8 trillion total parameters, a million token context window, and a sophisticated mixture of experts architecture. It is a genuinely impressive, genuinely open system. And the community’s response tells you everything. Developers were excited about the open release in principle, but they openly joked that running a model of this scale locally would be impractical to the point of absurdity, that even on capable consumer hardware the throughput would be unusably low. The gap between having the weights and being able to run them is not a small technicality. It is a chasm.
Now consider what happened on the demand side. Kimi K3 generated so much interest through its hosted service that Moonshot had to suspend new subscriptions within 48 hours, citing hardware overload. Sit with that for a moment. The company that built the model, with all its resources and infrastructure, hit capacity limits serving it. If Moonshot struggles to serve this model at scale, the notion that a small business could run it in a back office is not realistic.
Meanwhile, the API side of the market is racing in your favor. Alibaba priced its Qwen 3.8 preview at roughly a tenth of standard rates. Anthropic released Opus 5 at near flagship quality for about half the cost of its top tier. Providers are competing aggressively on price, and every new open model that arrives adds pressure to that competition. This is the crucial insight. Open models do not help you most by being something you run. They help you most by being competition that drives down the price of the APIs you actually use.
So the conventional narrative, that open equals free equals independence, collapses under its own weight. Open is real and valuable, but the value flows to you through lower API prices, not through a server in your closet.
The Solution
Here is the approach that actually works, and it is the one I use and teach. Rent, do not buy. Go API first.
An API first strategy means you access AI models on demand through their application interfaces, paying only for what you use, without owning or maintaining any of the underlying infrastructure. Someone else runs the data center. Someone else handles the hardware, the uptime, the security, the updates. You simply send your requests and get your results, and you always run on a current model because the provider keeps it current for you. This is the opposite of the self hosting fantasy, and it is dramatically cheaper and simpler for almost every business.
The magic of this approach is the flexibility it gives you. When you consume models through APIs and keep your workflows model neutral, you are not married to any single provider. If one lab drops its prices, you move volume there. If a new model proves better for your use case, you switch with a small change instead of a rebuild. You become a free agent in a market where providers are fighting to win your business by lowering prices. Every price war, every new open release, every competitive move works to your benefit, because you are positioned to capture it instantly.
Contrast that with the self hosted business. It is locked to the model it deployed, running aging infrastructure, unable to easily benefit from the next price drop or the next better model, and carrying all the cost and risk of ownership. The API first business is light, flexible, and always current. The self hosted business is heavy, brittle, and always one step behind.
This is what changed my thinking completely. I stopped seeing subscriptions as a leak and started seeing them as the smartest possible deal, a way to access world class intelligence without any of the burden of owning it. The independence I wanted was never going to come from a server. It comes from keeping my setup flexible enough that no single provider can trap me. Let the giants fight the expensive hosting war. I just harvest the results through my API bill.
Practical Steps
1. Cross self hosting off your plan. Review your AI strategy and remove any assumption that depends on running a large model yourself. For nearly every small business, that path costs more and delivers less than the alternative.
2. Adopt an API first mindset. Commit to accessing AI models on demand through their interfaces rather than owning infrastructure. Let the providers carry the cost and complexity of running the models while you focus on using them.
3. Identify which tasks truly need frontier quality. Sort your workflows and route only the genuinely demanding tasks to premium models. Everything else can run on cheaper options, which keeps your API costs low without sacrificing results where it counts.
4. Keep your workflows model neutral. Structure your integrations so you can swap the underlying model with a configuration change rather than a rebuild. Portability is what lets you capture every price drop the moment it arrives.
5. Track provider price movements. Watch for the price cuts that follow major open model releases. When a new open frontier model lands, API providers often lower their prices in response, and you want to be ready to move.
6. Review your provider mix quarterly. Every quarter, compare your providers on price and performance and shift volume toward the best value option. Treat this like re shopping insurance, a routine that keeps you on the best deal available.
7. Reinvest the savings into customers. Take the money you would have burned on hardware and engineering and put it into the parts of your business that actually serve customers. That is where the return lives, not in a server rack.
Frequently Asked Questions
What is the difference between open weights and free to run?
Open weights means a model’s parameters are released under a license that lets you use them, while free to run would mean you can operate it at no meaningful cost. Frontier open models are enormous and expensive to run, so open and free to run are very different things.
How much does it cost to self host a large AI model?
Far more than most owners expect. Beyond the specialized hardware, you need engineering talent to deploy and maintain it, plus ongoing costs for power, security, and uptime. For nearly all small businesses, these costs vastly exceed simply paying for API access to the same class of model.
What does API first mean for a small business?
API first means you access AI models on demand through their interfaces and pay only for what you use, without owning any infrastructure. The provider runs and maintains the model, so you always have current capability with none of the hosting burden, cost, or risk.
Do open source AI models actually help my business?
Yes, but usually not by being something you run. Their biggest benefit is competition. Every capable open model that launches pressures commercial providers to lower their prices, which reduces the cost of the APIs you already use, so you benefit as a customer, not a host.
When would self hosting ever make sense?
Only at very large scale, with strict data control requirements, or with genuine in house engineering capacity, where the volume and constraints justify the cost. For the vast majority of small businesses, that threshold is far away, and an API first approach remains cheaper and simpler.
The Close
Every few weeks the cycle repeats. A powerful open model drops, and a fresh wave of owners dream about finally owning their AI and escaping the subscriptions. I understand the dream, because independence is a worthy goal. But the path to it was never a server humming in the back office.
This week’s frontier open model has trillions of parameters and made its own creator hit capacity limits. The idea of running it yourself is not a strategy. It is a punchline. The real power move is quieter and far more profitable. Stay light. Consume the best models on demand. Keep your setup flexible enough that no provider can trap you. And let the relentless competition between labs, fueled by every open release, keep pushing your costs down.
Rent, do not buy. Let the giants spend fortunes fighting the hosting war. You have a business to run, and the smartest thing you can do is stay free enough to always be on the best deal available. That is the independence that actually matters.
Jonathan Mast is the founder of White Beard Strategies, where he helps tens of thousands of entrepreneurs make practical, profitable AI decisions. He is the creator of the Perfect Prompt Framework, a speaker on applied AI for business, and a firm believer that the smartest technology strategy is the one that keeps you light, flexible, and free to move.