Open-weight models are now matching frontier performance at a fraction of the price, and this article answers the question most business owners have not asked yet: how do you actually move your work onto cheaper intelligence without breaking the thing that works?
The Number in Your Budget Is Already Wrong
I looked at our AI spend last week and felt something I had not felt in a while: embarrassment.
Not because the number was huge. Because the number was built on assumptions I made in January and never revisited. I had priced our workflows, our client deliverables, and our internal automation against what intelligence cost me six months ago. In this market, six months is a geologic era.
Here is the direct answer. Yes, AI is getting dramatically cheaper, and the savings are real and available to you right now. Epoch AI, a research institute that tracks this stuff rigorously, found that the price to reach a given level of model performance has been falling between 9x and 900x per year depending on the capability, with a median of roughly 50x per year, and that median jumped to about 200x per year when they looked only at data after January 2024. That is not a forecast. That is a measurement of what already happened.
But the second half of that answer matters more. Cheaper intelligence does not automatically become cheaper operations. It only helps you if you can actually move your work onto it. And most small businesses have quietly built themselves into a position where switching is expensive, scary, or simply invisible to them.
This week made the point loudly. Moonshot AI released Kimi K3, a 2.8 trillion parameter open-weight model that took the number one spot on Arena’s Frontend Code evaluation, ahead of Claude Fable 5, in blind developer testing. Full weights go public by July 27. Meanwhile DeepSeek lists V4-Pro at $0.87 per million output tokens, and Fireworks AI just raised $1.5 billion at a $17.5 billion valuation to serve open models fast.
My thesis is simple. The price of intelligence has collapsed, your plan has not been updated, and the only question that matters now is whether your switching cost is low enough to capture the difference.
Key Takeaways
- The price to achieve a given AI performance level has fallen at a median of roughly 50x per year, and faster since 2024, according to Epoch AI.
- Open-weight models from Moonshot, DeepSeek, and others now compete at or near the frontier while publishing prices a fraction of closed alternatives.
- Cheaper models only help you if your business can actually switch, which means your real constraint is switching cost, not model cost.
- The companies capturing these savings are running evaluations and staged rollouts, not swapping model names and hoping.
- You should be re-pricing your AI assumptions quarterly, because a number from January is already a stale number in July.
The Problem: You Cannot Save Money You Cannot Measure
Here is the honest version of the challenge. Most small business owners have no idea what their AI actually costs them, what it is doing, or what it would take to change it.
You have a ChatGPT subscription. Somebody on your team has a Claude account. Your CRM has AI features baked into a tier you upgraded last year. Your content tool has credits. Your customer service widget bills per conversation. None of it appears on one line item. None of it is tied to an output you can price. So when the cost of intelligence drops by an order of magnitude, nothing happens in your business, because there is no lever to pull.
I have been where you are. When I started building AI into White Beard Strategies, I picked the model that was best at the moment, wired it into everything, and moved on to the next problem, because I had a business to run. That was reasonable at the time. It also meant that when better and cheaper options showed up, I had no clean way to evaluate them. My prompts were tuned to one model’s quirks. My workflows assumed one provider’s behavior. I had built a switching cost into my own company without ever deciding to.
The other half of the problem is emotional, and I want to name it because nobody else will. Switching feels risky. The AI you have works. It knows your voice, more or less. Your team trusts it. The idea of moving to something called DeepSeek V4 Flash to save money sounds like the business equivalent of buying tires from a guy in a parking lot. That instinct is not stupid. It is just incomplete.
And there are legitimate complications. In April, House committees opened an investigation into Airbnb and Anysphere, the maker of Cursor, over their use of Chinese-developed AI models. Airbnb had used Alibaba’s Qwen for a customer service agent, and Anysphere disclosed that its Composer 2 model was built on Moonshot’s Kimi. If you handle regulated data, this is a real consideration, not a talking point.
So here is the reframe. Stop asking whether cheaper models are good enough. That question has no answer, because “good enough” is not a property of a model. Start asking a much better question: where in my business is the work simple enough that cheaper intelligence is obviously sufficient? That question has an answer, and it is usually “most of it.”
The Evidence: What Actually Happened This Year
Let me give you the specifics, because I do not want you taking my word for any of this.
One. The price decline is measured, not hyped. Epoch AI analyzed state of the art models across six benchmarks over three years and fit price curves to performance thresholds. Their finding: prices declined between 9x and 900x per year, median 50x, with the fastest declines occurring after January 2024. When they restricted the data to post-2024 models, the median rate rose from 50x to 200x per year. Their explanation includes models getting smaller and hardware getting more cost effective.
Two. The gap between open and closed pricing is enormous and public. DeepSeek publishes its API pricing openly. V4-Flash is $0.14 per million input tokens and $0.28 per million output tokens. V4-Pro is $0.435 in and $0.87 out. For comparison, Kimi K3’s API is $3 per million on cache-miss input and $15 per million on output. You do not need a consultant to tell you what a 17x difference in output pricing means for a workflow that generates a lot of text.
Three. A real company made the switch and published the numbers. Lindy, the AI assistant startup led by Flo Crivello, moved most of its managed-agent model traffic off Claude and Gemini paths onto DeepSeek v4 Flash, served by Atlas Cloud. On the migrated traffic, inference costs fell by about 90 percent. Their engineering write-up is refreshingly honest. They tried Kimi K2.5 first, it did well in offline evaluations, then failed with real users almost immediately. One customer said it felt like their assistant had had brain surgery overnight. They also found the same model scored differently depending on which provider served it. Their conclusion is the sentence I want tattooed on every business owner’s forearm: do not ask whether a cheaper model is as good, ask where it is good enough.
Four. The infrastructure money is following open models. Fireworks AI closed a $1.5 billion Series D at a $17.5 billion valuation on July 16, with participation from Nvidia, Index Ventures, and Lightspeed. The company reports exceeding $1 billion in annualized revenue, five times the prior year, and says it handles 40 trillion AI tokens per day. CNBC reported the driver plainly: finance executives are anxious about model costs and have begun directing their teams to consider open-source alternatives.
Five. Small businesses are spending, but not strategically. Business.com’s 2026 Small Business AI Outlook, a survey of 1,009 US workers at companies under 250 employees conducted with Dialog, found 57 percent of small businesses now investing in AI, up from 36 percent in 2023. Among firms with 1 to 9 employees, that number drops to 24 percent. The average worker using AI saves 5.6 hours a week. The tools are working. The spending is growing. The optimization is not happening.
One honest caveat, because I refuse to sell you a clean story. Cheaper is not universal. Tom’s Hardware pointed out that Kimi K2 launched a year ago at $0.60 per million input tokens, which means uncached K3 input costs five times more than its predecessor. The frontier is not getting cheaper. The tier just below the frontier is, and that tier keeps eating more of the work.
The Solution: Build a Switching Muscle, Not a Model Preference
The framework I now use with entrepreneurs has three parts, and the order matters.
First, separate your work from your model. This is the single highest leverage change most businesses can make, and it is not technical. It is a documentation habit. For every AI-assisted task in your business, write down what goes in, what must come out, and what “correct” looks like. That description is your asset. The model is a rented engine. When your process lives in a document instead of inside one vendor’s chat history, you can test a new engine in an afternoon instead of a quarter.
I learned this the hard way. My Perfect Prompt Framework exists partly because I got tired of prompts that only worked in one place. When you specify role, task, context, format, and constraints explicitly, your prompt becomes portable. It stops being a magic incantation tuned to one model’s personality and starts being an instruction that any competent model can follow. That portability is worth real money now, and it will be worth more next year.
Second, tier your work by how much intelligence it actually requires. Not everything deserves the smartest model. Flo Crivello put it well when he wrote that you do not need God to write your emails. I sort work into three buckets. High-judgment work like strategy, positioning, and anything a client reads with my name on it stays on a top-tier model. Medium work like first drafts, summaries, and outlines runs on a cheaper tier with review. High-volume mechanical work like classification, tagging, extraction, and routine replies belongs on the cheapest capable model available. Most businesses have far more bucket three than they realize, and that is where the money is.
Third, test before you trust, and roll out in stages. Lindy’s process is the template, scaled down. They built offline evaluations from real tasks, ran candidates against them, tested the same model across different providers, tuned prompts, rolled out to internal users first, then watched retention over weeks before going to full traffic. You do not need their engineering team to borrow the logic. Your version is: keep twenty real examples of a task with known good outputs. Run any candidate model against all twenty. Compare side by side. Move one workflow, not ten.
The tools that make this practical are more accessible than a year ago. A routing layer like OpenRouter, which raised $113 million at a $1.3 billion valuation in May, lets you point one integration at many models and change your mind with a configuration change instead of a rebuild. US-hosted inference providers like Fireworks and Atlas Cloud will serve open-weight models on American infrastructure, which addresses the data sovereignty concern behind the Airbnb and Anysphere headlines. And for genuinely sensitive work, open weights mean you can eventually run the model yourself, the only arrangement where nobody can raise your price or change your terms.
Take the muscle, not the vendor recommendation. Businesses that can evaluate and switch in two weeks will ride this cost curve down for years. Businesses that cannot will pay January prices in December.
Practical Steps: What to Do in the Next Thirty Days
Find every dollar you spend on AI. Pull your last three months of card statements and subscription invoices and list every AI tool, seat, credit pack, and AI-enabled upgrade. Include the ones bundled into software you already pay for. You cannot manage a number you have never seen written down.
Inventory your top ten AI-assisted tasks. Write one line per task describing what goes in and what comes out. Rank them by volume, because volume is where price differences compound. A task you run twice a month is not worth optimizing; a task you run two hundred times a month is.
Sort those tasks into three intelligence tiers. High judgment, medium draft work, and high-volume mechanical work. Be honest. Most owners over-assign work to the top tier out of habit, not because the task requires it.
Build a twenty-example test set for your highest-volume task. Save twenty real inputs and the outputs you would consider correct. This takes an afternoon and becomes the most useful asset you own for every future model decision. Without it, you are judging models on vibes.
Run two challenger models against that test set. Pick one cheaper commercial option and one open-weight option served by a provider you trust. Compare outputs side by side against your known-good examples. The serving provider matters, not just the model name.
Move exactly one workflow and watch it for two weeks. Start with something internal where a mistake is embarrassing rather than expensive. Tell your team you are testing so they report problems instead of quietly working around them.
Put a recurring quarterly review on your calendar right now. Ninety minutes, four times a year, to re-price your assumptions. Given the rate of decline Epoch AI documented, a quarterly review is not paranoia. It is the minimum responsible interval.
Frequently Asked Questions
Are open-source AI models actually good enough for real business work?
For most business tasks, yes. Open-weight models now compete at the top of major benchmarks, with Moonshot’s Kimi K3 ranking first in Arena’s Frontend Code evaluation in July 2026. The practical answer depends on your task, not the model. Run cheaper models against twenty real examples of your own work and judge from the results.
How much can a small business realistically save by switching AI models?
Savings vary widely by workload. Lindy, an AI assistant company, reported inference costs falling roughly 90 percent on the traffic it migrated from Claude and Gemini to DeepSeek v4 Flash. Published API list prices show differences of ten times or more between frontier and open-weight options, so meaningful savings are realistic for high-volume work.
Is it safe to use Chinese AI models like DeepSeek or Kimi in my business?
It depends on your data. Hosted APIs from Chinese providers process your data under Chinese law, and US House committees opened an investigation in April 2026 into Airbnb and Anysphere over their use of Chinese models. For regulated or sensitive data, use US-hosted inference providers serving the open weights, or avoid these models entirely.
What does open weights actually mean, and why should I care?
Open weights means the model’s parameters are published for anyone to download and run. Moonshot committed to releasing Kimi K3’s full weights by July 27, 2026. For your business it means multiple companies can host the same model and compete on price, and you are never locked into one vendor’s pricing decisions.
How often should I review my AI spending and tools?
Quarterly at minimum. Epoch AI measured the price of reaching a given AI capability falling at a median of about 50 times per year, accelerating to roughly 200 times per year for post-2024 models. At that pace, pricing assumptions made six months ago are meaningfully out of date, and annual reviews are far too slow.
The Close: Stale Numbers Cost Real Money
I opened this by admitting I felt embarrassed looking at my own AI spend. I want to end there too, because the embarrassment was useful.
It was not about the money. It was about realizing I had let a fast-moving input become a fixed assumption. I had done the thing I warn other entrepreneurs about constantly: I made a good decision once and then stopped deciding.
You are probably doing the same thing right now. Not because you are careless, but because you are busy, and because your AI tools work well enough that nothing is screaming for your attention. That is exactly how this cost hides. It does not show up as a crisis. It shows up as a margin you never got, on work you were already doing, using a model that was replaced three times while you were not looking.
So here it is as plainly as I can say it. The intelligence you are buying today is cheaper than it was in January, and you are almost certainly not buying it at the new price. Nobody is going to send you a refund. Nobody from your AI vendor is going to call and suggest you spend less. The savings are sitting there, in public, on pricing pages you have never read, and the only thing between you and them is thirty days of unglamorous work: list your tools, rank your tasks, build a test set, move one workflow, and put a quarterly review on the calendar.
Do that and you will not just capture this cost drop. You will capture the next one, and the one after that, while your competitors pay January prices and wonder why their margins feel tight.
The price of intelligence keeps falling. Make sure your business is standing underneath it.
About the Author
Jonathan Mast is the founder of White Beard Strategies, where he provides AI coaching and mentorship to entrepreneurs and small business owners who want practical results instead of hype. He serves a community of tens of thousands of entrepreneurs, is the creator of the Perfect Prompt Framework, and speaks regularly on applying AI to real business problems. His focus is simple: help business owners use AI to make more money and get more time back, in plain English, without needing a technical background.
Sources
- LLM inference prices have fallen rapidly but unequally across tasks, Epoch AI
- China’s 2.8-trillion-parameter Kimi K3 beats Claude Fable 5 in Frontend Code Arena benchmark, Tom’s Hardware
- Models & Pricing, DeepSeek API Docs
- Migrating from Claude to DeepSeek, Lindy Engineering Blog
- Fireworks Secures $1.5 Billion in Series D Funding, Fireworks AI
- Nvidia-backed Fireworks hits $17.5 billion valuation as companies pursue cheaper AI models, CNBC
- 2026 Small Business AI Outlook Report, Business.com
- House committees probe Cursor parent, Airbnb over Chinese AI, Semafor
- Chairmen Moolenaar, Garbarino Announce Joint Investigation into Airbnb, Anysphere, House Select Committee on the CCP
- How OpenRouter Model Routing Works, OpenRouter
- OpenRouter Raises $113M for Enterprise AI Model Routing, Enterprise DNA