Our team will be out of office on Friday, May 1, 2026. We’ll be back and ready to assist you starting Monday, May 4th.

What Happens to My Business If My AI Tool Goes Down?

Contents

If a single vendor outage would stop something your customers can see, this article answers the question most owners have never actually worked through: which of your processes depend on AI being available, and what you do in the hour it is not.


Seven and a Half Hours

On August 5, Anthropic’s models went offline for roughly seven and a half hours. Four of them at once. Reported as the 164th service disruption this year.

The day before, the same company disclosed a second record chip financing deal, bringing committed compute financing to something like seventy-one billion dollars inside sixty days.

Read those two facts next to each other, because together they say something important. Seventy-one billion dollars of infrastructure spending did not prevent an outage. That is not a criticism of Anthropic. It is the nature of infrastructure. AWS goes down. Stripe goes down. The power company goes down. Spending more money buys you better odds, not certainty.

Here is the direct answer to the headline question. If your AI tool goes down today, what happens depends entirely on one thing you have probably never written down: whether a customer is standing on the other side of that dependency. If the AI sits in your drafting, thinking, or internal work, an outage is an annoying afternoon. If it sits in a live customer-facing path with no human in front of it, an outage is a service failure your customer experiences directly and attributes to you, not to your vendor.

Most small businesses have quietly moved AI into the second category over the last eighteen months without ever making that decision on purpose.

The thesis of this article is that the binding constraint on AI in a small business has changed. For three years it was capability, and the correct response was patience. Now it is availability, and the correct response is design.


Key Takeaways

  • The failure mode of AI in small business has shifted from insufficient capability to unplanned unavailability, and most continuity planning has not caught up.
  • Anthropic took a seven and a half hour multi-model outage in the same week it disclosed roughly seventy-one billion dollars in committed compute financing, which demonstrates that spending buys better odds, not certainty.
  • Downtime research puts typical small and mid-sized business costs between roughly $5,000 and $50,000 per hour, with Datto’s research placing the SMB average near $8,000 per hour.
  • Your customers attribute an outage to you, not to your vendor, which means vendor uptime is your operational problem regardless of whose fault it is.
  • The highest-value fix is usually not a second vendor. It is moving AI out of the live customer path and into a draft step a human reviews.

You Built a Dependency You Never Approved

Nobody sits down and decides to make their business dependent on a single AI vendor. It happens the way most dependencies happen, which is one useful decision at a time.

You start with AI helping you draft. That works, so you use it to handle first-pass responses. That works, so you connect it to your inbox. That works, so somebody wires it into the intake form. Each step was sensible. Each step was small. And somewhere in that sequence, without a meeting or a decision or a note in a file, the AI moved from being something that helped you work to being something your customers touch.

Then you find out on a random Wednesday, when it is not there.

I have been on the wrong end of this. I had a process where AI handled the first response to a specific category of inbound message. It was good. It was faster than me. And when the tool was unavailable one morning, the honest answer to “what happens now” was that I did not know, because I had never asked. Messages sat. Nobody noticed for a while. The failure was silent, which is the worst kind, because a loud failure at least tells you to do something.

That is the part I want to name clearly. The problem is rarely the outage itself. Outages are usually measured in hours and most businesses can absorb a bad few hours. The problem is that nobody knows what to do during them, so the default behavior is waiting, and waiting is only the right answer if you decided in advance that it was.

Let me be honest about the difficulty here too, because I do not want to make this sound easy. You cannot buy uptime at small-business scale. You are not going to negotiate an SLA with a frontier lab. You have no leverage over their infrastructure, no visibility into their roadmap, and no seat at the table when they decide to sunset a model. That is genuinely uncomfortable and I am not going to pretend a framework makes it go away.

But what if the goal was never to prevent the outage? What if the entire job is making sure that when it happens, it is boring?


The Numbers Are Worse and Better Than You Think

Four things worth knowing.

The downtime cost research is startling, and it applies more to you than you think. Estimates for small and mid-sized businesses in 2026 range from roughly $5,000 to $50,000 per hour depending on size, industry, and technology dependence, with Datto’s research placing the SMB average near $8,000 per hour. Those figures were built for IT downtime broadly rather than AI specifically, and I would treat the top of that range with skepticism for a very small operation. But the shape of the finding holds: the cost of a system being unavailable is consistently higher than owners estimate before they experience it, because the visible cost is only the work that stopped, and the invisible cost is the trust that leaked.

Frequency matters more than severity for AI specifically. A seven and a half hour outage reported as the 164th disruption of the year tells you something a single incident does not. This is not a rare catastrophic event you insure against. It is an ongoing operational reality you plan around, more like weather than like an earthquake.

Roadmap risk is real and currently visible. Google’s flagship Gemini release, planned for June, still had not shipped as of August 5, and the same week brought a leadership overhaul at DeepMind with senior departures including Jeff Dean. Alphabet shares fell about 4 percent. If you have promised anything to a customer or your team that depends on a capability that has not actually shipped, that promise is a guess wearing a plan’s clothing.

Concentration risk runs deeper than your own stack. A filing reported August 5 shows Microsoft recorded $24.1 billion in revenue from OpenAI in the year ended June, implying OpenAI accounts for more than half of Microsoft’s AI sales. You may believe you have diversified by using two vendors. Trace the dependencies far enough down and the diversification often turns out to be thinner than the logos suggest.

Now the better news, and it is genuinely better. Practitioner communities are reporting something that cuts against the entire industry narrative. Some of the highest-performing automations in production right now have no model in the critical path at all. A missed-call text-back system for a truck repair shop. A logistics dispatch workflow explicitly labeled as having no LLM node, running in production. These are not sophisticated. They are also not fragile, and they do not care whether a frontier lab is having a bad day.

That is the door out of this problem. Not redundancy. Reduction.


What Changed for Me: Move the Model, Do Not Duplicate It

My first instinct after getting burned was redundancy. Wire in a second vendor. Have somewhere to go.

That is not wrong and I did do it for my highest-value workflow. But it was not the change that mattered, and I want to be precise about why. Redundancy adds a moving part. It means two sets of prompts, two sets of quirks, two places where output format can drift, and a switching decision made under stress by whoever happens to be around. It buys real protection and it costs real complexity, so it is worth doing exactly once, on the thing that matters most.

The change that actually mattered was cheaper and more boring. I moved the AI out of the live customer path and into a draft position.

Concretely: instead of AI generating a first response that went to a customer, AI generates a draft that a person approves. The speed cost is small, usually minutes. The reliability gain is enormous, because now an outage does not produce a service failure. It produces a person doing the task manually, slightly slower, exactly as they did eighteen months ago. Nothing breaks. Nothing goes silent. A customer waits twenty minutes instead of four.

That single structural move converted my worst failure mode from invisible to trivial.

The second change was adding success signals rather than error alerts. This sounds like a small distinction and it is the whole thing. An error alert fires when something breaks loudly. A success signal confirms the thing actually happened. Silent failures never trigger error alerts, which is precisely what makes them silent. Now I get a daily count in a message I already read. If the number is zero and it should not be, I know within a day instead of within a week.

The third change was deciding my switch trigger in advance, in writing, before I needed it. Thirty minutes of confirmed outage on a customer-facing process, and we move to manual. Who decides, what they do first, how we switch back. One page. The point of writing it down is not that the plan is brilliant. The point is that nobody makes a good decision during an outage, so the decision has to be already made.

And I moved my prompts and context out of any single vendor’s interface into plain text files. A prompt living inside one tool is a prompt you will rewrite from memory under pressure. A prompt in a text file is a prompt you paste somewhere else in fifteen seconds.


Practical Steps

1. Walk your customer journey literally and mark every AI touchpoint. Not from memory. Actually move through the path a customer takes, from first contact to delivery, and note every place a model sits in the sequence. Owners routinely find one or two they had forgotten about entirely, usually inside an automation somebody set up months ago.

2. Sort those touchpoints into three tiers. Must never stop, can degrade gracefully, can wait. Be ruthless, because you cannot protect everything and pretending otherwise means protecting nothing well. Most businesses have fewer must-never-stop processes than they initially claim.

3. For each must-never-stop process, write the manual fallback. Not the elegant fallback. The one a person on your team could execute within ten minutes with no tools. Manual is not a failure state, it is how your business ran before, and it is what keeps the lights on during someone else’s incident.

4. Add a success signal to every automation that could fail silently. A daily count in a channel you already read beats a dashboard you will never open. The test is simple: if this stopped working entirely, how long before somebody noticed? If the answer is more than a day, you need a signal.

5. Move the model out of the live path wherever a review step is tolerable. This is the highest-leverage change available and it costs you minutes of speed. Convert customer-facing generation into draft-and-approve. Your worst-case outcome shifts from a service failure to a slower human.

6. Add one second vendor, to one workflow, and stop there. Pick your highest-value AI dependency, wire in an alternative, and make switching a matter of minutes. Resist doing this everywhere, because redundancy multiplies maintenance and unmaintained redundancy is worse than none.

7. Write your switch trigger on one page and run the drill once. Time threshold, decision owner, first action, how you switch back. Then test it for thirty minutes without touching real customers. An untested fallback is not a plan, it is a wish with formatting.


Frequently Asked Questions

How likely is an AI outage that actually affects my business?

More likely than most owners assume. A single vendor reported its 164th service disruption of the year in early August 2026. Treat AI availability the way you treat internet availability: not a rare catastrophe, but a recurring condition you build around. The relevant question is not whether, it is how boring you have made it.

Should I use two AI vendors for everything to be safe?

No. Redundancy adds maintenance burden, duplicate prompts, and a switching decision made under stress. Add a second vendor to your single highest-value workflow so you have somewhere to go, and address everything else by moving AI out of the live customer path instead.

What is the fastest way to reduce my AI outage risk this week?

Convert one customer-facing AI step into a draft-and-review step. It takes an afternoon, costs you minutes of speed per task, and changes your worst-case outcome from a visible service failure to a person working slightly slower. That single structural change outperforms most redundancy work.

How do I know if one of my automations is failing silently right now?

Ask how you would find out. If the answer involves a customer complaining or you noticing something felt quiet, it is a silent-failure candidate. Add a daily confirmation signal showing the thing happened, then deliberately break it in a controlled test to verify the signal fires.

My AI vendor promised a new model that has not shipped. What should I do?

Rebuild any commitment that depends on it around capability that exists today. Google’s flagship Gemini release promised for June had still not shipped by August 5. Treat vendor roadmaps as marketing, plan with what is stable this morning, and let anything better that arrives be upside rather than a requirement.


Make the Bad Tuesday Boring

Seventy-one billion dollars, and it still went dark for seven and a half hours.

I keep coming back to that pairing because of what it settles. You are never going to out-engineer this. You will not negotiate an SLA that matters, you will not get advance notice, and you will not be told when the model you depend on is being retired until the notice arrives with everyone else’s.

So stop trying to prevent the outage. That was never the assignment.

The assignment is to make sure that when it happens, and it will happen, the answer to “what do we do now” is already written down and mildly boring. Somebody picks up the manual version. A customer waits twenty minutes instead of four. Nothing goes silent. Nobody discovers on Friday that Wednesday’s messages never went out.

That is not a sophisticated ambition. It is the entire game. And the businesses that get hurt in the next year will not be the ones whose vendor had a bad day, because every vendor will have a bad day.

They will be the ones who never asked what happens next.


Jonathan Mast is the founder and CEO of White Beard Strategies, where he helps entrepreneurs and small business owners put AI to work without hype, guesswork, or wasted spend. He serves a community of more than 100,000 entrepreneurs learning to use AI practically, is the creator of the Perfect Prompt Framework, and speaks regularly on AI implementation for small business. He learned the value of success signals over error alerts the way most people do, which is by finding out several days late that nothing had been sent.

About the Author