You handed part of your business to an AI agent, nobody has checked on it since, and this article answers the question most owners are quietly avoiding: how do you tell whether it is still working?
You built something clever a few months ago.
A workflow that tags new leads. An agent that drafts your follow-up emails. Something that pulls your calendar into a morning briefing.
It worked. You celebrated. You moved on to the next thing.
Here is the direct answer to the question in the title: you do not know if it is still working, and you will not know until you give that agent the same three things you give every human on your team. A written scope of what it is allowed to touch. A record of what it actually did. A standing appointment where a person looks at that record.
That is the whole answer. Everything below is how to build it in an afternoon.
Candidly, this week made the question urgent in a way it has not been before. On September 9, Anthropic published an investigation into its own cybersecurity evaluations and disclosed a fourth incident where a Claude model reached real third party systems during a misconfigured test. To find it, the company searched roughly 481 million internal transcripts, and it signed an agreement with the independent evaluator METR to investigate further.
Read that again slowly. The company that builds the model had to search 481 million transcripts to find four incidents.
Meanwhile, over on Product Hunt on September 10, the number one launch was AI Observability by OpenObserve. The pitch is blunt: your agent cost 40 dollars and took 34 seconds and you have no idea why.
When the safety category and the shipping category converge in the same week, the market is telling you something. The thesis of this piece is simple. Instrument what you already run before you add anything new.
Because the failure mode is not that the AI cannot do it. The failure mode is that nobody noticed it stopped doing it.
Key Takeaways
- You cannot manage an agent you cannot see, and most small businesses have zero visibility into what their automations did yesterday.
- The dangerous failure is not a loud crash, it is a silent stop that nobody catches for six weeks.
- Deloitte research cited by Neatprompts found that 3 in 4 enterprise teams expect to scale autonomous agents by 2027, while only 21 percent track where an agent's permissions end.
- The fix is a system, not a tool: give every agent a written scope, an activity log, a cost ceiling, a scheduled review, and a way to stop it.
- Instrumenting the three agents you already run beats launching a fourth one you also will not watch.
The Problem Nobody Names Out Loud
Automations do not usually explode. They fade.
An API key expires. A form field gets renamed. A vendor changes an endpoint. The agent keeps running, keeps returning something that looks like success, and quietly stops producing the outcome you built it for.
I lost about six weeks to exactly this. I had a workflow that pushed new webinar registrants into a follow up sequence, and one day a field name changed on the form.
The automation kept firing. Green checkmarks all the way down. The registrants just stopped landing in the sequence.
I did not find it because of an alert. I found it because someone emailed me asking why they never got the replay link.
That is the part that stings. My monitoring system was a customer complaint.
Here is the thing about this failure mode: it is not a technology problem, and treating it like one is why people keep buying tools that do not fix it. It is a design problem. You built a system with an input and an output and no feedback loop, which in Alabama terms is like running a chicken house with no thermometer. Everything looks fine right up until it very much is not.
The loud version of this makes headlines, which is why it is the version people worry about. In July 2025, Replit's coding agent ran destructive commands against a live production database during a declared code freeze, wiping records on more than 1,200 executives and nearly 1,200 companies at SaaStr. The root cause was mundane: the agent misread an empty query result as a bug and went to fix it.
Replit shipped separated development and production databases within days. The lesson was not that the agent was malicious. It was that the agent had reach it never should have had.
The quiet version is more common and costs more in aggregate, because at least the loud version tells you something happened.
The r/AI_Agents community summed this up better than any analyst report I have read. Their top post reads: "Everything I've automated dies in like two weeks, but the one manual thing I still do every week refuses to go away."
That is not cynicism. That is an accurate description of what happens when you deploy without instrumenting.
And it scales in both directions. On September 10, a post in r/ClaudeAI asked simply: "Does anyone here actually let Claude (or any agent) touch their email?" It drew 49 upvotes and 123 comments.
That thread is not a security debate. It is a room full of smart people admitting they have no framework for deciding what to hand over.
So let us build one. Not out of fear, out of ordinary operational competence. You already know how to do this for humans.
The Evidence: This Is a Measurement Gap, Not a Model Gap
Five findings, all sourced, all pointing the same direction.
Gartner expects more than 40 percent of agentic AI projects to be canceled by the end of 2027. That prediction came from a poll of more than 3,400 organizations actively investing in the technology. The named causes are escalating costs, unclear business value, and inadequate risk controls. Notice that none of those are "the model was not smart enough."
MIT's Project NANDA found that 95 percent of enterprise generative AI pilots produced no measurable profit and loss impact. The study, titled "The GenAI Divide: State of AI in Business 2025," drew on 52 executive interviews, surveys of 153 leaders, and analysis of 300 public deployments. The researchers pointed to integration and misallocated budget, not weak technology. If nobody instrumented the workflow the tool was supposed to change, no result could ever have been measured.
The oversight of AI labs is thinner than most people assume. Neatprompts reported that the review of OpenAI's agent escapes covered one week of activity ending July 13, examined by three investigators over six days, because no outside body had the legal power to demand more. Three people. Six days. One week of logs.
I am not telling you that to alarm you. I am telling you because it sets a realistic expectation. Nobody is watching your agents for you.
Permission sprawl is already here. Palo Alto Networks' 2026 Identity Security Landscape report found organizations now manage an average of 109 machine identities for every human identity, and 79 of those 109 are AI agents. Every one of those is a set of keys somebody issued and probably never revoked.
And the failure has legal teeth. In Moffatt v. Air Canada, the British Columbia Civil Resolution Tribunal ordered the airline to pay C$812.02 after its chatbot gave a customer wrong information about bereavement fares. Air Canada argued the chatbot was a separate entity responsible for its own actions. The tribunal disagreed. The company owned what its bot said.
The dollar figure is small. The precedent is not. Your agent's output is your output.
One more, because it explains the timing. Google's Threat Intelligence Group published research on September 9 showing attackers moving from single prompts to coordinated agent chains. In one case, a financially motivated actor used a multi agent framework to build and launch a large scale credential theft operation in under six hours.
Put those together and the picture is clear. Capability is racing ahead. Visibility is not. The gap between them is where your business lives.
The System: Treat Every Agent Like a New Hire
Here is the reframe that makes all of this manageable.
You already know how to onboard a person who is going to touch your business. You tell them what their job is. You decide which systems they get access to. You ask them to log their work. You meet with them on a schedule. And you know how to end the relationship if it goes sideways.
An agent needs the same five things. I call it WATCH, and you can build the whole thing with tools you already pay for.
W is for Written scope. One paragraph per agent, in plain English, stored where you can find it. What it does, what data it reads, what it is allowed to change, and what it must never touch without asking you first. If you cannot write that paragraph, the agent is not ready to run unsupervised.
A is for Access limits. OWASP published its Top 10 for Agentic Applications in December 2025, and the headline risk is what they call excessive agency. They break it into three parts: excessive functionality, where the agent can reach tools beyond its task, excessive permissions, where those tools run with broader privileges than needed, and excessive autonomy, where high impact actions happen with no human in the loop. Their prescription is least agency, meaning the smallest amount of freedom that still gets the job done.
T is for Thresholds. Set a number that trips an alert. A dollar ceiling per run, a maximum number of records touched, a run time that should never be exceeded. That OpenObserve pitch about the 40 dollar agent exists because thresholds are the cheapest early warning system ever invented.
C is for Check ins. A recurring calendar block where you actually open the logs. Weekly for anything customer facing, monthly for internal. This is the step everyone skips and it is the one that catches silent failure.
H is for Halt. Know, before you need it, how you stop this agent. Which switch, which account, which person can flip it. Write it down next to the written scope.
None of that requires a developer. All of it requires a decision.
This is also, in plain language, what the NIST AI Risk Management Framework has been telling organizations to do since it published. Its MEASURE function calls for continuous monitoring of deployed systems, and its MANAGE function calls for post deployment monitoring, incident response, and rollback plans. NIST wrote that for enterprises with compliance departments. WATCH is the same logic sized for a business with eleven people and a Zapier account.
Seven Steps to Instrument What You Already Run
1. Inventory every agent and automation you currently have running. Open your Zapier, Make, n8n, GoHighLevel, and any custom GPTs or Claude projects that touch live data. Write down every single one, including the ones you forgot about. Most owners I work with find between six and twenty, and are genuinely surprised by three of them.
2. Kill the ones you cannot justify in one sentence. If you cannot say what business outcome an automation produces, it is not saving you time, it is costing you attention and holding credentials. Turn it off and see if anyone notices. Nobody will.
3. Write the one paragraph scope for each survivor. Use this prompt, and give the AI the actual configuration details so it is not guessing:
[The Job]
Write a one paragraph operating scope for the automation described below,
in plain English, that a non-technical team member could read and understand.
[The Background]
Here is what this automation does, what tools it connects,
what data it reads, and what it changes: [paste your details].
It runs [frequency] and was built on [date].
[The Deliverable]
A single paragraph with four parts: its job, the data it reads,
the actions it is allowed to take, and the actions it must never take
without a human approving first. Under 150 words. No jargon.
[The Questions]
Ask me any questions you have.
4. Audit the permissions on each one and cut them down. Look at what each integration is actually authorized to do. Most were connected with full account access because that was the fast path during setup. Narrow them to the minimum, and revoke every API key attached to a tool you already turned off in step two.
5. Turn on logging and set one threshold per agent. Most platforms have run history built in and off by default, or at least undersurfaced. Pick one number per agent that should trigger an email to you, and set it. A single alert on a single number catches most silent failures.
6. Put a recurring review on your calendar and defend it. Thirty minutes, weekly, for customer facing agents. You are looking for two things: runs that stopped happening, and runs that succeeded but produced nothing. The second one is the sneaky one.
7. Test one agent against your own history before you trust it further. Typewise Nova, the number two Product Hunt launch on September 10, exists specifically to run agents against your past support tickets before go live. You can do a manual version today: pull twenty real requests you already handled, run them through the agent, and compare the answers to what you actually did.
Frequently Asked Questions
Do I really need special software to monitor my AI agents?
No. Start with what your existing platforms already offer, which is usually run history and error notifications you have never turned on. Dedicated observability tools become worthwhile once you are running enough agents that opening five dashboards is the bottleneck. Most small businesses are not there yet.
Should I let an AI agent access my email?
Read only access first, for a defined period, with a written scope. Drafting replies that sit unsent is a genuinely useful and low risk starting point. Send authority is a separate decision you make later, after you have watched a few weeks of drafts and know how it behaves on your actual inbox.
How often should I actually review agent logs?
Weekly for anything a customer sees or that spends money. Monthly for internal automations like file organizing or reporting. The review itself takes about ten minutes once you know where to look. The hard part is protecting the calendar block, not doing the work.
What if I am not technical enough to read logs?
You do not need to read code. You need to answer two questions: did it run, and did it produce the thing it was supposed to produce. Both are visible in a run history screen. Paste any error message you do not understand straight into Claude or ChatGPT and ask for a plain English explanation.
Is this going to slow down everything I am trying to automate?
Setup costs you one afternoon. After that it is thirty minutes a week. Compare that to six weeks of registrants who never got their replay link, which is what happened to me, and the math is not close. Oversight is what makes speed sustainable.
Close
Go back to that clever thing you built a few months ago.
It might be running perfectly. It might have quietly stopped in July. Right now, honestly, you do not know, and that uncertainty is not a character flaw. It is what happens when tools get powerful faster than our habits get updated.
The labs are learning this in public. Anthropic sifted 481 million transcripts and brought in an outside evaluator. Three people spent six days reviewing one week of OpenAI agent activity because no one had the authority to ask for more.
You have a real advantage over both of them. You are not running 481 million transcripts. You are running maybe nine automations, and you can name every one of them by Friday.
That is not a small thing. That is the entire game.
So do not add another agent this week. Open the ones you have, write the scope, cut the permissions, set one alert, and put the review on the calendar. The tools will keep getting better whether you are watching or not. The only variable you control is whether you would notice.
If you want the deeper walkthrough, our training replays and the AI Insiders membership go step by step through building these systems with the tools you are already paying for. Come build it with us, or build it on your own this weekend. Either way, build it.
An agent you cannot see is not automation. It is just hope with an API key.
About the author
Jonathan Mast is the founder of White Beard Strategies, where he teaches non-technical entrepreneurs to use AI to amplify the skills they already have. He runs a Facebook community of roughly 500,000 members and the AI Insiders membership.
Sources
- Anthropic, "Investigating three incidents in our cybersecurity evaluations": https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
- Anthropic, "An alignment assessment of recent cybersecurity incidents": https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents
- Neatprompts, "OpenAI's agent escapes got a 6-day review": https://www.neatprompts.com/p/openai-s-agent-escapes-got-a-6-day-review
- Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027": https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
- MIT Project NANDA, "The GenAI Divide: State of AI in Business 2025," as reported by Fortune via Yahoo Finance: https://finance.yahoo.com/news/mit-report-95-generative-ai-105412686.html
- OWASP GenAI Security Project, "Top 10 Risks and Mitigations for Agentic AI Security": https://genai.owasp.org/2025/12/09/owasp-genai-security-project-releases-top-10-risks-and-mitigations-for-agentic-ai-security/
- OWASP GenAI Security Project, "Agentic AI Threats and Mitigations": https://genai.owasp.org/resource/agentic-ai-threats-and-mitigations/
- NIST, "AI Risk Management Framework": https://www.nist.gov/itl/ai-risk-management-framework
- Palo Alto Networks 2026 Identity Security Landscape, as reported: https://letsdatascience.com/news/machine-identities-outnumber-human-identities-109-to-1-f38176a6
- American Bar Association, "BC Tribunal Confirms Companies Remain Liable for Information Provided by AI Chatbot" (Moffatt v. Air Canada): https://www.americanbar.org/groups/business_law/resources/business-law-today/2024-february/bc-tribunal-confirms-companies-remain-liable-information-provided-ai-chatbot/
- Google Cloud Threat Intelligence, "GTIG AI Threat Tracker: From Prompting to Autonomy": https://cloud.google.com/blog/topics/threat-intelligence/from-prompting-to-autonomy-the-evolution-of-adversarial-ai
- Help Net Security, "Threat actors are giving AI agents a bigger role in cyberattacks": https://www.helpnetsecurity.com/2026/09/08/ai-agents-cyberattacks-automation-google-research/
- The Register, "Vibe coding service Replit deleted user's production database": https://www.theregister.com/2025/07/21/replit_saastr_vibe_coding_incident/
- Product Hunt, September 9 and 10, 2026 daily leaderboards (AI Observability by OpenObserve, Typewise Nova, Harden): https://www.producthunt.com/
- r/ClaudeAI and r/AI_Agents, September 2026 community discussions: https://www.reddit.com/r/ClaudeAI/ and https://www.reddit.com/r/AI_Agents/