If a vendor offers you the same AI model for a twelfth of the price in exchange for training on your prompts, this article answers the question underneath that offer: which of your work can safely go into a cheap tier, and which cannot.
The Pricing Page Just Started Telling the Truth
Meta released a coding agent this week with two prices for the same product.
The standard tier costs $1.25 per million input tokens and $4.25 per million output. The contributor tier costs ten cents and twenty cents. Same model. Same capability. Roughly a twelvefold discount, and the entire difference is one permission: you let Meta train on your prompts.
I want to be careful about how I frame that, because the reflexive reaction is alarm and alarm is not useful here. This is not a scandal. It is the first time a major vendor has priced the trade openly enough that you can actually see it.
Here is the direct answer to the question in the headline. Yes, it is safe to use AI tools that train on your data for a large portion of your work, and it is genuinely unsafe for a specific and identifiable minority of it. The line does not run between tools. It runs between categories of your own information. Throwaway experiments, format tests, public-facing drafts, and your own scratch thinking can go into a training-enabled tool without any real exposure. Client identities, unreleased positioning, pricing logic, and anything covered by a confidentiality agreement should not, regardless of how good the discount looks.
The problem is that almost nobody has drawn that line. In a July 2026 survey of 500 employed US adults commissioned by Kolmogorov Law, 38 percent admitted entering at least one type of work information into a personal AI account their employer does not control, and 64.4 percent did not know that doing so can, in some circumstances, break the law.
The thesis of this article is simple. The mistake is not choosing the cheap tier. The mistake is choosing it without knowing you chose it. What follows is how to know.
Key Takeaways
- The cost gap between AI tiers is now large enough that the data permission attached to the cheap tier is a real business decision, not fine print.
- Nearly two in five US workers have put company information into personal AI accounts their employer does not control, and most did not know it carried legal risk.
- The safe line runs between categories of your own information, not between vendors, which means you need an information tiering system before you need a tool policy.
- Redacting names and figures preserves most of the value of an AI task while removing most of the exposure, and it takes under two minutes per document.
- Asking your current vendor for a no-training option costs nothing and works more often than most owners expect.
The Problem: Nobody in Your Business Can Answer the Question
Try this at your next team meeting. Ask out loud which of the AI tools your company uses train on the data you put into them.
I have asked versions of this question in a lot of rooms, and the response is almost always the same. A pause. Then someone offers a guess about the one tool everybody knows about. Then quiet.
That quiet is the actual finding. It is not evidence that your team is careless. It is evidence that AI arrived in your business the way most useful things arrive in a small business: sideways, one person at a time, solving a real problem on a busy afternoon. Nobody convened a procurement committee. Somebody had a deadline and a free account, and it worked, and they told a colleague.
I did this myself, and I want to be honest about it because I think the honesty is more useful than the advice. For a long stretch, I had a working habit of pasting whatever I was struggling with into whatever AI tool I had open. Client emails I was drafting. Positioning I had not published. Pricing I was still deciding. My mental model was that I was talking to a very smart, very forgetful assistant. That model was comfortable and it was out of date by about two years.
The uncomfortable arithmetic is that the industry-wide numbers suggest this is normal rather than unusual. Research circulating in 2026 puts the share of employees who paste content directly into generative AI tools at 77 percent, with more than half of those paste events containing corporate information. The average user performs roughly 6.8 pastes per day, with 3.8 of them including sensitive content.
I am not going to tell you to stop using AI. That advice is useless and nobody follows it. I am also not going to tell you the risk is imaginary, because the survey data says otherwise and because you have signed agreements with clients that say otherwise.
But what if the problem is not the tools at all? What if it is that you have been making a data classification decision hundreds of times a week without ever having classified your data?
What the Research Actually Shows
Four findings shaped how I think about this.
Nearly two in five workers have already done it, and most did not know it mattered. The Kolmogorov Law survey of 500 employed US adults in July 2026 found 38 percent had entered work information into a personal AI account outside employer control. The more telling number is the second one: 64.4 percent did not know this can be illegal under some circumstances. That is not a discipline problem. That is a knowledge gap, and knowledge gaps are fixable with a one-page document, which is a much cheaper fix than most owners assume.
The volume is higher than any policy currently accounts for. At roughly 6.8 pastes per day with 3.8 containing sensitive content, we are not talking about occasional lapses. We are talking about a continuous, high-frequency flow of business information into systems whose terms almost nobody has read. Any policy built on the assumption that this happens rarely is built on sand.
The pricing gap is now large enough to change behavior at scale. This is the part that makes 2026 different from 2024. When the premium and free tiers were separated by a small margin, most businesses defaulted to whatever was convenient. Meta’s contributor tier at roughly one twelfth of standard pricing changes the calculation. Add DeepSeek V4 Flash at fourteen cents per million input tokens, and OpenAI cutting GPT-5.6 Luna pricing by 80 percent, and you have an environment where the cheap option is not just cheaper, it is dramatically cheaper. Discounts that large reliably override policy, especially in a business where one person controls the credit card and the workload.
Confidentiality exposure is asymmetric, which is what makes it manageable. Most of what you put into an AI tool is genuinely uninteresting to anyone. A prompt asking for five subject line variations reveals nothing. A prompt containing your unreleased pricing structure, your client’s internal financials, or the exact positioning language you spent a quarter developing is a different object entirely. This asymmetry is good news, because it means you do not need to protect everything. You need to identify the small percentage that matters and route it differently.
Put those four together and a clear picture emerges. The behavior is already widespread. The volume is high. The economic pressure toward the cheap tier is intensifying. And the actual risk is concentrated in a small, identifiable subset of your content.
That is a solvable problem. It is just not solvable by worrying about it.
What Changed for Me: A Tier, Not a Ban
The shift that made this workable was giving up on a single rule.
For a while I tried to operate on one policy: use the paid, no-training tier for everything. It felt responsible. It lasted about six weeks. It failed for the most boring reason imaginable, which is that I was paying premium rates to have an AI reformat a bullet list and check my grammar. The economics were absurd and eventually I stopped following my own rule, which is worse than never having had one.
What replaced it was three buckets.
Bucket one is throwaway. Format tests. Grammar. Brainstorming that involves no proprietary information. Rewriting a paragraph I have already published. Scratch code for something I will delete. This bucket is large, probably 60 to 70 percent of my actual AI usage by volume, and none of it carries exposure. It goes to the cheapest capable tier available and I do not think about it again. When Meta or anyone else offers me a twelvefold discount for this work, I take it without hesitation, because there is nothing in it worth protecting.
Bucket two is internal. Strategy I am still forming. Draft positioning. Financial modeling with my own numbers. Team documents. This bucket goes to tools where I have checked the terms and know my inputs are not used for training. It is smaller than bucket one but it is where most of the actual thinking happens.
Bucket three is client-facing and confidential. Anything covered by an agreement. Anything containing another business’s identifying information. This bucket gets redacted before it goes anywhere, and the redacted version goes to a no-training tool. That is belt and suspenders, and I am comfortable with the redundancy because a breach here does not cost me a subscription, it costs me a relationship.
The redaction step is the piece I resisted longest and now value most. I assumed replacing names and figures with placeholders would degrade the output. I tested it on real work and the difference was much smaller than my anxiety predicted. An AI helping you restructure an argument does not need to know the client is Acme Manufacturing. It needs to know the client is a mid-sized manufacturer with a specific objection. Substituting the first for the second costs about ninety seconds and removes most of the exposure.
The other thing that surprised me: asking vendors directly works more often than expected. Several tools have quiet no-training options, enterprise terms, or opt-out settings that are not advertised on the pricing page. A short, non-adversarial email asking whether an option exists for an account of your size is free to send. I have had it work more than once.
Practical Steps
1. Build the inventory before you build the policy. Open a document and list every AI tool that touches your business, including free accounts your team signed up for individually. Most owners find between two and three times as many as they expected. You cannot govern what you have not counted, and this step takes about thirty minutes.
2. Find the one sentence in each vendor’s terms about training. It exists, usually under a heading like Data Usage or How We Use Your Information. Quote it into your inventory sheet verbatim rather than paraphrasing, because paraphrase is where misunderstanding enters. If you cannot find it, that absence is itself information worth recording.
3. Classify your own information into three tiers before you touch the tools. Throwaway, internal, client-confidential. Do this based on one question: what happens if this content became public. Owners consistently misjudge this on the first pass, usually by over-protecting generic material and under-protecting the specific language that actually constitutes their advantage.
4. Match tiers to tools and accept that some pairings will fail. Anywhere client-confidential content is currently flowing into a training-enabled tool, you have a decision to make. The options are stay and accept, redact and stay, or migrate. All three are legitimate. Only the fourth option, which is not deciding, is a problem.
5. Build a redaction habit and time it. Create a standing substitution convention for names, figures, and identifying details. Time yourself doing it on a real document. If it takes more than two minutes, simplify the convention until it does not, because a habit that costs five minutes will be abandoned within a month.
6. Ask your existing vendors for a no-training option in writing. Keep it under 120 words, describe your usage volume, and do not sound adversarial. The worst outcome is a no. The realistic outcome, more often than owners expect, is a settings toggle or a tier you did not know existed.
7. Write a one-page rule and date it. Which tier of information may go into which class of tool. Put it where your team will actually encounter it, not in a folder nobody opens. Put a date on it and a review reminder ninety days out, because vendors change their terms and will not notify you.
Frequently Asked Questions
Does using an AI tool that trains on my data mean my information becomes public?
No. Training on your input means your text becomes part of a dataset used to improve the model. It is not published or made searchable. The realistic risk is subtler: your language, patterns, and specifics contribute to a system your competitors also use, and you cannot retrieve what you contributed.
Is it actually illegal to put client information into a personal AI account?
It can be, depending on your contracts and jurisdiction. Many client agreements contain confidentiality clauses that a personal AI account violates, and regulated industries have additional obligations. In the July 2026 Kolmogorov Law survey, 64.4 percent of workers did not know this. Check your client agreements, not your instincts.
Should I just pay for the premium tier for everything to be safe?
Probably not. I tried it and abandoned it, because paying premium rates to reformat bullet lists is economically absurd and abandoned policies are worse than no policy. Route genuinely non-sensitive work to cheap tiers and reserve premium pricing for content where exposure would cost you something real.
How do I know if redacting will ruin the AI’s output?
Test it rather than assuming. Run the same task twice, once with full detail and once with names and figures replaced by descriptive placeholders. Compare the results against your actual purpose. In most cases the difference is far smaller than the anxiety predicts, because the model needs the shape of the situation, not the identities.
What if my team is already pasting sensitive data and I only just found out?
Do not lead with blame. The survey data says this is normal behavior, not negligence, and 64 percent of people did not know it mattered. Build the inventory, write the one-page rule, explain the reasoning once, and make the compliant path the easiest path. Punishment produces hiding, not compliance.
The Ten Cents Is Not the Point
Go back to that pricing page. One dollar twenty-five, or ten cents.
The eleven cents is not really the story. The story is that for two years, the AI cost curve looked like a gift. Prices fell, capability climbed, and most of us quietly stopped reading the terms because the number kept getting smaller and smaller numbers do not feel like they require scrutiny.
That era just ended, and it ended in the most useful way possible: with a vendor stating the trade plainly enough that you cannot miss it. Every AI tool in your business has always had terms. This is the first one that put the price of those terms in a number sitting right next to the alternative.
So here is what I would actually ask you to do, and it is not complicated. Do not build a governance framework. Do not hire a consultant, including me. Just open a document tonight and write down every AI tool your business touches, and next to each one write yes, no, or I do not know.
The I-do-not-knows are your entire assignment. There will be more of them than you expect, and working through them will take an afternoon, not a quarter.
Because the businesses that get hurt over the next two years will not be the ones that used cheap AI. They will be the ones that never found out they were paying for it with something other than money.
Jonathan Mast is the founder and CEO of White Beard Strategies, where he helps entrepreneurs and small business owners put AI to work without hype, guesswork, or wasted spend. He serves a community of more than 100,000 entrepreneurs learning to use AI practically, is the creator of the Perfect Prompt Framework, and speaks regularly on AI implementation for small business. He has read more vendor terms of service than any reasonable person should, and started doing it only after realizing he had been pasting unreleased pricing into a tool he had never checked.