Our team will be out of office on Friday, May 1, 2026. We’ll be back and ready to assist you starting Monday, May 4th.

Why does the new AI model everyone is calling AGI feel worse at my actual work?

Contents

Five days after GPT-6 Astra shipped, one corner of the internet is calling it artificial general intelligence and another corner is asking for the old model back. Here is why both groups are right, and what to do about it.


The subreddit that broke my brain this week

I had two browser tabs open on Wednesday morning.

In the first tab, a post on r/singularity titled "AGI achieved" sitting at roughly 1,500 upvotes. In the second, a post on r/ChatGPT from a paying customer titled "Astra feels like a downgrade for writers, editors, and documentation work."

Same model. Same week. Opposite verdicts.

Here is the direct answer, and I want you to have it before you read another paragraph. The new model is genuinely better at hard, bounded, benchmark-shaped problems, and it may genuinely be worse at your Tuesday afternoon. That is not a contradiction and it is not a conspiracy. The score that made headlines came mostly from the software wrapped around the model, not from the model itself. If you did not get the wrapper, you did not get the result.

That wrapper has a name. Engineers call it a harness. It is the scaffolding around a model that decides what tools the model can reach, what it remembers between turns, and how its context gets managed.

And candidly, the harness is now the whole ballgame.

The gap between what a model does in a demo and what it does at your desk has never been wider. Closing that gap is the actual job now. Not picking the best model. Not waiting for the next release. Building the structure that makes a good model reliable inside your specific business.

Here is the thesis of this entire article: over the next twelve months, the people who win will not be the people with access to the best model. They will be the people who built the harness around it.


Key Takeaways

  • The headline AGI score for GPT-6 Astra came from a harness, not from raw model capability, and ARC Prize published both numbers on launch day.
  • Frontier capability and daily reliability are now moving in different directions, so a newer model can be smarter on benchmarks and worse on your workflow at the same time.
  • Research consistently shows that benchmark performance does not predict real world results, including a randomized trial where experienced developers got 19 percent slower with AI while believing they were 20 percent faster.
  • Enterprise AI failures are mostly failures of approach, not failures of model quality, according to MIT's Project NANDA.
  • You can build your own harness with context files, saved instructions, a small tool set, and a real evaluation habit, and none of it requires you to write code.

The problem: you are being sold a model when what you need is a system

Every launch cycle runs the same play. A lab publishes a number. The number is astonishing. Everyone upgrades. Then quietly, over the next ten days, half of them discover the thing they used to do easily now takes three tries.

I have lived this personally more than once. Last year I moved my entire content pipeline to a newly released model within about six hours of launch because the coding scores looked incredible. My content pipeline does not write code. It writes in my voice, follows my formatting rules, and refuses to use em dashes. The new model was measurably smarter and it broke every one of those constraints, because the constraints lived in scaffolding that I had tuned for the old model and never rebuilt.

That was not the model's fault. That was mine.

Here is the honest part. The reason this keeps happening is that "which model is best" is an easy question and "what system do I need around it" is a hard one. Easy questions get asked. Hard questions get postponed.

And the marketing does not help. The number in the press release is almost never the number you get. It is the number the lab got under conditions the lab controlled, with tooling the lab built, on tasks the lab selected.

That is not necessarily dishonest. But it is not your desk.

I want to reframe this, because I do not think this is bad news. I think it is the best news a small business owner has gotten in two years.

When raw model capability was the differentiator, you could not compete. OpenAI has more compute than you will ever have. But when the harness is the differentiator, the advantage moves to whoever understands the actual work best. That is you. You know your clients, your voice, your process, your edge cases. Nobody at a frontier lab knows those things. They cannot build your harness. You can.

The moat moved from the model to the implementation. And implementation is a skill you can learn.


The evidence: five findings that explain the split screen

1. The 99.9 percent came from the harness, and the benchmark's own authors said so.

ARC Prize ran GPT-6 Astra two different ways on ARC-AGI-3 and published both results. Under its own standard harness at maximum reasoning, Astra scored 62.7 percent at a cost of $26,098. Under OpenAI's Provider Adapter at high reasoning, it scored 99.9 percent for $18,817. Same weights. Different scaffolding. A 37 point swing, and the better score was the cheaper run. (The Next Web, September 6, 2026, reporting on ARC Prize's published results)

One row in that table is the whole article in a single line. Set the reasoning effort to none inside OpenAI's adapter and Astra still scored 96.7 percent. The scaffolding beat the reasoning dial outright.

ARC Prize itself declined the AGI conclusion. Co-founder Mike Knoop wrote that "we lack evidence to call this AGI yet."

2. This is not an OpenAI quirk. Harnesses are moving everyone's numbers.

The same reporting notes that Nvidia built a harness that took Claude Opus 5 from 30.2 percent on ARC-AGI-3 to clearing every level. Anthropic, Google, and Microsoft all sell harnesses as products with their own pricing. Plain English translation: the industry has already figured out that the wrapper is the product. Most buyers have not.

3. Benchmark brilliance and everyday competence are genuinely different things.

Stanford HAI's 2026 AI Index describes capability gains as jagged. Gemini Deep Think won gold at the 2025 International Mathematical Olympiad. The best model on ClockBench read analog clocks correctly only 50.1 percent of the time, against 90.1 percent for humans. AI agents jumped from roughly 12 percent to 66.3 percent accuracy on OSWorld computer-use tasks, which still means failing about one attempt in three. (Stanford HAI 2026 AI Index, summarized by UNU C3)

Olympiad gold and a coin flip on wall clocks, in the same field, in the same year.

You can see the same jaggedness in Astra's own robotics numbers this week. It scored 19 out of 20 on putting a block in a bowl and 2 out of 20 on inserting a puzzle piece. (reported in TLDR AI, September 7, 2026, from OpenAI's robotics evaluation page)

4. People are bad at judging whether AI is helping them.

METR ran a randomized controlled trial with 16 experienced open source developers across 246 real issues in their own repositories. Developers forecast a 24 percent speedup. Afterward they estimated they had been sped up 20 percent. The measured result was 19 percent slower. (METR, July 2025; METR published updated data in February 2026 and notes the original figures reflect early-2025 tools)

That study is about coding, and METR is careful to say it does not generalize to every setting. I am not citing it to claim AI slows you down. I am citing it for the perception gap, which is the part that generalizes hard. Feeling faster and being faster are separate measurements.

5. When AI projects fail, the approach is usually the culprit, not the model.

MIT's Project NANDA studied enterprise generative AI deployment and found that 95 percent of pilots delivered no measurable impact on profit and loss. The authors were direct about the cause: the divide "does not seem to be driven by model quality or regulation, but seems to be determined by approach." (MIT Project NANDA, The GenAI Divide: State of AI in Business 2025)

One note on the Reddit numbers I opened with. Those upvote counts are community signal, not research. They tell you what people are feeling this week, not what is true. I treat them as a thermometer, never as a scale.


The solution: stop shopping for models, start building a harness

Here is how I think about this now, and it is the same architecture I use inside my own business.

A harness has four parts. Context, instructions, tools, and evaluation. Miss any one of them and the whole thing wobbles.

Context is everything the model needs to know that is not in its training data. Your voice. Your offers. Your pricing. Your clients. Your rules. In my business this lives in plain markdown files that I attach to every project. One file describes who I am and how I write. One file describes the company, the team, and exact pricing. When I switch models, those files come with me, and about eighty percent of my quality holds steady across the switch. That is the harness doing its job.

Instructions are the standing orders. Not the prompt you type today, but the persistent rules that apply to every output. Mine include no em dashes, no income claims, never fabricate a statistic, and every number needs a named source or it comes out. Those live in project instructions, not in my head and not retyped every session.

Tools are what the model can reach. Web search so it can verify a claim. File access so it can read your actual documents. Nothing exotic. The Astra result showed that memory management and tool access moved the score more than the reasoning setting did, and that scales down to your desk perfectly well.

Evaluation is the part almost everyone skips. It is also the part that separates the 5 percent from the 95 percent in the MIT data. You need a small, fixed set of real tasks from your own business that you re-run whenever something changes. Mine is seven tasks. A client email, a sales page section, a blog opening, a spreadsheet cleanup, a meeting summary, a social post, and a research brief. When a new model ships, I run all seven before I move anything.

That is the whole system. Context, instructions, tools, evaluation.

The reason it works is that it makes the model interchangeable. When Astra launched, I did not have to guess whether it was better. I ran my seven. Two got better, four stayed flat, one got worse. So I moved five workflows and kept one on the old model. Total time invested: about ninety minutes.

Ninety minutes is what stands between the person posting "AGI achieved" and the person posting "this feels like a downgrade." Neither of them ran the test. Both of them are generalizing from a vibe.

You do not need to be technical to build this. Every part of it is a document, a setting, or a habit.


Practical steps: build your harness this week

1. Write your context files first. Create two plain documents. One is about you: how you write, what you believe, what you refuse to say. One is about the business: offers, exact prices, team names, common client situations. Attach both to every AI project you run. This is the single highest-leverage hour you will spend.

2. Move your rules out of your prompts and into project instructions. Anything you find yourself typing more than twice belongs in the standing instructions of a saved project or custom assistant. In ChatGPT that is a Project. In Claude that is a Project. In Gemini it is a Gem. Same idea everywhere.

3. Build a seven-task evaluation set from real work. Pull seven actual jobs you have already done well, by hand, in the last month. Save the inputs and your finished versions. This is your ruler. Without a ruler you are guessing, and the METR study shows how badly people guess.

4. Score before you switch. When a new model ships, run all seven tasks and grade each one better, same, or worse against your saved version. Move only the workflows that improved. Keeping an old model for one specific job is a completely legitimate answer, not a failure of nerve.

5. Give the model one or two real tools, then stop. Turn on web search and file access. Resist the urge to bolt on twelve integrations in week one. More surface area means more places for the system to fail quietly, and quiet failure is the expensive kind.

6. Use a consistent prompt structure so results are comparable. Here is the framework I use for anything that matters. Fill in the brackets.

[The Job]
Help me build an evaluation set I can use to test any AI model against my real work.

[The Background]
I run [YOUR BUSINESS TYPE] and my most common AI-assisted tasks are [LIST 3 TO 5 TASKS]. My non-negotiable standards are [LIST YOUR RULES, FOR EXAMPLE TONE, FORMATTING, THINGS YOU NEVER SAY]. I currently use [MODEL OR TOOL] and I am considering [NEW MODEL OR TOOL].

[The Deliverable]
A seven-task evaluation checklist, written from the perspective of a careful operations manager who cares about consistency more than novelty. For each task, give me the input I should paste, the specific qualities I should grade, and a simple better, same, or worse scoring line.

[The Questions]
Ask me any questions you have.

7. Re-run the whole thing quarterly, not on launch day. Launch weeks are the worst time to evaluate anything, because the signal is buried under excitement and pricing changes and half-finished tooling. Put a recurring block on the calendar and let the noise settle first.


Frequently Asked Questions

What exactly is an AI harness, in plain English?

A harness is the software wrapped around an AI model that controls what tools it can use, what it remembers between messages, and how it handles long conversations. The model is the engine. The harness is the rest of the car. Same engine in two different cars produces very different driving experiences, which is exactly what ARC Prize documented with GPT-6 Astra.

Should I upgrade to the newest model as soon as it launches?

No. Test first, then move the workflows that actually improved. ARC Prize's own numbers put GPT-6 Astra between 62.7 and 99.9 percent on the same benchmark depending purely on scaffolding, so a headline score tells you almost nothing about your work. A seven-task evaluation set takes about an hour to build and answers the question directly.

Why does a smarter model sometimes produce worse writing for me?

Because writing quality for your business is mostly a function of context and constraints, not raw reasoning power. When you switch models, the scaffolding you tuned for the old one stops fitting. The model did not forget how to write. It never knew your rules, and the structure that used to carry them across did not come with you.

Is the 95 percent enterprise AI failure rate real?

MIT's Project NANDA reported that 95 percent of enterprise generative AI pilots produced no measurable profit and loss impact, based on executive interviews, leader surveys, and analysis of public deployments. The important detail is the cause. The researchers attributed the divide to approach rather than to model quality or regulation, which is good news for anyone willing to fix their approach.

Do I need technical skills to build a harness?

No. Every component is a document, a setting, or a habit. Context files are written in plain language. Standing instructions go in a saved project. Tools are toggles. The evaluation set is seven pieces of your own past work. If you can write an onboarding document for a new hire, you can build a harness.


The close: two tabs, one decision

Go back to those two browser tabs.

One person is watching a benchmark and calling it the future. Another person is watching their own workflow get worse and calling it a scam. Neither one is lying. Neither one is measuring.

That is the entire gap, and it is going to keep widening. Frontier capability is accelerating. Daily reliability is not moving at the same speed, in the same direction, or for the same reasons. Stanford's index calls the frontier jagged. I would call it something blunter: the demo and the desk have stopped being the same place.

So here is what I want for you, and I mean this.

I want you to stop being a spectator to model releases. I want you to stop reading launch posts like they are weather reports about your business. The score in the headline was produced by scaffolding somebody else built for tasks somebody else chose. It was never about you.

But the scaffolding around your work? That is yours to build. Nobody at OpenAI knows your clients. Nobody at Anthropic knows why your third paragraph always needs to land a certain way. Nobody at Google can write your context file. That knowledge is the thing you have been accumulating for years, and for the first time it is the scarce input rather than the obsolete one.

The model is a commodity. Your understanding of your own work is not.

Build the harness. Then it will not matter much what ships next month, because you will be able to test it in ninety minutes and know.

If you want the harness built alongside people doing the same thing, that is exactly what happens inside AI Insiders. It is $247 per month, and we build these systems together in the open, with real files, real tests, and real numbers.

P.S. The most useful thing I did this week was not upgrading. It was running seven old tasks against a new model and keeping one workflow on the old one. Unglamorous. Nobody upvotes that. It works.


About the Author

Jonathan Mast is the founder of White Beard Strategies LLC, where he provides AI coaching and mentorship for entrepreneurs. He teaches non-technical business owners how to use AI to amplify the skill and experience they already have, in plain language, without the hype. He runs a Facebook community of more than 500,000 members, the AI Insiders membership at $247 per month, and ongoing live training programs. He gives away the knowledge and charges for the access and the implementation.

About the Author