The Same AI Model Scored 30%, Then 100%. Nvidia Only Changed Its Setup.
Nvidia changed the tools, memory and rules around Claude Opus 5 and nothing else. For marketing teams, that moves the highest-return hour of the week.
Nvidia published a result on August 21 that should change what marketing teams argue about in their next tooling meeting. Researchers took Anthropic’s Claude Opus 5, which scores about 30% on a reasoning test called ARC-AGI-3, and got it to 100%.
They didn’t use a better model. The model was identical in both runs. What changed was the setup around it.
That setup has a name in the industry. It’s called a harness: the tools, memory, rules and supervision you build around a model so it can work on its own for more than a moment. Adel El Hallak, Nvidia’s VP of product for AI, put the point bluntly:
“It is the model. It is the scaffolding around the model, which we call the harness, i.e. the set of tools that it utilizes.”
Nvidia’s system is called AVO, short for Agentic Variation Operators. It finished all 183 levels across 25 test environments using 6,624 moves, against 7,542 for the previous best system. Better score, roughly 12% less work.
Why this lands on a marketing desk
Most teams have spent the last two years treating AI quality as a shopping decision. Which model, which vendor, which tier. That framing survives because it’s easy to act on. You can switch a dropdown in an afternoon.
This result says the dropdown was never the biggest lever.
If the same model can go from a third right to everything right based purely on what surrounds it, then the gap between your AI output and someone else’s is mostly not about which company you pay. It’s about what you built around the thing you’re paying for: the instructions, the reference material it can reach, the checks that catch it when it drifts, the memory of what it did last time.
Marketing teams already own all of that. Nobody calls it a harness. They call it the brand guidelines doc, the approved-claims list, the review step, the brief.
Addy Osmani, an engineering leader at Google, wrote the plainest definition of the term in a piece on harness engineering in May. A harness is:
“every piece of code, configuration, and execution logic that isn’t the model itself.”
His summary of what it buys you is the sentence to carry into the meeting: a decent model with a great harness beats a great model with a bad harness.
The caveats Nvidia printed and most coverage dropped
This is where the story gets more useful, not less.
Nvidia states in its own post that the 100% covers the public set of tests only, and not the semi-private or private sets held back for real competition. The company also warns that comparing its system against a bare model “should not be interpreted as a controlled ablation,” meaning the two setups differ in too many ways to isolate one cause. And it says directly that Claude Opus 5’s standalone 30% is not a measurement of how much AVO contributed.
The honest version of the finding is narrower than the headline number. Nvidia is careful not to claim the harness is worth 70 points. Its claim is that judging a model on its own tells you very little about how the finished system performs.
That weaker claim is still the one that matters to you. It means benchmark scores, the numbers vendors put in launch posts, don’t predict what you’ll get. We found the same thing from the other direction in July, when a benchmark jump never reached the marketing copy.
The mechanism, in plain terms
The reason a wrapper can matter this much comes down to what a model actually is. It answers one prompt at a time and remembers nothing between calls unless something hands it the history.
Nvidia’s system adds a supervisor: a second agent whose only job is to watch the first one, spot a dead end, and tell it to change strategy. Most of the gain has that shape.
A marketing team’s version of a supervisor is a person reading the output before it ships, which most teams have. The difference is that Nvidia’s runs on every step rather than at the end, and it changes the plan rather than fixing the draft.
What this costs you if you ignore it
Budgets are already tight on this. Businesses have stopped raising what they’ll spend on AI, which means the next gain has to come from something other than buying a better tier.
So the practical read is a reallocation, not a purchase.
Stop A/B testing models as your first move. If two models give you similar mediocre output, the shared cause sits upstream of both.
Write the brief down where the tool can read it. Most teams keep positioning, banned claims and tone rules in a doc a human opens and a model never sees. That’s an unbuilt harness.
Add one check that runs every time, not at the end. A claims check, a source check, a tone check. The Nvidia result says mid-process correction is worth more than final review.
Keep it short. Osmani’s guidance is to keep the instruction file under 60 lines and make every rule trace to an actual failure. Long rule sets read as thorough and behave as noise.
The honest caveat on all of this: Nvidia sells the computers that agent systems run on, and a result arguing that system design matters is a result arguing for more of what Nvidia sells. The finding still holds, because the company undercut its own headline in its own post, which isn’t what a pure marketing claim does.
The model you picked is probably fine. The hour you were going to spend comparing it to another one is better spent writing down what your team already knows and never told the machine.
Quoted in this story
- Adel El Hallak, VP of Product, AI, Nvidia (source)
- Addy Osmani, Engineering Leader and Author, Google (source)
Want your perspective in coverage like this? Get quoted.
Sources
This story is part of our running coverage: the full picture →
Get Net Effect.
The net effect of AI on your marketing: the stories that matter, twice a week, in five minutes.


