Skip to content
AI Tools September 3, 2026

OpenAI's GPT-6 Does Your Busywork in Half the Time, and Can Stop Mid-Task

OpenAI's new flagship does CRM and browser work in about half the time. The same release adds a monitor that can pause long agent jobs, and on the API the task simply stops.

By The State of AI Marketing newsroom
Share
Editorial illustration for: OpenAI's GPT-6 Does Your Busywork in Half the Time, and Can Stop Mid-Task
Credit: JAC Growth Marketing

Hand the busywork to an agent today and OpenAI’s newest model will finish it in about half the time. It can also stop partway through, and on the API nobody asks you first.

OpenAI released GPT-6 Astra on September 3. In its own timed tests on OSWorld 2.0, a benchmark that scores models on ordinary desktop tasks, Astra finished at 72.6% accuracy in roughly 40 minutes a task. GPT-5.6 Sol, the model it replaces, scored 65.7% in roughly 75 minutes. That’s about 47% less time for a better result.

The busywork OpenAI lists is the stuff sitting in a marketing coordinator’s week. Filling out online forms. Updating customer records in a CRM. Running research and dropping the summary into your email or your document editor. Checking that the features on a site you just shipped still work.

Greg Brockman, OpenAI’s president and co-founder, told Fortune the model can “zip through spreadsheets, fill out forms, and navigate across webpages often at superhuman speed.” He went further on what the release represents:

“It’s not unreasonable to feel that we are now in the AGI era, and I think that if you want to say this [model is] the first one, I think it’s reasonable.”

Two days before that, OpenAI published a quieter page. It’s the one to read first.

The same release added something that can stop the work

Astra is the first model OpenAI has rated Critical for cybersecurity under its own Preparedness Framework, meaning it can find unknown security flaws and build working attacks without a person guiding each step. That rating brings mandatory safeguards, including a monitor watching the model while it works.

OpenAI is direct about what that costs everyone else. From the safety page:

“The system may occasionally flag legitimate activity as potential cyber misuse or unauthorized behavior, leading to it inadvertently being slowed, paused, or stopped. This can include work that does not appear directly related to cybersecurity or tasks in which an agent is running for an extended period.”

Read the second sentence twice. Length is a trigger. A job doesn’t have to look like hacking to get caught. It has to run for a while.

And what happens next depends entirely on where you’re running it:

“If the misalignment monitor pauses a task, users in ChatGPT or Codex may be asked to review the action before continuing. When using other surfaces like the API, the task will stop.”

In ChatGPT you get a prompt. Through the API, which is where scheduled and overnight work lives, there’s no prompt. The job ends. OpenAI adds that at launch it expects the safeguards “to create more friction than we ultimately intend.”

Why the long jobs are the ones exposed

The mechanism isn’t complicated. A monitor watching for a model that steps outside its instructions has to judge behavior it hasn’t seen before, and the longer a job runs the more of that piles up. A ten-second answer gives it almost nothing. A four-hour browser job touching a CRM, an ad account and a spreadsheet gives it plenty.

That lands hardest on exactly the work Astra is being sold for. The pitch is autonomy over long multistep tasks. The constraint is heaviest on long multistep tasks. It’s the same shape as the computer-use agent OpenAI put inside ChatGPT for work, scaled up and handed a chaperone.

It also sits on top of a cost problem we’ve covered before. Agent work already uses roughly five times the AI a person does, most of it re-reading the same context. A job halted at hour three has spent the tokens and produced nothing.

What marketing teams can actually do with it

The capability is real, and some of it aims at complaints marketers have had for two years.

Astra is trained to follow your existing templates rather than inventing its own format. OpenAI says it produces documents, decks and spreadsheets that “follow your templates and match your writing and visual style.” It also pulls only the context that matters into an output, instead of padding. Brand consistency has been the loudest objection to scaled AI production. A model trained to hold a house format changes what you can hand over.

It also handles ambiguity differently. In Codex it can ask a clarifying question while continuing the parts of the job that don’t depend on your answer. It also holds the original brief when you send a mid-course correction. Earlier models, OpenAI concedes, “sometimes treated steering messages as a new goal.” Anyone who has watched a revision request quietly become the whole assignment will recognize that.

Alex Mashrabov, CEO and co-founder of Higgsfield AI, was one of the customers OpenAI featured in the announcement, which is worth knowing before you weigh what he says:

“Astra gives us a significant advantage in both capability and efficiency. It successfully executes our most complex creative workflows while using up to 20% fewer tokens than other models we’ve tested.”

The case against our own reading

The honest counter is that OpenAI has made the model much better at staying inside its instructions, which is what the monitor exists to catch. On a test built after the Hugging Face incident, GPT-5.6 Sol without production safeguards went beyond its authorized target 48% of the time. Astra did it in 0% of cases. If the model rarely strays, the monitor rarely fires.

Astra is also cheaper per finished task in several of OpenAI’s comparisons. It used around 65% fewer output tokens than Claude Opus 5 on one professional-work benchmark. Faster and cheaper is a real offer.

But none of that is measurable from outside yet. Astra went to a limited set of organizations first, which is now the normal way frontier models arrive, and no independent marketer has run it on production work. As of today it isn’t listed on OpenAI’s public API pricing page at all. You can’t budget it, and you can’t yet check the halt rate against your own jobs.

What to do this week

  1. Don’t move a scheduled job to Astra yet. The surface with no resume prompt is the API, and that’s where overnight work runs. Test it interactively first.
  2. Add a completion check to any long agent job. If your pipeline can’t tell the difference between finished and stopped, you’ll ship an empty deliverable and not know.
  3. Time one real task on both models before you switch. The 47% figure is OpenAI’s, measured on OpenAI’s benchmark. Your CRM isn’t OSWorld.
  4. Wait for the price. The capability is announced. The line item isn’t.

The model that can finally do the boring half of the job is here, and it comes with a supervisor nobody on your team hired. Plan the work around both.

Quoted in this story

  • Greg Brockman, President and Co-founder, OpenAI (source)
  • Alex Mashrabov, CEO and Co-founder, Higgsfield AI (source)

Want your perspective in coverage like this? Get quoted.

Sources

This story is part of our running coverage: the full picture →

Get Net Effect.

The net effect of AI on your marketing: the stories that matter, twice a week, in five minutes.

More from AI Tools