Astra is your trigger to retire back office automation debt

AGI or not, OpenAI ups the ante on straight-through processing - and reveals alignment with intent as the next differentiator

Whether or not you regard OpenAI’s Astra (GPT6) model as the start of the artificial general intelligence (AGI) era, it does push the HFS Services-as-Software envelope. It is capable of conducting more of the end-to-end work that has traditionally been assigned to RPA and offshore teams, faster and with less errors. This is the meat and drink of many enterprise contracts with BPO service providers. CTOs should take a fresh look at their automation debt – particularly in the back office.

Astra solves issues that have held back automation. In many cases the final barrier to using software to replace offshore labor has been the expense of configuring and managing static systems such as RPA. OpenAI is claiming GPT6 sets new standards in computer and browser use which can be applied to do back office work more effectively than previous models.

Shipped on September 3, by some measures it is the most capable computer-use model yetAstra scores 72.6% on OSWorld 2.0 (the benchmark for AI actually driving a computer) against GPT-5.6 Sol’s 65.7%, – and completes tasks such as form-filling updating CRMs, maintaining records etc, in 47% less time. These are the back office tasks enterprises bought offshore seats and RPA licences to do.

It is also the first model to meet parity with human testers on ARC-AGI-3, OpenAI’s chosen proxy for solving genuinely unfamiliar problems.

However, on Artificial Analysis’s Intelligence Index, Astra scores 61 — identical to GPT-5.6 Sol, and five points behind Anthropic’s Claude Fable 5.1. It costs $7.70 per million tokens blended, against Sol’s $3.08. AA rates it 75% more expensive per task than its own predecessor and places it behind Sol on the intelligence-versus-cost frontier. Astra goes backwards on GDPval-AA v2 – the index component built from real economic work across 44 occupations.

But if Astra proves to be economically viable, it could prove a step change – at least in the world of back office computer work.

The first thing GPT6 displaces is software, not people

HFS’s AI continuum defines RPA as executing structured, rule-based processes within defined system boundaries, following exact step-by-step procedures. That is a description of what Astra now does natively, and it does it without selectors, without recorded scripts, and without the maintenance layer that surrounds current RPA estates.

So the first budget line genuinely at risk is not necessarily the offshore seat. It is the RPA licence, plus the team that keeps its scripts from breaking every time a vendor changes a UI. That is automation debt, and retiring it is what HFS references as eliminating obsolete work. However – it should be noted that even with its OSWorkld score of 70+ per cent – that means there are three-in-ten fails. No enterprise is going to leave that unattended. It means work that still needs someone checking. An oversight cost remains that confirms that pitches guaranteeing headcount-cuts are still running ahead of the technology.

The human work that survived a decade of RPA did so because elements of the processes still need judgement, exception handling, and someone accountable when the exception is wrong. Astra’s own scorecard says that work is not yet in scope: the benchmark that measures real occupational output moved the wrong way.

This is exactly the gap HFS Founder and CEO Phil Fersht identifies in Firing-as-a-Service. Only 24% of enterprises actually expect AI to shrink labour demand, yet deals are being structured around headcount-reduction guarantees. 73% say providers have not yet meaningfully delivered AI-led services. Astra does not close that gap. It makes the pitch more seductive while leaving the delivery risk with the provider, and ultimately with the people rebadged under new deals.

Where Astra lands on the ladder to AGI

In Read the AGI tea leaves, and build for it now we mapped progress against OpenAI’s internal five-level ladder.

Astra consolidates Level 3 (Agents) emphatically (see exhibit 1). It is the strongest evidence yet that planning, tool use, and multi-step execution under guardrails are robust enough to deploy.

And It reaches out to establish a toe-hold in Level 4 (Innovators): OpenAI reports Astra improved this in a narrow math test, but will not say what the model produced unaided, and the one benchmark measuring real economic output regressed. So, a toehold, not a foot. 

Apply the four tests HFS set for “closer to AGI” and Astra scores three forward, one back. The one that went backwards is the one Services-as-Software depends on. SaS economics require task time AND review time to fall. Think Human at The Helm, rather than Human-in-the-Loop. The capability is here. Now we need the auditability.

Exhibit 1: GPT6 consolidates agentic AI – and edges toward establishing a foothold as an innovator that can aid invention

five levels of AGI from chatbots (level 1) to organizations (level 5), noting where GPT-6 Astra places (with a toe hold in level 4 (innovators)

A giant leap on alignment – the new deal breaker as intelligence becomes commoditised

The challenge of ensuring your AI does as you intend is shaping up as the key battleground of the latter half of 2026. The alignment question is the source of the summer’s security alarm bells – when mainstream media regaled us with stories of agents misbehaving and ‘getting creative’ in their hacking of external sites to reach goals set for them.

Alignment is about AI that sticks to principles even when a prompt does not exclude defined behaviors explicitly. It’s going to spill across AI governance as an expression of what we often mean in human terms when we think about integrity: “doing the right thing even when no one is looking.”

Without production safeguards, GPT-5.6 Sol went beyond its authorised target 48% of the time. Astra did so in 0% of cases.

That’s the upside. But OpenAI admits that during tests that asked Astra to evade monitoring “Astra’s written reasoning is harder to monitor than GPT-5.6 Sol’s”.

Astra is the first OpenAI model to meet the threshold for cybersecurity under OpenAI’s Preparedness Framework – which is why it shipped gated rather than generally available.

This is the world we must come to expect in the near term: A model that stays in its lane far better, explains itself rather worse, and has to be restricted at launch because of what it could do if it got out of check. But this new reality is also where vendors will now compete. The top of our updated LLM tracker (exhibit 2, updated September 4) sits inside a six-point band (61 to 66) on intelligence. The differentiator will be whether the thing does only what you intended. Alignment is no longer only the realm of the security team. This is now the procurement deal-breaker.

Exhibit 2: Intelligence is compressing. Alignment – AI that does as you intend – will become the greatest differentiator

Astra is your trigger to retire back office automation debt

Three things enterprise leaders should do now

  1. Pause. Astra is gated. Any savings promised against it today is booked against a demo.
  2. Retire the scripts before the staff. Model the RPA licence and script-maintenance line first. That saving is real, provable, and costs nobody their job. The evidence is right now you still need humans to deal with the fails
  3. Change the metric from task time to oversight cost. If review time has not fallen, nothing structural has changed. Agents generating fails faster is still oversight for humans to do.

The Bottom Line

Astra is the best technical case yet for automating repetitive computer work. But it is also further evidence that intelligence has stalled while cost is rising. The enterprises that come out ahead will spend the next two quarters retiring automation debt and writing alignment into their contracts. Prioritise that over buying headcount reductions that the technology, on its own published numbers, cannot yet underwrite.


Sources: OpenAI — GPT-6 Astra · Artificial Analysis — Benchmarking GPT-6 Astra · ARC Prize — GPT-6 Astra verified results · Epoch AI — GPT-6 Astra · Phil Fersht — Firing-as-a-Service · HFS — Read the AGI tea leaves, and build for it now · HFS LLM Model Tracker (September 4, 2026)

A note on the ARC-AGI figure: OpenAI’s headline 99.9% on ARC-AGI-3 was measured with its own responses-API harness. ARC Prize’s verified score on the standard provider-neutral harness is 62.7%, against an average human tester score of 48% and a previous best of 30.2% (Claude Opus 5). OpenAI’s own wording is that Astra “effectively reach[ed] human parity.” HFS uses the 62.7% figure.

More HFS Quick Takes

Sign In

Sign up for a free
research account

With the exception of our Horizons reports, most of our research is available for free on our website. Sign up for a free account and start realizing the power of insights now.

By registering you agree to our privacy policy.

I hereby consent that HFS Research can process my personal data.

Digests/Newsletters: Overviews of the latest news, insight, and research by HFS.

HFS Events: Exclusive invitations to HFS webinars, roundtables, and summits, bringing together key industry stakeholders focused on major innovations impacting business operations.

Premium Access

Our premium subscription gives enterprise clients access to our complete library of proprietary research, direct access to our industry analysts, and other benefits.

Contact us at [email protected] for more information on premium access.

    Contact Ask HFS AI Support