Which may suit your business
The AI release cycle has become a business problem.
Not because the models are getting worse.
Because they are getting better, faster than most business owners can sensibly evaluate them.
Two frontier models landed within three weeks of each other in September 2026:
- GPT-6 Astra from OpenAI, announced on 3 September.
- Claude Opus 5.5 from Anthropic, announced on 22 September.
Both are exceptionally capable.
Both can reason, research, write, code, use computers and work across long, complicated tasks.
So which one should your business use?
The honest answer is: it depends on the work.
And before you spend another week comparing shiny tools, remember this:
The model matters less than the workflow you apply it to.
Meet the contenders
GPT-6 Astra: built for the frontier
OpenAI describes GPT-6 Astra as its most capable and aligned model to date.
It is designed for demanding work involving:
- Advanced computer use.
- Software engineering.
- Cybersecurity.
- Scientific research.
- Long, multi-stage reasoning.
- Large document and codebase analysis.
OpenAI reports a context window of around 1.05 million tokens, with an output limit of roughly 128,000 tokens.
In plain English, it can work across an enormous amount of information in one task.
OpenAI has used the phrase 'AGI era' when discussing Astra. That is OpenAI's characterisation of the technology's direction, not an established legal, scientific or commercial standard, and it should not be treated as a guarantee of artificial general intelligence or business performance.
Astra is rolling out across ChatGPT Pro, Business and Enterprise, the OpenAI API, Microsoft Azure and AWS Bedrock.
One important detail for Australian SMEs: Astra is not included in the standard ChatGPT Plus plan.
Published list pricing is approximately:
- $10 per million input tokens
- $50 per million output tokens
That makes Astra powerful, but not necessarily the sensible default for routine emails, meeting summaries or basic admin.
Claude Opus 5.5: frontier capability with a sharper cost profile
Anthropic launched Claude Opus 5.5 as the first model in the Claude 5.5 family.
Its headline achievement is not simply raw intelligence.
It is the combination of high-end performance, speed and token efficiency.
Anthropic says Opus 5.5 performs at roughly the level of Claude Fable 5.1 on most work, while costing around 40% less to run than the previous Opus generation.
It is particularly strong in:
- Agentic coding.
- Knowledge work.
- Computer use.
- Business workflows.
- Long-running research and analysis.
- Clear, practical communication.
Its published list pricing is approximately:
- $4 per million input tokens
- $20 per million output tokens
- $0.20 per million cache reads
Anthropic also says Opus 5.5 generates output more than 30% faster than Opus 5.
That matters when you are running AI workflow automation repeatedly across a business.
Opus 5.5 is available through the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry.
Where Fable 5.1 fits
Claude Fable 5.1 is not the main event in this comparison, but it helps explain the landscape.
Released on 1 September 2026, Fable 5.1 remains one of Anthropic's higher-capability models for demanding, long-running work.
It has a one-million-token context window and a 128,000-token maximum output.
It is the benchmark Opus 5.5 is being measured against.
Fable can still make sense for the most ambitious, long-running work. But Opus 5.5 is designed to deliver much of that capability more efficiently.
Head-to-head at a glance
Important benchmark note: The benchmark results mentioned in this article are reported by the relevant model vendors or third parties and have not been independently reproduced by Evolve With AI. The results may use different model versions, prompting methods, tools, safeguards, effort settings, evaluation harnesses, sample sizes and scoring methods. They should not be treated as a like-for-like ranking or as a prediction of performance in your workflow. The figures and product details reflect information available as at September 2026 and may change.
| GPT-6 Astra | Claude Opus 5.5 | Claude Fable 5.1 | |
|---|---|---|---|
| Release date | 3 September 2026 | 22 September 2026 | 1 September 2026 |
| Maker | OpenAI | Anthropic | Anthropic |
| Signature strength | Computer use, science, cybersecurity and broad agentic work | Coding, knowledge work, computer use and efficiency | Long-running, high-complexity reasoning and research |
| Context window | Around 1.05 million tokens | Large context; confirm current limit for your selected platform | 1 million tokens |
| Published list pricing | Approx. $10 input / $50 output per million tokens | Approx. $4 input / $20 output per million tokens | $10 input / $50 output per million tokens |
| Availability | ChatGPT Pro, Business and Enterprise, API, Azure and AWS Bedrock | Claude API, AWS, Google Cloud and Microsoft Foundry | Claude API and major cloud platforms |
| Best suited to | Technical, scientific and computer-use-heavy work | High-volume coding, research and business workflows | The hardest long-running knowledge and coding work |
All prices above are approximate USD API list prices per million tokens and can change.
Benchmark results, context limits, availability and pricing are not necessarily directly comparable. 'Large context' and similar descriptions may vary by platform and configuration. Confirm current specifications with the relevant provider before making a purchasing decision.
Pricing shown is indicative provider list pricing, not a quotation from Evolve With AI. Pricing and access may differ between direct APIs and third-party platforms such as Azure, AWS Bedrock, Google Cloud or Microsoft Foundry.
Where Astra pulls ahead
Astra’s strongest advantage is breadth.
It is not built for one narrow category of work. It is designed to operate across computers, code, research environments and technical systems.
1. Advanced computer use
Astra is built to work with software interfaces and browsers at a high level.
That could make it useful for workflows involving:
- Moving information between systems.
- Navigating complex business software.
- Testing websites or applications.
- Conducting multi-step online research.
- Handling repetitive digital processes.
This does not mean you should give it unrestricted access to your systems.
It means Astra is particularly interesting when the work involves doing, not just answering.
2. Scientific and research tasks
OpenAI reports strong results in scientific and agentic research evaluations.
For businesses working with technical documentation, complex research or specialist analysis, that breadth could be valuable.
3. Cybersecurity capability
OpenAI says Astra is the first model to reach the Critical cybersecurity threshold under its Preparedness Framework.
OpenAI reports a 100% result on ExploitBench under its published testing conditions. That result should not be treated as a measure of the publicly deployed model's unrestricted capabilities, a guarantee of cybersecurity performance or evidence that the model can safely operate as an autonomous security team.
What it does mean is that access, monitoring and human review matter more than ever.
4. Distribution
Astra’s availability across ChatGPT, the API, Azure and AWS Bedrock gives it a broad enterprise footprint.
For organisations already invested in the OpenAI ecosystem, that distribution may reduce integration friction.
Where Opus 5.5 pulls ahead
Opus 5.5’s advantage is more focused.
It is designed to do serious work without consuming as much time, money or output as previous frontier models.
1. Agentic coding
Anthropic reports that Opus 5.5 achieved a leading result in its selected agentic coding evaluations, including a 66.4% result on Terminal-Bench 4.0.
The practical message is more useful than the score:
Opus 5.5 is designed for large, messy, multi-step technical jobs.
That includes codebase audits, migrations, debugging and software work that runs for hours rather than minutes.
2. Knowledge work
Opus 5.5 reports an 1846 Elo score on GDPval-AA v2.1, covering professional work across 44 occupations.
That points to strong performance in research, analysis, reporting and other business tasks.
For an SME, this could translate into better support for:
- Preparing decision briefs.
- Analysing documents.
- Producing first-draft reports.
- Comparing business options.
- Turning scattered information into a usable recommendation.
3. Computer use
Opus 5.5 reports an 81.8% partial result on OSWorld 2.0 under Anthropic's published testing conditions. Anthropic also reports a lower strict score, illustrating why the scoring method and evaluation conditions matter.
Again, the practical point is not that the model should run your business unattended.
It is that it may be capable of handling more complex digital workflows with fewer steps.
4. Token efficiency
This is where Opus 5.5 is especially compelling.
Anthropic reports that Opus 5.5 matched Astra on a selected coding evaluation at approximately 40% of the reported cost per task, and outperformed Astra on a selected knowledge-work evaluation at approximately one-fifth of the reported cost per task. These are vendor-reported comparisons under specified testing conditions, not independent forecasts of customer savings. Actual total cost will depend on prompts, context, tool calls, retries, model settings, platform charges, human review and workflow engineering. Businesses should test their own representative tasks before relying on a cost comparison.
But if you are running thousands of automated tasks, efficiency may become a strategic advantage.
The uncomfortable truth about benchmarks
Benchmark figures are useful.
They are also easy to misuse.
The results above are vendor-reported or third-party reported. Testing conditions differ. Models may use different effort settings, tools, safeguards, harnesses and numbers of trials.
A result from one evaluation is not a guarantee about your workflow.
Benchmark margins at this level may be less reliable as a guide to real-world differences.
The gap between Astra and Opus 5.5 may look large on a chart, while feeling surprisingly small in day-to-day use.
The real question is not:
“Which model is smarter?”
It is:
“Which model fits the work we need done, at an acceptable cost and risk?”
So which one may suit your business?
The following observations are general information only. They are not a recommendation that any business select a particular model. Suitability depends on the business's objectives, data, systems, budget, risk profile, contractual arrangements and testing results.
Use this as a general starting framework for discussion, not a recommendation or substitute for assessing your own requirements.
| Your priority | General starting point | Matters to test |
|---|---|---|
| Heavy document and research work | Test both against your own documents. Opus 5.5 looks particularly compelling for efficient knowledge work, while Astra may suit broader technical research. | Accuracy depends heavily on source quality, retrieval and review controls. |
| Coding and technical build work | Start with Opus 5.5 for large, repeated coding workflows. Consider Astra where computer use, science or cybersecurity breadth matters. | Neither should merge code or change production systems without appropriate testing. |
| Everyday admin and productivity | Do not automatically use either frontier model. A lower-cost model may be enough. | Paying frontier prices for simple drafting is tool-chasing, not strategy. |
| Regulated or high-stakes decisions | Focus first on governance, data handling, human review and auditability. | A powerful model does not remove your legal, professional or operational responsibility. |
| Cost sensitivity | Opus 5.5 deserves close attention because of its token efficiency. | The cheapest token price is not always the lowest total workflow cost. |
The part that matters more than the model
A brilliant model with no business context is a fast way to produce confident nonsense.
The same model can be excellent in one workflow and almost useless in another.
Why?
Because performance depends on:
- The quality of the information it receives.
- Whether your documents are current.
- How clearly the task is defined.
- Which systems it can access.
- Whether it has permission to take action.
- How outputs are checked.
- Where a human must intervene.
Most SMEs will get more return from properly grounding one model in their business than from switching between four different tools every week.
That is the difference between Claude AI automation and randomly asking Claude questions.
It is the difference between a useful Claude ecosystem and another disconnected subscription.
It is also why “Claude for Small Business” should not be treated as a magic button. The value comes from the process around the model.
Start with the workflow.
Then choose the software.
Australian considerations
Both Astra and Opus 5.5 are available through major cloud platforms and offer enterprise data-handling options.
Before putting customer information into either system, check:
- Where data is processed and stored.
- Whether inputs are used for training.
- Retention periods.
- Available zero-data-retention settings.
- Sub-processors and overseas disclosures.
- Access controls and audit logs.
- Whether your proposed use complies with your own Privacy Act obligations.
From 10 December 2026, relevant APP entities will have additional privacy-policy transparency obligations where a computer program uses personal information to make, or substantially and directly assist with, a decision that could reasonably be expected to significantly affect an individual's rights or interests. Whether the obligation applies depends on the organisation, the information used and the decision-making process. Using a frontier model does not by itself determine whether the obligation applies. Before connecting customer or employee data, businesses should assess their Privacy Act coverage, collection notices, overseas disclosures, vendor terms, retention settings, security controls and human-review processes. Many small businesses may be exempt from the Privacy Act, subject to important exceptions, so the position should be checked rather than assumed; our detailed guide covers Privacy Act considerations for Australian SMEs.
Do not assume that an enterprise plan, a no-training setting or Australian-facing availability by itself resolves privacy, security or APP 8 issues.
The model is not your compliance program.
A remarkable moment, but not a strategy
Astra and Opus 5.5 are genuinely remarkable.
Astra is positioned by OpenAI as a broad-capability model across computer use, technical work, cybersecurity and scientific research.
Opus 5.5 is positioned as a frontier model with a strong reported efficiency profile.
Fable 5.1 remains relevant for the most demanding long-running work.
There is no honest overall winner.
And there probably should not be one.
Your business does not need the model with the most impressive launch announcement.
It needs the model that can improve a real workflow, safely, repeatedly and at a cost that makes sense.
That is why our approach is strategy first, software second.
Our free AI Discovery Audit — your first session is free, with no obligation, no minimum spend and no credit card required — is designed to help identify which workflows may benefit from a frontier model and which may be better served by a simpler approach.
We can also help with AI strategy and roadmap planning, practical AI business automation Australia projects and hands-on AI productivity training.
If your team wants practical guidance rather than another slide deck, our workshops, including Claude workshops where relevant, can help people use these tools confidently in their actual day-to-day work.
No pressure and no vendor preference — just a practical discussion about what may fit your business.
This article provides general information only and is not legal, financial, accounting, privacy, cybersecurity or technology procurement advice. It does not take into account your objectives, circumstances, systems, data, budget or needs, and should not be relied upon in place of advice from a qualified professional. Product features, availability, plan entitlements, pricing and benchmark results are based on vendor or third-party information available as at September 2026 and may change. Benchmark results have not been independently reproduced by Evolve With AI and are not a guarantee of performance, savings, accuracy, compliance or suitability in your workflow. Evolve With AI does not represent or warrant that any particular model or platform is suitable for your business. Nothing in this article creates a client, advisory or professional relationship with Evolve With AI. Any AI output must be reviewed by an appropriately authorised person before it is relied upon or sent to a customer.
OpenAI, GPT, ChatGPT, Astra, Anthropic, Claude, Opus, Fable and related names and logos are trademarks or brand assets of their respective owners. Evolve With AI is independent of, and is not affiliated with, endorsed by or partnered with OpenAI or Anthropic unless expressly stated otherwise.
This comparison is a snapshot, not a permanent verdict. These models, prices and availability details will continue to change.
