Analysis · 12 September 2026
There is a moment when an AI stops feeling like a search box. You give it an untidy brief, some files and an outcome you want. It researches, makes something useful, responds to corrections and carries on. The natural question becomes: if this is not artificial general intelligence, what exactly are we waiting for?
ChatGPT makes that question increasingly reasonable. But a convincing experience and a scientific conclusion require different kinds of evidence. Our assessment is that the latest developments strengthen the case for broadly useful AI. They do not, by themselves, establish that AGI has arrived.
Why the question feels different now
The latest reason to revisit the question is OpenAI’s GPT-6 Astra announcement on 3 September. Its reported gains include computer use and professional tasks. In a follow-up on 8 September, OpenAI described expanding opportunities to delegate work, while acknowledging that people still choose research priorities and judge results.
That change matters to ordinary users. A helpful answer leaves you to do the work. A useful deliverable removes part of the work. When the same interface can support research, analysis and creation, the experience starts to resemble working with a versatile assistant.
OpenAI’s ChatGPT documentation describes workflows that turn goals, files and context into documents, spreadsheets and presentations. That is a broader proposition than fluent conversation. Yet ChatGPT is a product containing models, tools and configuration choices. An outcome from one model in one environment should not automatically be attributed to every ChatGPT session.
The feeling of general intelligence can arrive before the evidence needed to establish it.
What would count as AGI?
Before declaring a finish line crossed, we need to say where it is. Does AGI mean doing many useful tasks, matching skilled humans across a wide range of cognitive work, or independently learning unfamiliar tasks? Those are different standards.
The research paper Levels of AGI, by Meredith Ringel Morris and colleagues, offers a helpful framework: consider the depth of performance and the breadth of capability, while treating autonomy as a related deployment dimension. It is a proposed framework, not a universally accepted certification.
This gives us a better set of questions. How well does the system perform? Across how many different situations? How much human help does it need? A brilliant specialist and a dependable generalist may both be valuable, but they are not interchangeable.
Nor does this require settling whether a machine has a private inner experience. The article’s question is about capabilities and their evidence. A fluent explanation of intelligence is not a test of intelligence, and sounding human is not proof of consciousness.
The numbers show progress—and limits
OpenAI reports Astra scores of 41.4% on AutomationBench, 59.3% on Agents’ Last Exam and 72.6% on an offline, partially scored OSWorld 2.0 evaluation. All three exceed its reported GPT-5.6 Sol results. The chart compares those specific evaluations.

The testing conditions matter. OpenAI says these are maximum scores across effort levels, obtained in research or API environments that may differ from production ChatGPT. The OSWorld result uses the v2026.08.08 offline set with partial scoring. These numbers therefore cannot be read as your probability of success on an arbitrary office task.
The strongest headline deserves attention too: OpenAI reports 99.9% on ARC-AGI-3. That is a striking result on a test associated with generalisation. Having “AGI” in a benchmark’s name, however, does not make its score a certification of general intelligence.
A benchmark is useful because it makes a defined question testable. The danger comes when we quietly replace that question with a much larger one. Success on a particular suite can support a claim about capability; establishing broad, dependable general intelligence requires evidence beyond that suite.
Why useful work can feel like a breakthrough
Consider a hypothetical content workflow: turn a messy source document into a clear article, create supporting visuals and prepare it for editorial review. The impressive part is the connection between activities. The system must carry the brief from reading to writing to design, rather than treating each step as unrelated.
But inspect the whole process. Did someone repair the brief? Were the sources checked? Did a human catch a misleading chart? Was the output merely saved, or was it actually reviewed? Those details determine how much responsibility the system carried.

There is an equally important trap in the opposite direction. A blocked login or unavailable tool is not proof that a model lacks intelligence. OpenAI’s browser documentation identifies access restrictions, CAPTCHA and deployment availability as potential obstacles. Product access, reasoning quality and reliability deserve separate examination.
Test the work you actually need
Choose a small set of representative tasks and define success before starting. Record the quality of the result, the corrections required and the human time spent checking it. Repeat with unfamiliar inputs. Count failed attempts and review time, not just the most impressive demonstration.
So, has AGI arrived?
“It helped me do something difficult” is strong evidence of usefulness. “It can reliably handle unfamiliar work across a broad range of domains” is a much larger claim. The first should encourage practical experimentation; the second needs a clear standard, reproducible testing and scrutiny beyond the company selling the system.
We should also be willing to update our assumptions. Our earlier article on AI and human copywriters focused on the guidance content tools required. That remains a useful question, but historical limitations should be retested against current systems rather than repeated as permanent truths.
A meaningful shift, with an open question
ChatGPT has become useful across enough activities to make the AGI question feel immediate. The evidence reviewed here supports substantial progress, not a settled declaration. For readers and businesses, the next step is concrete: test what it can complete, measure what you must correct, and keep the label separate from the results.
Editorial note: This is a source-based analysis, not an independent benchmark or AGI certification. Sources were reviewed on 12 September 2026. Featured and inline illustrations were generated with the assistance of AI models; the chart was constructed from the cited evaluation table.
Let's Explore What's Possible
Whether you're tackling a complex AI challenge or exploring new opportunities, we're here to help turn interesting problems into innovative solutions.