research - 2026-09-04
New AGI model — GPT-6 Astra?
Compare GPT-6 Astra, GPT-5.6 Sol and Claude Fable metrics for computer use, coding and research, and read Perfectory’s case for migrating its apps.
Should we migrate all our apps to GPT-6 Astra?
Our answer is yes: Astra is the direction we want to take across our applications. The reason is practical. In our early hands-on experience, it is fast and effective when working on our PC, handles routine tasks well, and brings a deeper understanding to research.
The “new AGI model?” question makes a good headline. It is not a conclusion we can establish from a benchmark. For our team, the more useful question is whether the model improves the work our applications actually do.
Astra vs Sol vs Fable: the published numbers
The next three tables reproduce selected results from OpenAI’s Astra announcement. These are vendor-reported scores, not Perfectory measurements. Higher is better. NR means not reported. Scores use the best evaluated effort; production results can differ.
Computer use and routine work
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | Fable 5.1 | Fable 5 |
|---|
| Agents’ Last Exam | 59.3% | 53.6% | NR | 48.7% |
| ScreenSpot-Pro (no tools) | 92.7% | 76.9% | NR | 87.3% |
| AutomationBench | 41.4% | 18.1% | 31.4% | 17.4% |
Source: OpenAI — computer use and professional evaluations. Fable columns refer to Claude Fable.
Software engineering
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | Fable 5.1 | Fable 5 |
|---|
| Terminal-Bench 4.0 | 57.9% | 37.3% | 55.8% | 42.0% |
| DeepSWE v1.1 | 74.1% | 72.7% | 67.4% | 69.9% |
Source: OpenAI — coding evaluations.
Research and reasoning
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | Fable 5.1 | Fable 5 |
|---|
| Terminal-Bench Science 0.1 | 64.6% | 22.4% | 52.6% | 21.4% |
| GPQA Diamond | 96.0% | 94.6% | 93.7% | 92.6% |
| Humanity’s Last Exam (with tools) | 57.2% | NR | 65.0% | 63.8% |
Source: OpenAI — academic evaluations. Fable leads on Humanity’s Last Exam; Astra does not win every comparison.
What we see in our own work
Our early assessment is qualitative: this article does not publish an internal timing dataset, sample size, or controlled Sol/Fable comparison. The observations below explain our decision without turning impressions into invented percentages.
| Area we assess | Our early Astra observation | Measurement to record across Astra, Sol and Fable |
|---|
| Working on our PC | Fast and effective in hands-on use | Median and p95 time to a verified completed task |
| Routine tasks | Useful execution of everyday work | Completion rate and human corrections per task |
| Research | Deeper understanding of the subject | Evidence quality and reviewer-rated correctness |
| Clarifying questions | Raises questions we had not seen other models ask in our use | Material ambiguities resolved before delivery |
The most interesting difference for us is the questioning. Astra raises follow-up questions that expose missing context and help us understand the subject more deeply. In our experience, these are questions other models often missed. That is an observation about the tasks we tried, not a claim that another model could never ask them.
For a research assistant, a good follow-up question can change the value of the final answer. For an app that operates a computer, the result we care about is a task completed correctly with little supervision. Those are the outcomes we want our migration to improve.
What migrating all our apps means
We want Astra to become our default across applications. We will move each application through its existing evaluation and release process, starting with PC workflows, routine automation, and research assistance. That gives us a concrete way to check the gains we see in everyday use.
We will compare the same tasks, inputs, tools, and acceptance criteria across Astra, Sol, and Fable. Alongside completion time, we will record correctness, human intervention, and cost per successful task. A fast response is useful only when the complete workflow produces the right result.
Our earlier model-migration work reinforced the value of checking the actual task. The current decision is to move forward with Astra; the application-level checks tell us when each integration is ready.
Our conclusion: yes, we should migrate
Yes, we should migrate our applications to GPT-6 Astra. Our early experience is that it is fast and effective on our PC and on routine tasks. It also brings deeper understanding to research, asks valuable extra questions that other models have missed in our work, and helps us understand the subject more thoroughly.
That combination is why we want to make the change. The published comparisons inform the decision, while our practical experience gives it direction. We will quantify those observations as each app moves through evaluation, then roll out the migration application by application.
FAQ
Is GPT-6 Astra an AGI model?
This article poses that as a question. Benchmark performance alone does not establish AGI.
Are these Perfectory’s own benchmark results?
The numerical comparisons come from OpenAI’s announcement. Perfectory’s observations are qualitative; we have not published an internal timing dataset or controlled comparison here.
Planning an AI automation initiative? Read the AI automation POC guide.