One year. A different class of AI.
One Year. A Different Class of AI. From GPT-5 to GPT-6 Astra. By Keith Lawrence Miller, M.A.

Review period: September 9, 2025–September 9, 2026. Scenario horizon: September 9, 2027. Evidence rechecked September 11, 2026.

Your opinion of AI may be one year out of date.

So may the way you organize your work.

The question worth reopening is how much of an assignment an AI system can carry from instructions to an acceptable result, and what you must still supply, supervise, and verify along the way.

GPT-5 was already capable of reasoning, coding, using tools, and handling multistep requests when this review period began. Treating that starting point as little more than a chatbot would erase capabilities that were already documented. [S01]

The more consequential question is how far the scope of delegation has expanded.

This review follows a selected OpenAI release sequence, with independent METR research providing context for longer assignments and future scenarios. It is not a census of the entire AI industry.

The central argument is straightforward: AI progress should change the assignments we test, while the evidence should determine the responsibility we delegate.

A year of changes to the work itself

The progression becomes clearer when releases are connected to the work they were designed to support.

Autumn 2025: more sustained execution. GPT-5-Codex arrived on September 15 with an emphasis on software work. GPT-5.1 followed in November, including developer tools for applying code changes and running shell commands. [S03, S04, S05]

Winter 2025–2026: stronger professional deliverables. GPT-5.2, released December 11, emphasized documents, spreadsheets, presentations, and longer projects. GPT-5.3-Codex followed February 5, extending coding capabilities into broader computer-based professional work. [S06, S07]

Spring 2026: broader use of applications. GPT-5.4 arrived March 5 with native computer use in a general-purpose model for Codex and the API. GPT-5.5 followed April 23 with further advances in professional-work evaluations. [S08, S09]

Summer into September: another expansion. GPT-5.6 Sol launched for general availability July 9. GPT-6 Astra followed September 3. [S10, S15]

My interpretation of this sequence is a widening boundary around the assignment: more reading, analysis, production, testing, and revision can potentially remain within one coordinated process.

That possibility matters to a professional deciding what to learn. It also matters to a manager deciding how to organize a team. Both need evidence about the actual work, beyond the release announcement.

One year: nine milestones
Selected release milestones from the GPT-5 baseline to GPT-6 Astra. GPT-5 was already available at the start of the review period. Sources: S01, S03–S11, S15.

Three ways to measure the change

1. Better answers on a bounded test

OpenAI reports 85.7% for GPT-5 at high reasoning effort on GPQA Diamond and 96.0% for Astra in its release table. The reported-score gain is 10.3 percentage points. [S05, S11]

Expressed as remaining error, the change is from 14.3% to 4.0%, a calculated reduction of approximately 72%.

Those are two descriptions of the same endpoints. Neither makes Astra “72% more intelligent.” These are reported evaluation configurations, rather than an independently rerun, matched-compute experiment.

The baseline also matters. GPT-5 Pro's published GPQA result was 88.4%; comparing that with 96.0% gives approximately 65.5% remaining-error reduction instead. Model variants and reasoning budgets belong beside the numbers. [S01]

GPQA Diamond by selected release
Reported GPQA Diamond results across selected releases. Configurations and reasoning effort vary; this is not a matched-compute time series. Sources S05, S06, S08–S11.

2. Better professional work products

Original GDPval measures performance on specified professional tasks. OpenAI's published win-or-tie rates against professional reference work progressed as follows. [S06, S08, S09]

ModelOriginal GDPval win-or-tie rate
GPT-538.8%
GPT-5.270.9%
GPT-5.483.0%
GPT-5.584.9%

The endpoint gain is 46.1 percentage points.

An 84.9% win-or-tie rate does not mean that 84.9% of jobs can be automated. The evaluation concerns defined work products across selected tasks, with reported configurations that vary. I have left Astra out of this series because the reviewed Astra table does not report a directly matching original-GDPval result. [S06, S08, S09, S11]

A similar discipline applies to coding. GPT-5's launch SWE-bench Verified result used 477 problems; a later comparison used all 500. Their 74.9% and 72.8% results should not be treated as interchangeable observations. [S01, S05]

Original GDPval professional-work results
Original GDPval win-or-tie rates against professional reference work. These are work-product evaluations, not percentages of jobs automated. Sources S06, S08, S09.

3. Uneven gains across different tasks

The same Astra release table reports these selected comparisons with GPT-5.6 Sol. Each row measures something different. [S11]

EvaluationSolAstraGain in percentage points
Terminal-Bench Science 0.122.4%64.6%+42.2
AutomationBench18.1%41.4%+23.3
MRCR v2, eight-needle, 512K–1M context73.8%96.3%+22.5
Terminal-Bench 4.037.3%57.9%+20.6
ScreenSpot-Pro, no tools76.9%92.7%+15.8
Agents' Last Exam53.6%59.3%+5.7
DeepSWE v1.172.7%74.1%+1.4

Scores are published maxima across reasoning efforts, evaluated in research or API environments that can differ from production ChatGPT. The rows should not be averaged into a general intelligence score. [S11]

For a business, the practical question is which evaluation resembles the failure it needs to solve. A broad claim that one model is “better” tells you less than a specific account of where the improvement occurs.

Sol and Astra: seven separate benchmark comparisons
Seven distinct benchmarks, using Sol and Astra values from the same Astra release table. Do not average them into an intelligence score. Source S11.
Benchmark gain breakdown
Percentage-point gains vary across seven distinct benchmarks. Separate tests must not be averaged into a general intelligence score. Source: S11.

The important test: how much correction remains?

An attractive first draft can still require substantial repair. A useful evaluation follows the work through acceptance.

OpenAI's Legora case study reports nearly 40% improvement on a particular financial-statement-review workflow, while the average improvement across Legora's broader Benchmark for Agentic Reasoning was about 3%. Those are vendor-published, workflow-specific results, rather than a universal gain for legal work. [S16]

In another vendor-published case, Playco reports 50% fewer manual fixes while prototyping games with Astra. That describes its integration and experience; it does not establish a general productivity effect across developers. [S17]

Speed also needs context. OpenAI reports OSWorld 2.0 latency simulations of roughly 75 minutes per task for Sol and 40 for Astra, with scores of 65.7% and 72.6%. The simulated time reduction is about 47%; a reader's workflow may behave differently. [S11]

These examples suggest a better testing question: How much usable work remains after review, correction, and recovery have been counted?

METR's survey of 349 technical workers illustrates another measurement distinction. Median self-reported speed change was around 3×, while self-reported work-value changes were around 1.4–2×. Selection and self-report limitations prevent treating those estimates as experimentally established gains for the workforce. [S21]

A faster low-value output and a better consequential output are different achievements. Decide which one your organization is trying to produce before choosing a success measure.

OSWorld 2.0 simulated task time
Published latency simulations, not measured speedups for the reader’s workflow. Source S11.

More context does not settle the cost question

The API specifications add useful perspective. The following are listed specifications and standard text-token prices checked September 11, 2026, rather than measured project costs. [S02, S12]

SpecificationGPT-5GPT-6 Astra
Context window400,000 tokens1,050,000 tokens
Maximum output128,000 tokens128,000 tokens
Listed knowledge cutoffSeptember 30, 2024April 30, 2026
Standard input price per million tokens$1.25$10
Standard output price per million tokens$10$50

Both model pages list text and image input, with native text output. Tool-enabled applications can add other functions around the model. [S02, S12]

Astra's context window is 2.625 times larger, while the stated output limit is unchanged. More room for source material does not establish permanent memory or perfect retrieval.

The displayed input and output unit prices are respectively eight and five times higher. Astra also has separate cache and processing rates; prompts exceeding 272,000 input tokens trigger higher rates for the full request. These are API charges, not ChatGPT subscription prices. [S02, S12]

The relevant business calculation is cost per accepted result: include failed attempts, tool charges, review time, and rework. That cost is UNKNOWN for the reader's assignment until measured.

Before committing to a paid workflow or scaling usage, run a small proof of concept with permitted data. No organization-specific implementation or return on investment was tested for this article.

GPT-5 and Astra API specification comparison
GPT-5 and GPT-6 Astra API specifications. Listed September 9, 2026. Unit prices are not measured project costs. Sources: S02, S12.

What another year could bring

A forecast becomes misleading when its mathematics are unclear.

Accuracy is bounded by 100%. The size of a task that a system can handle at a particular success probability is a different measure. Neither supplies a general growth rate for “AI.”

The following calculations preserve two separate models. They are illustrative scenarios, without assigned probabilities or statistical confidence intervals. They should not be read as predictions of what a September 2027 product will deliver.

Scenario model 1: reduce the remaining errors again

Using the reported GPQA endpoints, the remaining-error multiplier is:

4.0 ÷ 14.3 ≈ 0.2797. [S05, S11]

Apply that multiplier to Astra's 4.0% remaining error. Repeating the same proportional reduction for another year gives approximately 1.12% error, equivalent to 98.88% accuracy.

Applying the modeled improvement process at twice or three times the rate produces:

September 2027 arithmetic scenarioImplied accuracyRemaining error
Repeat the same error-reduction rate98.9%1.12%
Twice that modeled rate99.7%0.31%
Three times that modeled rate99.9%0.09%

The formula is 100 − 4 × (4/14.3)^a, where a equals 1, 2, or 3.

This avoids adding another 10.3 points to 96% and reaching an impossible 106.3%. It does not make the extrapolation reliable. Benchmark noise, saturation, changing evaluation conditions, and the difficulty of the remaining questions remain unresolved.

Treat the decimals as calculation outputs, not evidence that future performance can be predicted with that precision.

Remaining-error scenarios
Reported endpoint errors and mathematical remaining-error scenarios. Future values are illustrative, not calibrated forecasts. Sources S05 and S11; original calculations.

Scenario model 2: expand the size of the assignment

METR's task horizon relates task difficulty to the time a human expert would need. A 50%-success horizon marks a duration where the fitted model predicts 50% task success. It does not mean the AI runs independently for that many hours. [S18]

METR's January 2026 update reported a long-run hybrid doubling time of approximately 196 days. Fits since 2023 and 2024 were faster, around 131 and 89 days. That sensitivity matters: the selected time window and task suite affect the estimate. [S19]

For an illustration, normalize the September 2026 starting level to 1.0 and apply the rounded 196-day historical reference for 365 days:

Task-horizon multiple = 2^(a × 365/196).

Illustrative rate assumptionAssumed doubling timeSeptember 2027 multiple
Half the historical reference rate392 days1.9×
Historical reference rate196 days3.6×
Twice the reference rate98 days13.2×
Three times the reference rateAbout 65 days48.1×

These are original calculations using the historical reference, not a new estimate of the September 2025–September 2026 trend. The starting value is a normalization, not a measured Astra horizon. [S19]

The large endpoints come from compounding. Doubling the rate is a much stronger assumption than doubling the final capability level.

A 48.1× hypothetical task horizon does not mean 48.1× employee productivity, revenue, intelligence, or job displacement.

METR warns about transferring technical-task results to broader work and about uncertainty in its measurements. Its reviewed dashboard also flags estimates above 16 hours as unreliable with the current suite. These limitations become more consequential as an extrapolation reaches further beyond observed tasks. [S18, S20]

Illustrative task-horizon growth paths
ESTIMATED scenarios, all normalized to 1.0. The 196-day reference is historical, not an Astra-specific measurement. Rate multipliers are assumptions, without assigned probabilities. Source S19; original calculations.

Three possibilities worth preparing to test

The mathematics cannot tell us which profession will change first or which application will work. The following are INFERRED possibilities, conditional on improvements transferring to practical environments.

At the historical reference rate, investigate larger bounded assignments. A system might complete more of a software feature, a source-grounded research update, or a reconciled analysis before requiring intervention. The test is whether a larger package reaches the same quality threshold with less correction. Approval should remain explicit for consequential decisions and changes.

At twice the reference rate, investigate coordination across workstreams. Several specialized processes might gather evidence, reconcile records, build an output, and check dependencies under one plan. The difficult question is whether handoff failures fall enough to make that coordination dependable.

At three times the reference rate, investigate larger research and engineering programs. Systems might help formulate experiments, build evaluation tools, inspect results, and assemble reproducible evidence across more substantial projects. This is a conditional possibility, not proof of autonomous scientific judgment, recursive self-improvement, or a date for artificial general intelligence.

Prepare by making the work more testable: define inputs, acceptance criteria, permissions, and intervention points. Those foundations remain useful even if progress slows.

Three conditional capability pathways
Conditional possibilities for September 2027. These are INFERRED interpretations of ESTIMATED scenarios, not announced capabilities or probability-weighted forecasts.

Reliability changes what deserves delegation

Consider a hypothetical assignment with 20 necessary steps. Assume independent errors, the same success probability at each step, and no recovery.

At 95% reliability per step, all 20 succeed only about 35.8% of the time.

At 99%, the result becomes 81.8%.

At 99.9%, it reaches 98.0%.

These are calculations from p^20, not measured results for any AI model. Real errors can be correlated, and checking, retries, or recovery can alter the outcome.

The illustration shows why a system can look capable in individual interactions and still fail to finish a longer assignment. Evaluate the complete process, including what happens when a step goes wrong.

OpenAI's developer guidance describes asynchronous tool calling and mid-turn steering, while making clear that the application executes tools and manages pending work. Documentation establishes an available design capability; a specific integration still needs testing. [S13]

Safety also deserves separate attention. Astra's system card reports an estimated 8.5% indirect prompt-injection attack success rate versus Sol's 27.0% in a Gray Swan evaluation using up to 15 attempts per scenario. These are adversarial evaluation results, not production incident rates. The same system card reports reduced reasoning monitorability, with findings largely based on adversarial tests. [S14]

A responsible delegation decision therefore asks two questions together: can the system do the work, and will the surrounding process keep its actions within the authority granted?

20-step workflow reliability
Calculated illustration: 20 necessary steps, independent equal-probability errors, no recovery. These are not measurements of model reliability.

What this means for your career

The career implication I draw is a shift toward demonstrable judgment across a larger piece of work.

Consider an illustrative finance assignment. Drafting commentary is one contribution. Defining the reconciliation rules, identifying unsupported explanations, verifying the figures, and deciding what management needs to know is a broader contribution.

Consider an illustrative operations assignment. Generating a process document is useful. Establishing how exceptions are handled, testing whether the process works, and explaining why an apparent improvement should be rejected demonstrates a different level of ownership.

These examples are proposed career-development directions, not verified vacancies, implemented workflows, or predictions about hiring.

Ask what you can demonstrate across four activities: framing the problem, coordinating the work, evaluating the result, and explaining its consequences. Then identify the additional skill needed to strengthen the weakest part.

Keep evidence of your actual contribution. Record the problem, the tools used, the checks performed, the failures encountered, and the measured outcome. A résumé claim about responsible implementation should rest on work you actually completed.

For leaders, the parallel task is to make responsibility explicit. An assignment needs a clear definition of success, permissible actions, and a person accountable for acceptance. As the potential scope of delegated work grows, those decisions deserve greater attention.

A practical response: one assignment, one measured pilot

Here is a proposed evaluation approach. It remains UNTESTED for any particular organization in this article.

Choose one recurring, bounded assignment that a qualified person can evaluate. Establish the current process and its quality standard before comparing an AI-assisted alternative.

Use public, synthetic, or explicitly authorized information. Limit access to what the assignment needs. Set approval requirements before any external communication, financial commitment, or production change.

Measure the full process: preparation, execution, waiting, review, corrections, failures, and actual charges. Record whether the work met the acceptance criteria and remained within its permissions. Include unsuccessful attempts in the accounting.

Then make a decision supported by the result. Expand the pilot, narrow it, redesign it, or stop. A well-documented failure can prevent a much larger and more expensive mistake.

The next useful investment might be training. It might be cleaner source data, better task design, or a more appropriate integration. Let the test identify the gap.

The next year should be judged by completed work

A year is long enough to justify reopening an old assumption. It is also long enough for an impressive demonstration to become an outdated basis for a major decision.

The opportunity is to test larger and more useful assignments without abandoning the standards that make work trustworthy.

For professionals, that means developing evidence of useful contribution. For leaders, it means designing delegation around acceptance and accountability.

The question to carry into September 2027 is concrete:

What can the system finish correctly, how much intervention does it require, what does the accepted result cost, and can we verify what it did?

That is how a different class of AI should earn a different level of responsibility.

Turn the analysis into a career question

Identify one part of your work where stronger AI might change what you can contribute. MyTopMatch's Professional Passport is described as a place to organize skills, achievements, qualifications, goals, preferences, and dealbreakers. Start with the evidence you already have, then choose what to develop next. [S22]

Explore the Professional Passport → https://mytopmatch.com/professional-passport

Evidence note: Published figures and documentation were checked against the cited sources. Model evaluations were not independently rerun. Scenario values are illustrative estimates without probabilities; workplace interpretations are inferences. Customer examples are attributed vendor-published reports, not MyTopMatch platform outcomes. Costs for readers' workflows remain unknown until measured. No job, earnings, or productivity outcome is guaranteed.