🔍 Read the full analysis: My AI Workflow For September 2026: Build, Dig, Decide on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Analyst Thorsten Meyer published his September 2026 AI workflow, one day after GPT-6.1 Sol’s release: Claude Opus 5.5 builds, Sol reviews and digs into details at roughly one-eighth the per-task cost of rivals. Six frontier models now sit within about 20 index points while per-task costs differ by about 100×, shifting the selection question from capability to cost per task.
Independent AI analyst Thorsten Meyer published his model workflow for September 2026 on 29 September 2026, recommending Claude Opus 5.5 as the primary model for software development and the same-day-released GPT-6.1 Sol as a low-cost second model for detailed review. The recommendation rests on his reading of Artificial Analysis benchmark data showing six frontier models clustered within roughly 20 index points of each other while their cost per task differs by about 100× — a spread he says has changed the central question in model selection from “which model is smartest?” to “which model clears my quality bar at the lowest cost per task?”
Meyer’s stack assigns specific roles by cost and measured capability. Opus 5.5 (released 22 September, index score 58 at max) serves as the main builder at high or xhigh effort, costing $1.82 to $3.46 per task. GPT-6.1 Sol at xhigh scores 51 and costs $0.39 per task, making a review pass on every meaningful change affordable, according to Meyer. Sonnet 5.5 and GPT-6 Luna handle side work, while Astra and Fable are reserved as tie-breakers when the two main models disagree. All scores cited come from the Artificial Analysis Intelligence Index v4.3.x.
Three price-performance findings anchor the piece. First, Opus 5.5 outscores its more expensive sibling Claude Fable 5.1 (58 vs. 53) while costing less per task ($5.98 vs. $7.63 at max). Second, Sonnet 5.5 at max effort costs $7.60 per task — more than Opus at max — for two fewer points, which Meyer argues disqualifies it at that setting. Third, GPT-6.1 Sol costs roughly one-eighth of Astra and one-twentieth of Fable per task while scoring only one to two points lower.
Meyer also reports that the effort-level setting, not model choice, is the largest cost lever: on Opus 5.5, moving from xhigh to max adds two index points but raises cost per task by 73%; from medium to max, cost rises 4.46× for seven points. He runs Opus at high (54 points, $1.82) for everyday development and xhigh only for architecture, migrations and trust-boundary work.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
Why Cost Per Task Now Drives Model Choice
The piece documents a practical shift for anyone spending on AI infrastructure: as top-tier scores compress, pricing has become the main differentiator. A developer who routes review and classification work to cheap models like Sol or Luna — $0.39 and $0.07 per task respectively — can run those steps on every change rather than selectively. Meyer notes that Luna delivers 1,429 tasks per $100, against 17 for Opus at max, making high-volume routing and classification jobs dramatically cheaper.
The workflow also carries a governance argument: a different model family reviewing output is a stronger check than a model reviewing itself, and low cost makes independent review routine rather than exceptional. Meyer pairs this with four operating rules, including that effort settings do not add capability and that passing tests is not approval to ship. He also cautions that halving model price saves only about 12.5% of real project cost — an illustrative figure, not a measured one — because a single extra minute of human review can erase the saving.
A Month of Back-to-Back Frontier Releases
: “September 2026 saw near-weekly frontier releases, according to the piece’s timeline: Claude Fable 5.1 on 1 September, GPT-6 Astra on 3 September, Opus 5.5 and GPT-6 Luna on 22 September, Sonnet 5.5 on 28 September, and GPT-6.1 Sol on 29 September — the day the workflow was published. Sol launched at the same published token price as its week-old predecessor, $2 input / $10 output per million tokens.
Despite the fast cadence, Meyer observes that capability gains have narrowed: one index point, he notes, is inside the noise. Artificial Analysis has not yet published low or max effort settings for Sol, listing only medium, high and xhigh as of the publication date. Meyer also flags latency trade-offs: Sol takes 57 to 69 seconds to produce a first token at high and xhigh settings, and Opus still leads it by five points at xhigh (56 vs. 51).
“In four weeks, the AI frontier stopped being a leaderboard and became a price curve.”
— Thorsten Meyer, ThorstenMeyerAI.com
Benchmark Limits and Missing Settings
Several points remain unresolved. Artificial Analysis has not published low or max settings for GPT-6.1 Sol as of 29 September, so its full cost-performance range is unknown. Meyer himself notes that one index point is within measurement noise, and that the index measures general capability rather than performance on any specific workload — he advises shadow-testing before switching models. His claim that cheaper tokens save only 12.5% of real cost is explicitly labeled illustrative, not measured. Sol’s 57-to-69-second time to first token at high and xhigh also means it is unsuitable for interactive use at those settings. This is one analyst’s workflow based on third-party benchmarks, not an independently verified evaluation.
October Benchmarks and Sol’s Full Range
The next data points to watch are Artificial Analysis’s pending low and max effort settings for GPT-6.1 Sol, which will complete its price-performance picture. Meyer’s workflow implies continued divergence between premium builders and cheap reviewers, so upcoming releases will likely be judged on cost per task at a given quality bar rather than headline scores. Readers applying the approach are advised by the author to shadow-test candidate models against their own workloads before switching, and to expect effort-level pricing to remain the dominant cost lever in the near term.
Key Questions
What is the core recommendation in this September 2026 workflow?
Use Claude Opus 5.5 at high or xhigh effort as the main building model, and GPT-6.1 Sol at high or xhigh as a cheap reviewer and detail-digger at $0.32–$0.39 per task, with Sonnet 5.5, Luna, Astra and Fable reserved for specific roles rather than as defaults.
Why is GPT-6.1 Sol positioned as a review model rather than a builder?
Per Meyer, Sol scores 51 at xhigh — five points below Opus 5.5 — and takes 57 to 69 seconds to produce a first token at high and xhigh, making it non-interactive. But at roughly one-eighth of Astra’s and one-twentieth of Fable’s per-task cost, it is affordable to run as an independent review pass on every meaningful change.
Where do the benchmark scores and prices come from?
All scores come from the Artificial Analysis Intelligence Index v4.3.x, and token prices are the vendors’ published rates (e.g., Opus 5.5 at $4/$20 and Sol at $2/$10 per million input/output tokens). Meyer cautions the index is a general-capability map, not a workload verdict.
Is Sonnet 5.5 still worth using at its maximum effort setting?
According to Meyer, no: at max it costs $7.60 per task — more than Opus 5.5 at max — for two fewer index points, and it produces about 193k output tokens per task, the most Artificial Analysis has measured. He places its best value at high effort: 47 points for $1.08.
Does switching to cheaper models actually reduce overall project costs?
Only modestly, per Meyer’s illustrative example: halving model price saves about 12.5% of real cost, and one extra minute of human review can erase that saving. He presents this as an illustration rather than a measured result.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
