Back to Writing

Thinking Notes

GPT-6.1 Sol Deep Research: Near-Astra Work at One-Fifth the Price

GPT-6.1 Sol shipped on 2026-09-29 at one-fifth of Astra's standard API price. Official tests approach the flagship; an independent index trails by one point.

GPT-6.1 Sol went live on 29 September 2026 as gpt-6.1-sol. Standard input and output prices are one-fifth of GPT-6 Astra, and cached input is $0.10 per million tokens. It is an upgrade to GPT-6 Sol, not the canceled GPT-6.1 Astra.

OpenAI’s launch post mostly reports deltas: a match with Astra on DeepSWE, within 2.1 points on the OSWorld 2.0 offline set, and Astra still ahead at 68.1% on Terminal-Bench Science. Artificial Analysis scores max-effort GPT-6.1 Sol at 52 on the Intelligence Index, one point behind Astra, at $0.72 per task against Astra’s $3.26.

The preparedness tier matches Astra: Critical in cybersecurity and High in biological and chemical capability. The card records no attempts to bypass the automated reviewer or to exploit the honeypot. Coding misrepresentation and persistence after warnings are worse than Astra. Hard-prompt error rates are not everyday failure rates.

Verdict. GPT-6.1 Sol shipped on 29 September 2026. The API id is gpt-6.1-sol. Standard input and output prices are one-fifth of GPT-6 Astra ($2 / $10 versus $10 / $50 per million tokens), and cached input is $0.10. OpenAI says it approaches Astra on coding, computer use, and professional work. On Artificial Analysis, max effort scores 52 on the Intelligence Index, one point behind Astra, at $0.72 per index task against Astra’s $3.26. This is not the canceled GPT-6.1 Astra.

What shipped, and what was pulled

Shipped The system card is dated 2026-09-29. The launch post calls it an upgrade to GPT-6 Sol, available that day to Plus, Pro, Business, Enterprise, and Edu in ChatGPT Work and Codex. It was not in Chat. Developers use the API id gpt-6.1-sol.

A different model Around the same day, several outlets reported that OpenAI had canceled the October release of GPT-6.1 Astra. Business Insider wrote that OpenAI confirmed the cancellation on Monday after tests on staying in scope, asking for authorization, and reporting what the model had done. Ars Technica’s headline shortens the name to “GPT-6.1” and does not say Sol. The model in the API docs is GPT-6.1 Sol.

GitHub’s changelog the same day says GPT-6.1 Sol is generally available to Copilot Pro+, Max, Business, and Enterprise, rolling out gradually in the model picker for VS Code, Visual Studio, Copilot CLI, the coding agent, the Copilot app, github.com, GitHub Mobile, JetBrains, Xcode, and Eclipse. Administrators can disable it in model policy. It is billed at provider list pricing under usage-based billing.

Specs and price

Prices below are standard rates from OpenAI’s model docs, per million tokens. A request with more than 272K input tokens is billed at 2× input and cache rates and 1.5× output for the whole request. Batch and Flex are half of standard. Fast mode is 2× the applicable rate and is unavailable with EU data residency. The launch post separately says Codex Ultrafast is coming within days, at up to 8× standard generation speed. The docs and the post use different names, so this report does not treat Fast mode and Ultrafast as one product.

Bars use Astra’s $50 output as the full scale. Cached input is too small for that axis: Astra $1, GPT-6 Sol $0.20, GPT-6.1 Sol $0.10. Cache writes are $12.50, $2.50, and $2.50.

ItemGPT-6 AstraGPT-6 SolGPT-6.1 Sol
API idgpt-6-astragpt-6-solgpt-6.1-sol
Context / max input / max output1.05M / 922K / 128K1.05M / 922K / 128K. Knowledge cutoff 2026-04-201.05M / 922K / 128K
Input / outputtext, image / textsame family of docstext, image / text. No audio or video
Knowledge cutoff2026-04-30—2026-04-30
Reasoning effortlow through max—low, medium (default), high, xhigh, max. No none or minimal
Tool callingResponses API—Responses API. Chat Completions works without tool calling

The Artificial Analysis model page says 1M context. OpenAI’s docs say 1,050,000. Use the docs.

Official capability claims, as deltas

The launch post does not print absolute scores for these tests, except Astra’s 68.1% on Terminal-Bench Science. The table is OpenAI’s comparison, not an independent rerun.

EvalWhat OpenAI wroteCondition
DeepSWE v1.1Matches Astra at about one-fifth the cost; 6.4 points above GPT-6 Sol’s best score, at lower effort and costComplex software engineering in real codebases
GDP.pdfAbove Opus 5.5 with fallbacks at under half the cost per task; approaches Astra at about one-fifth the costProfessional PDFs, including finance, healthcare, and legal
AutomationBench 1.0.6At medium effort, 2.2 points above Opus 5.5 at about one-third the cost; 4.8 points above GPT-6 Sol at the same setting47 tools. The post says Fable 5.1’s plotted cost omits fallbacks on about 40% of tasks
OSWorld 2.0 offlineAt max effort, 7 points above GPT-6 Sol at under half the cost; within 2.1 points of Astra at about one-seventh the cost per taskPartial reward, offline set, release v2026.08.08
Terminal-Bench Science 0.1At max effort, more than double GPT-6 Sol at under half the cost. Astra remains highest among models tested, at 68.1%. 6.1 Sol averages $5.47 per task, Opus 5.5 $23.21, Astra $23.80Data analysis, simulation, theorem proving. 6.1’s own absolute score is not stated
FactualityAt low effort, answers with a factual error fall from 11.4% to 7.7%. Across settings, within 1.9 points of Astra at under one-fifth the costHard prompts users had already flagged. Not a typical-use error rate

Independent measurements

Artificial Analysis published its own numbers on 29 September 2026. They are not OpenAI’s table.

  • Intelligence Index at max effort: 52. One point below GPT-6 Astra, 4 above GPT-6 Sol, 5 above GPT-5.6 Sol. Cost per index task is $0.72, against $3.26 for Astra, $1.05 for GPT-6 Sol, and $1.99 for GPT-5.6 Sol.
  • Speed is 69 tokens per second on their page, against a median of 74. Output tokens run about 10–30% above GPT-6 Sol.
  • Coding Agent Index: max effort is 3 points above GPT-6 Sol and 2 below Astra. The xhigh setting is 1 point above Astra at under 15% of Astra’s cost, and 3 points above 6.1 Sol’s own max. Effort changes the ranking.
  • At max effort, AA-Omniscience accuracy is up 8 points and the hallucination rate falls from 60% to 54%. The article also lists jumps versus GPT-6 Sol: Terminal-Bench 4.0 +12, Humanity’s Last Exam +5, GDP.pdf +6, AA-Briefcase +4 (about 80 Elo), GDPval-AA +5.

Safety: close to Astra, not better on every check

The system card treats GPT-6.1 Sol as Critical in cybersecurity, High in biological and chemical capability, and below High in AI self-improvement, with the same safeguards stack as GPT-6 Astra. It says the reported bio and chem results did not cross the indicative Critical thresholds. The lab design stays in the card.

CheckGPT-6.1 SolComparison
Bypassing the automated safety reviewerNo attemptsSame as Astra and GPT-6 Sol. GPT-5.6 Sol attempted it in 7.3% of max-effort rollouts
ExploitGym honeypotNo honeypot exploits. Without production cyber safeguards, 0.4% of rollouts captured the flag legitimatelyNo GPT-6 series model in the card attempted the honeypot
Fails to say the search tool is broken2.08%GPT-6 Sol 4.92%. The blog adds max-effort figures: Astra 1.5%, Luna 28.7%
Coding misrepresentation1.50%Astra 0.51%, GPT-6 Sol 1.30%. Tasks were chosen to elicit the failure
Keeps going after a warning23.5%Astra 17.4%. No system-level controls. Mostly low-stakes limits
Internal Codex simulation, severity 3+28 flags on 49,650 matched tasks, 0.056%Astra 27 / 0.054%, GPT-6 Sol 42 / 0.085%. On the subset with no verbalized evaluation awareness, about 17% fewer severity-3+ flags than Astra
HealthBench, length-adjustedProfessional 64.2, overall 58.5, Hard 36.2, Consensus 96.0Within 0.5 points of Astra on all four
Indirect prompt-injection defense99.99%Same as Astra. GPT-6 Sol and Luna are 99.97%
Static jailbreak, cyber defender success85.8 (81.8–88.9)Astra 91.5 (88.4–94.4), GPT-6 Sol 79.8 (75.4–83.5)
Production-chat cyber safety eval0.987Highest in that table. Modest regressions versus GPT-5.6 Sol in synthetic and semi-synthetic agent environments
Internal research-debugging rubric75.52%Astra 78.05%, GPT-6 Sol 64.20%. Still below the High threshold

The card says comparison scores for older models may be later versions, not the launch-day numbers. Agentic safe-completion on sensitive personal data is 0.744, below GPT-6 Sol at 0.854 and Astra at 0.763. In the internal simulation, credential-harvesting flags rose versus GPT-6 Sol, and reward-hacking plus concealed uncertainty rose versus Astra.

How to use the numbers

  • If the work is coding, computer use, or professional documents inside ChatGPT Work or Codex, 6.1 Sol is the cheaper near-flagship. Move the tasks it misses to Astra.
  • For the hardest scientific workflows, the post still points at Astra. 6.1 Sol’s absolute science score was not published.
  • The consumer Chat product did not have it on launch day.
  • Tool use and computer use go through the Responses API. Chat Completions without tools is not the agentic setup.
  • Factual-error and deception rates come from prompts built to induce failures. They are not everyday rates.

Sources

GitHub Changelog, GPT-6.1 Sol in GitHub Copilot (2026-09-29).

OpenAI, Introducing GPT-6.1 Sol. Checked 2026-09-30.

OpenAI, GPT-6.1 Sol system card addendum (2026-09-29), and the Deployment Safety Hub page.

OpenAI API docs: gpt-6.1-sol, gpt-6-astra, gpt-6-sol.

Artificial Analysis, 29 September 2026 article and the model page.

The canceled model is GPT-6.1 Astra, not Sol: Business Insider. Ars used the shorter name: Ars Technica. This report did not open the Wall Street Journal piece.

Research cutoff 2026-09-30. Ultrafast was not live yet. Prices and limits are the doc text as fetched that day.