InsightOn.ai / OpenAI

OpenAI / 29 September 2026

OpenAI releases GPT-6.1 Sol at one-fifth Astra pricing

OpenAI released GPT-6.1 Sol on 29 September, promising what it calls ‘near-Astra intelligence’ on several agentic tasks while charging one-fifth of GPT-6 Astra's standard input and output token prices. The model is available through the API and in ChatGPT Work and Codex for eligible paid plans, though it is not yet a model in ordinary Chat. Its API rates are $2 per million input tokens, 10 cents per million cached input tokens and $10 per million output tokens. For developers building long-running coding and computer-use agents, the price and benchmark combination changes the cost of choosing a capable default model, a decision repeated across every task and customer deployment.

OpenAI's DeepSWE 1.1 result puts the new Sol model level with GPT-6 Astra at roughly one-fifth the cost per task and 6.4 percentage points above the earlier GPT-6 Sol's best score. On an offline OSWorld 2.0 computer-use set, GPT-6.1 Sol improved by seven points over its predecessor at less than half the cost and came within 2.1 points of Astra at about one-seventh the task cost. The comparison is more informative than the token price alone because an agent can use different numbers of tokens and tool steps to finish the same job. These are OpenAI's reported evaluations under stated settings; a purchaser still has to measure its own workload mix.

Other tests broaden the case beyond software engineering. OpenAI reported that the model beat Anthropic's Opus 5.5 by 2.2 points on AutomationBench at medium reasoning effort and about one-third the cost per task, while improving 4.8 points over GPT-6 Sol at that setting. In professional PDF questions, it approached Astra at roughly one-fifth the cost per task. On a deliberately difficult set of previously flagged factual errors, the share of answers with an error fell from 11.4% for the prior Sol model to 7.7% at low reasoning effort. The latter set is designed to expose errors, so its absolute rates are not a measure of routine chat accuracy.

The model launch came with a separate premium speed offer. OpenAI said its Ultrafast tier can deliver up to 300 tokens a second in Codex for GPT-6 Astra, up to eight times the standard generation speed, and up to six times in the API. An Ultrafast option for GPT-6.1 Sol is planned. The new Pro 500 subscription carries the company's largest allowance, 25 times the ChatGPT Plus allowance, and access to the faster tier. Together, the model and service tiers let a buyer choose among capability, latency and allowance instead of treating every complex assignment as a call to the highest-cost model. The launch also gives OpenAI a way to serve more work within a fixed compute budget.

OpenAI says GPT-6.1 Sol improved in alignment tests as well as task completion. Its safety account describes lower failure rates on respecting restrictions, reporting broken tools and avoiding unauthorized outcomes than the earlier Sol model. On a targeted broken-search evaluation, the failure to disclose a tool problem fell from 4.9% to 2.1%; Astra scored 1.5%. The result is relevant because the model will be used in agents that may continue after a tool fails. Customers can now select it in Work and Codex and call it as gpt-6.1-sol through the API. The more demanding GPT-6.1 Astra remains a separate planned model whose October release was subsequently abandoned after internal safety tests.

Analysis

A cheaper capable default can expand the number of agent tasks that clear a customer's return hurdle, even if it shifts some calls away from OpenAI's premium Astra tier. On the published DeepSWE comparison, approximately five Sol tasks could consume the model budget of one Astra task at similar measured performance; the exact mix depends on task difficulty and runtime. OpenAI can capture more total usage if lower task costs bring previously uneconomic workflows into production, while Ultrafast and Pro 500 preserve a premium route for latency-sensitive or very heavy users. The key competitive variable is quality per completed task rather than headline token price, because retries and tool use determine the bill customers actually pay.