GPT-6.1 Sol had reached 503 Hacker News points and 403 comments when MrKeyoor's feed captured it on September 29. Yet the release's most useful detail for developers is easy to miss: OpenAI did not cut the normal Sol input or output rates. The new model still costs $2 per million input tokens and $10 per million output tokens. The price change is cached input, which falls from $0.20 on GPT-6 Sol to $0.10 on GPT-6.1 Sol.
That distinction decides who gets a meaningful saving. An agent that repeatedly sends the same large prefix, such as repository instructions, tool definitions, or a long project brief, can benefit from the lower cache-read rate. A one-off request with a fresh prompt cannot. OpenAI's launch announcement leads with near-Astra performance at one-fifth of Astra's standard token prices. Compared with the model GPT-6.1 Sol directly replaces, the economic case rests on better results at familiar rates and cheaper reuse of context.
The standard rates did not move
The price table needs two comparisons. Against GPT-6 Astra, GPT-6.1 Sol is straightforward: Astra costs $10 per million input tokens and $50 per million output tokens, so Sol's $2 and $10 rates are exactly one-fifth as high. Against GPT-6 Sol, the standard rates are unchanged. Both generations charge $2 for input, $2.50 for cache writes, and $10 for output per million tokens.
Cache reads are the exception. The GPT-6.1 Sol model page lists cached input at five percent of its normal input rate. The GPT-6 Sol model page lists ten percent. OpenAI therefore cut the cost of reading a reusable prompt prefix in half while leaving the first write at $2.50 per million tokens. A team has to reuse that prefix enough times before the cheaper reads outweigh the write and the work needed to keep the prefix stable.
Consider a request that reads 200,000 cached input tokens and produces 10,000 output tokens. Ignoring tool fees, GPT-6.1 Sol charges about $0.12: two cents for the cached prompt and ten cents for the answer. GPT-6 Sol comes to about $0.14. Astra comes to roughly $0.70 at its listed cache and output rates. The new Sol saves two cents against old Sol on that request, not 58 cents. The larger difference belongs to the existing Sol-versus-Astra tiering.
The calculation changes again for very long prompts. OpenAI says requests above 272,000 input tokens cost twice the listed input and cache rates, while output rises to 1.5 times the listed rate for the whole request. Batch and Flex processing are priced at half of Standard, Fast mode at twice Standard, and eligible regional processing adds ten percent. Any cost model built from the headline $0.10 figure will be wrong if it ignores those conditions.
Better work at the same rate is the larger claim
OpenAI's argument for upgrading is mostly about how much useful work one request completes. On DeepSWE v1.1, the company says GPT-6.1 Sol matched Astra at about one-fifth of the cost and beat the best GPT-6 Sol result by 6.4 percentage points while using a lower reasoning setting. DeepSWE uses original, long-running software tasks in real codebases, which makes the comparison more relevant to coding agents than a short answer test. It still does not predict performance on a particular repository.
The pattern continues in other tests reported in the announcement. On AutomationBench, GPT-6.1 Sol at medium reasoning finished 4.8 percentage points above GPT-6 Sol at the same setting. On the offline portion of OSWorld 2.0, which measures computer use, the new model gained seven points over GPT-6 Sol at maximum reasoning and came within 2.1 points of Astra. Those are different tasks with different cost profiles, so the gaps should not be blended into one general quality score.
The scientific-work result shows why cost per completed task can be more useful than token price alone. At maximum reasoning on Terminal-Bench Science 0.1, OpenAI reports an average task cost of $5.47 for GPT-6.1 Sol, compared with $23.80 for Astra. Astra still posted the highest score, 68.1 percent. A team choosing between them is buying a probability of completion, not merely tokens. Sol can be cheaper per attempt while Astra remains worth its price on work where a failed run costs more than the model call.
OpenAI also reports a factuality gain at low reasoning. On a set of de-identified conversations where users had flagged an earlier model error, the share of GPT-6.1 Sol answers containing at least one factual error fell from 11.4 percent to 7.7 percent. The company explicitly warns that these prompts were selected to provoke failures and do not represent normal traffic. That caveat belongs beside the percentage. It stops a difficult test set from being mistaken for a general hallucination rate.
All of these numbers come from OpenAI's research environment or API runs. The company notes that production results can differ because system prompts, available tools, and reasoning settings change. It also says competitor scores came from public reports. Developers should read the charts as a reason to rerun their own evaluation set, not as a substitute for one.
Migration has a few sharp edges
The API model name is gpt-6.1-sol. The official GPT-6 guide says it supports low, medium, high, xhigh, and max reasoning, with medium as the default. Unlike GPT-6 Sol, the new model does not accept none or minimal. An application using the old model with reasoning disabled cannot change only the model string and assume identical behavior, latency, or cost.
Tool use is another dividing line. OpenAI directs developers to the Responses API for GPT-6.1 Sol tool calling. Chat Completions remains available for requests without tools. The model accepts text and image input, produces text, and has a 1.05 million-token context window with up to 128,000 output tokens. Audio and video are not supported inputs on the model page, and fine-tuning is not offered.
For an agent service, a sensible evaluation keeps the workload fixed and records more than pass rate. Teams should measure total input, the share served from cache, output tokens, tool calls, wall time, retries, and completed-task cost. A higher benchmark score can still produce a worse bill if the model reasons longer, emits more text, or repeats failed tool actions. A lower cache price helps only when the request structure earns cache hits.
Prompt organization now carries a visible price consequence. Stable instructions and tool schemas belong at the front of a request where they can form a reusable prefix. Volatile user data belongs later. That was already good cache hygiene. At $0.10 per million cached tokens, OpenAI is putting a clearer discount on it. The model page also lists cache writes at 1.25 times the uncached input rate, so constantly changing the prefix throws away much of the benefit.
Work and Codex get it before Chat
Availability is split across OpenAI's products. The launch post says GPT-6.1 Sol is available to Plus, Pro, Business, Enterprise, and Edu customers in ChatGPT Work and Codex, and to developers through the API. OpenAI says it is not yet available in Chat. That wording matters for readers expecting to select it in an ordinary chat model picker.
OpenAI also says an Ultrafast option is due in the coming days, with token generation up to eight times faster than standard speed in Codex. That is a forward-looking availability claim, not a shipping feature developers can benchmark today. The current model documentation says Fast mode costs twice the Standard rate and is unavailable with EU data residency. US and EU data residency are otherwise supported.
The immediate decision is narrow. Existing GPT-6 Sol users can test whether the 6.1 model completes more of their real coding, document, or computer-use jobs without raising standard token rates. Teams with large repeated prefixes should separately verify cache-hit behavior and include cache writes in the bill. Astra still has the top reported score on some demanding work, while Luna remains the cheaper tier for focused, high-volume tasks.
What comes next should be visible in production traces rather than another chart: cache-hit rates after migration, task completion at each reasoning setting, tail latency, and the number of jobs that still need Astra. If GPT-6.1 Sol changes developer spending, it will happen through fewer failed runs and more reused context. The unchanged $2 and $10 rates make those two measurements the ones worth watching.