GPT-5.6: compare the models by coding and agent tasks
Main chat
A chat for vibe coders: news, guides, live cases, marketplace, and finding executors.
Quick answer
A model comparison is useful only when it matches the work you actually do. For GPT-5.6, separate coding quality, agent reliability, latency, context handling and price instead of choosing from a single benchmark.
Run a small private evaluation: five representative tasks, the same instructions, a fixed acceptance checklist and a record of retries or manual fixes. That result will be more useful than a generic leaderboard for your project.
OpenAI opened GPT-5.6 in ChatGPT, Codex and API. We compare Sol, Terra and Luna to GPT-5.5 in terms of code, agent tasks, price and main benchmarks.
On July 9, OpenAI shared the GPT-5.6 family in ChatGPT, Codex and API. This time, the release consists of three models at once: the flagship Sol, the balanced Terra and the more affordable Luna.
At first glance, the choice looks linear: the more expensive the model, the smarter it is. OpenAI's final tables show a more interesting picture. Terra overtakes GPT-5.5 at half the price. The Luna fare is a fifth of the Sol fare, but on complex code and long context, it is sometimes inferior even to the previous generation. Sol breaks several records, but doesn’t win every test either.
Prior to the wide launch we are писали о GPT‑5.6 по данным ограниченного превью. Now OpenAI has published полную страницу релиза with results on code, browser, science, cybersecurity and long context. Below is a comparison based on these final data.
What did OpenAI release
The number in the title now stands for generation, and Sol, Terra and Luna are constant levels of capability. OpenAI plans to update them at its own pace.
- GPT-5.6 Sol is a flagship for complex code, professional analysis and long-term agency work. The short identifier
gpt-5.6in the API leads to Sol. - GPT-5.6 Terra is a mid-range model. In its role, it is closer to the previous mini-models, but in many tests it competes with the GPT-5.5.
- GPT-5.6 Luna is the cheapest option for bulk queries, classification, data extraction and short automation.
According to current model cards, all three versions have a context window of 1.05 million tokens and a maximum response of 128,000 tokens. The date of the knowledge slice is February 16, 2026. GPT-5.5 has the same window and maximum response, but the knowledge is limited on December 1, 2025.
| Модель | Вход за 1 млн токенов | Кэшированный вход | Выход за 1 млн токенов | Контекст |
|---|---|---|---|---|
| GPT‑5.5 | $5 | $0,50 | $30 | 1,05 млн |
| GPT‑5.6 Sol | $5 | $0,50 | $30 | 1,05 млн |
| GPT‑5.6 Terra | $2,50 | $0,25 | $15 | 1,05 млн |
| GPT‑5.6 Luna | $1 | $0,10 | $6 | 1,05 млн |
The nominal price of Sol coincides with GPT-5.5. Terra is half cheaper, Luna is five times cheaper. For requests longer than 272 thousand input tokens, there is an increased tariff for the entire request: entrance costs twice as much, output - 1.5 times. Recording a new cache is charged at a factor of 1.25, and reading from the cache saves a 90% discount.
The main benchmarks are GPT-5.6 against GPT-5. 5
The table below includes tests from different categories. This is not an attempt to reduce model quality to a single number, but a quick way to see the nature of the upgrade.
| Бенчмарк | Что проверяет | Sol | Terra | Luna | GPT‑5.5 |
|---|---|---|---|---|---|
| Agents’ Last Exam | Долгие профессиональные задачи | 52,7% | 50,4% | 50,3% | 46,9% |
| Coding Agent Index | Работа кодового агента | 80,0 | 77,4 | 74,6 | 76,4 |
| SWE‑Bench Pro | Исправление реальных репозиториев | 64,6% | 63,4% | 62,7% | 59,4% |
| Terminal‑Bench 2.1 | Сложная работа в терминале | 88,8% | 87,4% | 84,7% | 85,6% |
| BrowseComp | Поиск и исследование в интернете | 90,4% | 87,5% | 83,3% | 84,4% |
| OSWorld 2.0 | Управление приложениями и компьютером | 62,6% | 50,2% | 45,6% | 47,5% |
| HealthBench Professional | Профессиональные медицинские задачи | 60,5% | 57,7% | 55,7% | 49,5% |
| GeneBench Pro | Долгий анализ геномных данных | 28,7% | 23,3% | 10,8% | 12,0% |
| SEC‑Bench Pro | Рабочие примеры эксплуатации уязвимостей | 71,2% | 57,7% | 48,9% | 45,8% |
| MRCR 512K–1M | Поиск фактов в очень длинном контексте | 73,8% | 72,5% | 41,3% | 74,0% |
| Toolathlon | Использование разных инструментов | 58,0% | 53,1% | 53,4% | 55,6% |
Sol breaks off the most not in the usual issues, but where the model should work for a long time: control the computer, conduct research, work with vulnerabilities or collect results from many steps. On OSWorld 2.0, the increase against GPT-5.5 is 15.1 percentage points, on GeneBench Pro – 16.7, on SEC-Bench Pro – 25.4.
In the code, the growth is noticeable, but not so dramatic. Sol adds 5.2 points on the SWE-Bench Pro and 3.2 points on Terminal-Bench 2.1. The main change here is not only the maximum result, but also the density of the family: Terra is very close to Sol, and on the SWE-Bench Pro even Luna surpasses GPT-5.5.
Sol: Agential tasks have grown the most
GPT-5.6 Sol is the safest choice when a model error costs more than the tokens themselves. These are complex changes in the codebase, research with dozens of sources, analysis of documents, work with the browser and tasks where the agent must check his own result several times.
On the Artificial Analysis Coding Agent Index, the model scored 80 against 76.4 for GPT-5.5. On BrowseComp, the result rose from 84.4% to 90.4%, and on BenchCAD - from 44.4% to 70.6%. The last test is especially indicative: it checks the work with engineering drawings, where it is not enough to write a plausible text - you need to control the tool correctly and get a valid artifact.
OpenAI also claims higher efficiency. In early testing, Lovable reported about 25 percent fewer steps, a 35 to 48 percent reduction in tool calls, and a 15 percent decrease in stuck launches. This is the data of a single partner, not a universal guarantee, but it explains why the same price for a token does not mean the same cost of the finished task.
Terra looks like a major API update
Sol collects records, but Terra is more interesting for product teams. It costs exactly half the price of the GPT-5.5, while overtaking the previous model on most of the key tests in the table above.
On SWE-Bench Pro Terra receives 63.4% against 59.4% for GPT-5.5. Terminal-Bench 2.1 has 87.4% compared to 85.6%. On BrowseComp - 87.5% against 84.4%. For Agents’ Last Exam, the difference is 50.4% versus 46.9%.
This makes Terra a logical starting point for bots, content generation, document analysis, internal assistants, and background agents. Previously, such routes often had to be sent to the flagship model, because the mini-level too quickly lost quality. Now there are several points between Terra and Sol in many tests, and the difference in tariff is twofold.
There are exceptions. Terra is inferior to GPT-5.5 on Toolathlon, several academic tests and extracting information from context with a length of 512,000 to 1 million tokens. Therefore, it is too early to change the model in all routes with one line.
Luna is cheap, but not universal
Luna costs $1 per million input and $6 per million output tokens. It is suitable for tasks that are easy to check automatically: classification of requests, field extraction, routing, text normalization, short resumes and mass processing of the same type of records.
On SWE-Bench Pro Luna shows 62.7% – above GPT-5.5 from its 59.4%. On HealthBench Professional, she scores 55.7 percent versus 49.5 percent. On SEC-Bench Pro – 48.9% against 45.8%.
But complex agent scenarios quickly reveal the limits of savings. Luna trails GPT-5.5 on the Coding Agent Index, Terminal-Bench, BrowseComp and OSWorld. The gap is particularly noticeable in the long context: 41.3% on MRCR 512K-1M versus 74% on GPT-5.5. Formally, the model accepts 1.05 million tokens, but the size of the window does not guarantee that it will find and connect the necessary facts inside it equally well.
The practical conclusion is simple: the Luna is good as a standalone low-cost route, not a global replacement for the older model.
What Max and Ultra give
In GPT‐5.6, a new level of max reasoning appeared, standing above xhigh. It gives the model more time to search for options, check and correct the result. The basic levels of none, low, medium, high and xhigh were previously described in отдельном материале о reasoning effort. OpenAI does not recommend turning on max for each query: first compare the current level of reasoning with the same level on GPT-5.6, and then test adjacent settings on real-world problems.
Ultra goes further and runs multiple agents in parallel. The standard configuration uses four agents that divide the work into independent directions and then combine the results.
| Бенчмарк | Sol | Sol Ultra | Прирост |
|---|---|---|---|
| Terminal‑Bench 2.1 | 88,8% | 91,9% | +3,1 п. п. |
| BrowseComp | 90,4% | 92,2% | +1,8 п. п. |
| SEC‑Bench Pro | 71,2% | 74,3% | +3,1 п. п. |
Ultra spends more total tokens because the work of all agents is paid, but independent branches are executed simultaneously. Therefore, the mode is designed for tasks where the maximum probability of success and less waiting time are important, rather than the minimum bill for the request.
Codex Ultra is available starting with Plus. In ChatGPT Work, it is open to Pro and Enterprise users. Developers can collect similar scenarios through the multi-agent beta in the Responses API.
Where GPT-5.5 is still stronger
The final table is also useful because it does not show the perfect victory of the new generation.
- On MRCR with a context of 512 thousand – 1 million tokens, GPT-5.5 receives 74%, Sol – 73.8%, Terra – 72.5%, Luna – 41.3%.
- Toolathlon GPT-5.5 scores 55.6, Terra 53.1, Luna 53.4. Only Sol is higher with 58.
- The GPQA Diamond GPT-5.5 is stronger than Terra and Luna: 93.6% against 92.9% and 92.3%. Sol gets 94.6 percent.
- On the MMMU Pro with tools GPT-5.5 shows 83.2%, Terra – 82%, Luna – 79.5%. Sol leads with 84.6 percent.
- On the FrontierMath Tier 4, the Terra and Luna are also inferior to the previous model: 68.3% and 58.5% against 72.5% for the GPT-5.5. Sol goes up to 83%.
This is a good argument against automatic migration. The name of the new generation says nothing about a specific route: the model for short code correction, PDF analysis, and fact-finding in a millionth context can be different.
How much does a typical task cost
Imagine an agent launch with 20,000 input and 5,000 output tokens. Without cache and additional tools, its base cost will be:
| Модель | Один запуск | 100 запусков |
|---|---|---|
| GPT‑5.5 | $0,25 | $25 |
| GPT‑5.6 Sol | $0,25 | $25 |
| GPT‑5.6 Terra | $0,125 | $12,50 |
| GPT‑5.6 Luna | $0,05 | $5 |
The calculation does not take into account hidden reasoning tokens, repeated attempts, calls to paid instruments and the difference in the number of steps. Therefore, it is necessary to compare not the price of one token, but the cost of a successfully completed task.
For such a comparison, a set of 20-50 real examples is enough. On each model, it is worth measuring the share of accepted results, the number of repeated launches, the total number of tokens, the delay and the total cost. If Terra solves the problem on the first try, it is actually twice as profitable as Sol. If Luna requires three fixes, its low fare is no longer an advantage.
Other changes to GPT-5. 6
Benchmark is just one part of the release. The family has several new mechanisms for agent applications:
- *Programmatic Tool Calling allows a model to write a small program, call up permitted tools from it, and shorten intermediate results before they return to mainstream context.
- Multi-agent in the Responses API runs parallel subagents and combines their work in a single request.
- Obvious caching of the prompt allows you to mark stable parts of the context. The minimum life of the cache is 30 minutes.
- **Preserving reasoning between moves helps to continue long work without completely restarting logic on every query.
- *Pro mode is now enabled by the
reasoning.mode: "pro"parameter of the selected model, rather than the separate suffix in its name.
For a regular update of the application, all this is not required to include. The safest order is to first replace the model with the current level of reasoning, check the behavior, and then separately test the new cache, max, Pro or several agents.
Accessibility
OpenAI began the global deployment of GPT-5.6 on July 9 and spent about a day on it.
- In conventional ChatGPT, Plus, Pro, Business, and Enterprise users receive Sol at medium and higher levels of reasoning. Sol Pro is available on both Pro and Enterprise.
- In ChatGPT Work and Codex, Free and Go users get Terra. Starting with Plus, you can choose Sol, Terra or Luna and change the level of reasoning.
- All three identifiers are available in the API:
gpt-5.6-sol,gpt-5.6-terraandgpt-5.6-luna. The short identifiergpt-5.6directs requests to Sol.
What's the bottom line
The most important part of the release of GPT-5.6 is not a new Sol record, but a change in the price of the operating level of models. Terra costs half the price of GPT-5.5 and surpasses it in many code, browser and professional analysis tasks. For most API products, it is the first candidate for testing.
Sol is needed where the model acts on its own for a long time, manages the tools and should give the finished result with a minimum probability of error. Luna is useful for cheap mass operations with understandable automatic verification. It is too early to remove GPT-5.5 from the router: it retains an advantage in individual tests of long context, academic reasoning and toolwork.
Switching to GPT-5.6 is better understood not as a replacement for one line, but as a route adjustment. Sol, Terra and Luna are designed for different roles, and the final benchmarks for the first time provide enough data to allocate these roles not by marketing description, but by measurable outcome.
*The material is current on July 10, 2026. *
Sources
- Официальный релиз GPT‑5.6XX
- Руководство OpenAI по семейству GPT‑5.6XX
- Карточка GPT‑5.6 SolXX
- Карточка GPT‑5.6 TerraXX
- Карточка GPT‑5.6 LunaXX
- GPT‑5.6 System CardXX
- Цены OpenAI APIXX
FAQ
Which model is best? The one that reaches your acceptance criteria with the least correction cost for the target task.
Can benchmarks predict production results? They provide a signal, not a substitute for your repository, constraints and review process.
What should an evaluation record? Task success, tests, omissions, latency, token cost and manual intervention.