Essay · 2026-08-19
How I use OpenAI's GPT-5.6, based on 175,000+ tests
Surprising results from looking at the cost, intelligence, and latency tradeoffs between GPT-5.6 capability tiers and reasoning effort.
GPT-5.6 was recently released with three capability tiers: Sol, Terra, and Luna. Each tier supports six reasoning-effort settings: non-reasoning, low, medium, high, xhigh, and max. That gives us a fairly large menu of models and naturally begs the question: should you increase reasoning effort, or move up to a more capable tier?
To compare these 18 configurations, I used an Artificial Analysis snapshot. Its Intelligence Index v4.3 incorporates 10 evaluations across four categories—Agents (30%), Coding (20%), Scientific Reasoning (20%), and General (30%). The suite includes 91 AA-Briefcase tasks across four scenarios, 100 AA-LCR v1.1 questions, 70 CritPt tasks, and 2,158 text-only Humanity's Last Exam questions, among others. The 95% confidence interval for the aggregate index of less than ±1%, though individual evaluations may have wider intervals
Basic comparison
As expected, cost and performance move together. More reasoning effort and higher capability tier produce better performance at the expense of more tokens.

The important detail is that the two levers are not equally powerful. On the intelligence index, moving gpt-5.6-luna from medium reasoning effort to max reasoning effort adds 13 points. Moving from gpt-5.6-luna at max reasoning effort to gpt-5.6-terra at max reasoning effort adds only 5 points.
Cost versus intelligence
The tradeoff is less clear when cost enters the picture. The below shows us that, generally speaking, improved intelligence costs us more in terms of token usage. Note that the cost axis is in terms of dollars. This is the right metric because it is the most realistic, taking into account that:
- different models may use different numbers of tokens on a given benchmark task, not just the published token cost for the model;
- different models may use different amounts of reasoning tokens, which costs tokens, even if they don't always show up in the final result;
- model outputs and inputs can be cached, which can have lower cost on read than new input.
These effects are hard to predict so utilizing dollars encapsulates all this complexity into a single metric.

Intuitively, we can tell that Terra is expensive for its performance. In several settings, it performs worse than Sol while costing only a little less. It lacks a clear niche:
- Luna is a cheap, capable everyday model.
- Sol is the high-end model when accuracy matters most.
- Terra sits awkwardly between them.
This becomes clearer with a Pareto frontier. A model is on the cost-intelligence Pareto frontier when no other model is both cheaper and more intelligent. In other words, each frontier option represents a point where you cannot improve one dimension without giving up something in the other.

Terra never reaches this frontier. There is always a cheaper, more intelligent alternative. For example, gpt-5.6-luna with max reasoning effort is both smarter and cheaper than gpt-5.6-terra with medium reasoning effort. Likewise, gpt-5.6-sol with medium reasoning effort is both cheaper and smarter than gpt-5.6-terra with xhigh reasoning effort.
It also does not make sense to move to Sol before medium reasoning effort. Until then, there is always a Luna configuration that is both cheaper and more intelligent. This is the recommendation I have been following:
- Start with
gpt-5.6-lunaat the lowest effort that works. - Increase Luna's reasoning effort as the task gets harder.
- Switch to
gpt-5.6-solatmediumreasoning effort when Luna is no longer enough. - Skip Terra unless a particular workload gives you a reason not captured by this snapshot.
In fact, I've usually settled at medium reasoning effort as my workhorse (it's a replacement for deepseek v4 that's probably equally performant and cheaper, but that's an article for another day).
Time versus intelligence
Cost is not the only constraint. We can also compare response time with intelligence.

The time-intelligence Pareto frontier looks similar, but some of the details have shifted.

Terra starts to play a role here: you might choose gpt-5.6-terra at low reasoning effort over gpt-5.6-luna at medium reasoning effort because it takes about the same time and is more intelligent. It is not cheaper, but it is faster than the Luna settings that reach comparable accuracy.
Similarly, Sol starts to become useful at lower reasoning effort. Using gpt-5.6-sol at low reasoning effort is useful when you need Sol-level accuracy quickly. Nothing else reaches that level of accuracy in as little time, even though other options cost less.
Time is usually less of a constraint than cost, so I generally optimize for cost. Sometimes, though, speed matters more. When I am well under quota for the week and my subscription is about to restart, I am not very token-sensitive. In those cases, I will switch to gpt-5.6-terra at low reasoning effort instead for faster performance.
The simple rule
For most work, stay on Luna and increase effort before changing tiers. Luna gives you the best range of inexpensive options, while Sol provides a clear upgrade path for difficult tasks. Terra can be reasonable when latency matters, but it is not the best choice if cost and intelligence are the main concerns.