In this article+
01 / ARTIFICIAL INTELLIGENCE
GPT-5.6 Luna, Terra or Sol: What Is the Difference?
I went into GPT-5.6 expecting Sol to win by default. It does win some jobs. It also chews through a five-hour work window surprisingly quickly, and it still asks for another round when the brief is vague. After a few months, my answer became less exciting and more useful: I reach for Luna Max first.
That does not make the three models interchangeable. Luna, Terra and Sol have different prices, different amounts of reasoning effort and, judging from the Reddit comparison, very different ways of spending time. The model name alone is not enough; max, xhigh and high change the deal quite a bit.

| Model | Where it fits | My rough starting point |
|---|---|---|
| GPT-5.6 Luna | Cheap, repeatable work that a person or test can check | Drafts, small coding changes, extraction and routine fixes. |
| GPT-5.6 Terra | A larger, clearly planned implementation | Several files, broader tracing or execution after the plan is settled. |
| GPT-5.6 Sol | The difficult, unclear or expensive-to-get-wrong job | Architecture, stubborn debugging and deep analysis. |
That is the whole article in miniature. The rest is the receipt: what the Reddit chart actually shows, where Terra sits, why Luna Max works for me and when I would still pay for Sol.
02 / ARTIFICIAL INTELLIGENCE
The Reddit Cost Chart, Read Properly

The chart comes from a Reddit post comparing Pass@1, average task cost, output tokens and steps. Pass@1 is the percentage of tasks passed on the first attempt in that evaluation. It is not a guarantee that Sol will nail your next repository or that Luna will mess up your next email.
The 15 rows shown in the chart are listed below. A screenshot is handy for the overall shape, but a table makes the numbers much easier to search and compare. One footnote matters: this is a community comparison, not an official OpenAI price table and not a universal test.
| Model / setting | Pass@1 | Average cost | Output | Steps |
|---|---|---|---|---|
| GPT-5.6 Sol max | 73% ± 3% | $8.39 | 60k | 61 |
| GPT-5.6 Sol xhigh | 71% ± 1% | $4.70 | 41k | 44 |
| GPT-5.6 Terra max | 70% ± 3% | $3.96 | 72k | 76 |
| GPT-5.6 Sol high | 69% ± 1% | $3.47 | 28k | 37 |
| GPT-5.6 Luna max | 67% ± 4% | $0.61 | 73k | 102 |
| GPT-5.6 Sol medium | 61% ± 2% | $1.86 | 18k | 31 |
| GPT-5.6 Terra xhigh | 60% ± 2% | $1.70 | 40k | 43 |
| GPT-5.6 Luna xhigh | 57% ± 2% | $0.31 | 45k | 71 |
| GPT-5.6 Terra high | 54% ± 4% | $0.91 | 22k | 34 |
| GPT-5.6 Sol low | 45% ± 2% | $1.07 | 11k | 23 |
| GPT-5.6 Luna high | 44% ± 3% | $0.16 | 26k | 49 |
| GPT-5.6 Terra medium | 35% ± 3% | $0.47 | 12k | 25 |
| GPT-5.6 Terra low | 24% ± 1% | $0.34 | 8.6k | 21 |
| GPT-5.6 Luna medium | 11% ± 1% | $0.04 | 8.2k | 24 |
| GPT-5.6 Luna low | 2% ± 1% | $0.01 | 3.1k | 12 |
The first thing that jumps out is Luna max: 67% at $0.61. Sol max is 73% at $8.39. So the chart gives Sol a six-point lead and roughly fourteen times the average cost. That is a big gap to pay for if the task is easy to review. It is a much smaller gap if the task can quietly break a production system.
The second thing is less flattering for simple rankings. Luna max uses 102 steps, while Sol max uses 61. Cheap does not mean quick, and a high score does not mean short. The model can arrive at the right answer by taking the long way round. Sometimes that is perfectly fine. Sometimes you have already gone for coffee twice.
03 / ARTIFICIAL INTELLIGENCE
What Each Model Is Good At
On the current GPT-5.6 page, OpenAI lists Sol at $5 per million input tokens and $30 per million output tokens, Terra at $2.50 and $15, and Luna at $1 and $6. Those numbers matter when you run the API. If you are using a work product with a rolling five-hour allowance, the effect is felt more as quota and waiting time than as a neat little invoice.
Here is how I read the three names in day-to-day work. Luna is the model I want when a human can spot the remaining error. Terra is useful once the job has a shape. Sol earns its keep when the hard part is deciding what to do, not merely doing more of it.

| Task in front of you | Try first | Move up when… |
|---|---|---|
| Short edit, extraction or small patch | Luna high / xhigh | It keeps missing a requirement or the checks are no longer cheap. |
| Defined work across several files | Terra high / xhigh | The plan is wrong, the scope widens or the judgement becomes the hard part. |
| Architecture, deep debugging or research | Sol high / xhigh | The risk of a subtle mistake is larger than the extra usage. |
| A genuinely critical problem | Sol max | Only when the avoided failure is worth the premium route. |
One small trap: people say they tested “Luna” or “Sol” as if each name describes one fixed experience. It does not. Luna low and Luna max are miles apart in the chart. The setting belongs in the comparison, otherwise you are comparing a car in first gear with the same car on the motorway and calling the result surprising.
04 / ARTIFICIAL INTELLIGENCE
The Chart Cost Is Not Your Account Limit
This is where a lot of model comparisons quietly go off the rails. A $0.61 Luna Max row does not mean your account will deduct sixty-one cents. An $8.39 Sol Max row does not arrive as an API bill either. Depending on the product, you may be dealing with messages, reasoning usage, a rolling window, plan-specific routing or several of those at once.
You still feel the difference. Heavy Sol runs leave less room for retries and other work. That is why I keep coming back to the five-hour limit: it is not an abstract pricing detail when you are in the middle of a workday and the model has just spent a large chunk of your available time thinking through a task that still needs a correction.

| Number | What it tells you | What it does not tell you |
|---|---|---|
| API token price | The published cost when you control the request | Your exact ChatGPT or Codex quota deduction. |
| Average task cost | A relative cost from this particular evaluation | The price of your next task. |
| Output tokens | How much generated work came out | Whether the work was good. |
| Steps | How many turns/actions the run used | How many edits you will personally need. |
| Five-hour limit | The practical amount of work left in your window | A benchmark score. |
I think of it as three clocks. Money, model effort and available working time. They do not always tick together, which is why the cheapest row on a screenshot can still feel slow, while a premium run can feel expensive before you have seen a single dollar figure.
05 / ARTIFICIAL INTELLIGENCE
Why Luna Max Is the Interesting Option
The chart puts Luna Max at 67% Pass@1, $0.61, 73k output tokens and 102 steps. That is a strange combination. It uses more steps than Sol Max, but costs dramatically less. It is not elegant; it is useful.
Compare the nearby rows and the picture gets clearer. Luna xhigh is 57% at $0.31. Luna max is 67% at $0.61. In this test, an extra thirty cents buys ten points. Terra max reaches 70% at $3.96, while Sol high reaches 69% at $3.47. There is no smooth staircase where every extra dollar buys the same improvement.
| Comparison | What the chart says | What I would ask next |
|---|---|---|
| Luna max / Sol max | 6 points apart; roughly 14× apart in displayed cost | Can I check Luna’s mistakes cheaply? |
| Luna xhigh / Luna max | 10 points apart for $0.30 more | Has this small task become sticky enough for max? |
| Terra max / Sol high | Terra is 1 point higher but costs $0.49 more here | Do I need Terra’s execution style or Sol’s stronger planning? |
So yes, Luna Max can wander. That is the trade. If the job is a draft, a testable patch or a table a person will inspect, I can live with the scenic route. If it is changing live records or making a decision nobody will review, I would rather spend more and add a human checkpoint anyway.
06 / ARTIFICIAL INTELLIGENCE
I Used Luna Max for Months
This is the part that matters most to me, because it came from work I was actually trying to finish. Luna Max does not always understand the brief the first time. I sometimes need to restate the tone, point at the exact section that went wrong or make the requirement painfully explicit. Two or three rounds is normal for demanding work.
The surprise was that Sol did not remove that relationship. It produced stronger first attempts on some tasks, especially when the problem was genuinely difficult, but I still had to steer it. And while I was doing that, my five-hour allowance disappeared much faster. That is a poor bargain for work where the final result still needs my judgement.

| What I noticed | My answer |
|---|---|
| Can Luna Max reach a usable result? | Usually, if I give it boundaries and an example. |
| Does it follow every small instruction first time? | No. I expect a couple of targeted revisions. |
| Does Sol always save time? | No. It can still need several rounds in my workflow. |
| Why keep Luna as default? | It leaves enough room to finish other work and retry when needed. |
“Workability” is my own word here. I mean: can I keep moving without babysitting every sentence, and do I still have enough usage left for the rest of the day? Luna Max wins that test for me. Someone working on a difficult codebase may reasonably land somewhere else. Fair enough—different job, different bill.
07 / ARTIFICIAL INTELLIGENCE
When Sol Is Actually Worth It
The Reddit author’s routing idea is sensible: start with the lighter Luna settings, use Luna Max when the work gets sticky, then move through Sol high, Sol xhigh and Sol Max as the consequences get bigger. I would not open Sol Max just because the prompt looks important. A vague prompt can make a powerful model confidently spend time on the wrong job.
Sol high is probably the most interesting premium middle. The chart gives it 69% Pass@1, $3.47, 28k output tokens and 37 steps. Sol xhigh moves to 71% at $4.70. Sol Max reaches 73% at $8.39. That is a sensible ladder; you do not need to jump straight to the most expensive rung.

| Sol setting | Chart row | I would use it for |
|---|---|---|
| Sol high | 69% / $3.47 / 37 steps | Hard work where cost still matters. |
| Sol xhigh | 71% / $4.70 / 44 steps | A task that has already resisted the lighter route. |
| Sol max | 73% / $8.39 / 61 steps | The difficult edge case, not every Tuesday afternoon. |
There is one catch, and it is not very dramatic: give Sol a proper brief. Files, limits, acceptance criteria, tests. Otherwise you are paying a specialist to improvise. Specialists can improvise, of course. That does not mean you will like the result.
08 / ARTIFICIAL INTELLIGENCE
Where Terra Fits
Terra Max scores 70% at $3.96 in the chart, with 72k output tokens and 76 steps. Terra xhigh is 60% at $1.70 and 43 steps. On those numbers, Sol high looks like a better middle deal for broad reasoning. Terra Max spends more than Sol high for one extra point in this particular evaluation.
The comments tell a different story for some users. One says Terra xhigh is faster than Sol. Another uses Terra for implementation and Sol for planning. Someone else prefers Luna xhigh because Sol tends to overthink straightforward tasks. I would not smooth those comments into one verdict; they are describing different kinds of work.
| If you already have… | Terra may suit you | Check this before committing |
|---|---|---|
| A plan and a defined set of files | It can carry out a broader implementation without jumping straight to Sol Max. | Ask for a change log and tests; extra steps add extra places to drift. |
| A need for more reach than Luna | Terra xhigh is worth a controlled trial. | A speed win in one repository may disappear in another. |
| An architecture that is still undecided | Use Terra to map the options, then use Sol for the hard judgement if needed. | Fast execution of the wrong plan is still the wrong result. |
So I see Terra as the practical middle, not a consolation prize. It is the model I would try when the job has been planned but is too wide for a lightweight run. If the plan is still foggy, I would spend the budget on thinking first.
09 / ARTIFICIAL INTELLIGENCE
How I Choose a Model Now
- Small and obviousLuna high or xhigh. The work is easy to inspect, so I do not need the expensive route.
- Normal daily workLuna Max. This is where my own balance has held up.
- Defined multi-file workTerra high or xhigh. I want more reach, but I already know the plan.
- Unclear or high-risk workSol high. The hard part is judgement, not typing speed.
- A stubborn or costly failureSol xhigh or Max. The premium is easier to justify once I know what the cheaper route could not solve.
| Situation | My first choice | Why |
|---|---|---|
| A short task with an obvious check | Luna high / xhigh | Cheap corrections. |
| My regular work | Luna Max | Best balance in my actual workflow. |
| A planned implementation | Terra high / xhigh | More execution range without paying Sol Max immediately. |
| Architecture or difficult debugging | Sol high | The judgement is worth the extra spend. |
| A high-cost-to-fail task | Sol xhigh / Max | The avoided mistake is worth more than the quota. |
This will not be everyone’s routing order. It is mine today. I would rather have a rule that can change after five real tests than a favourite model I defend because I saw one impressive screenshot.
10 / ARTIFICIAL INTELLIGENCE
If Luna Needs Two More Tries, Is It Still Cheaper?
If Luna gives me a workable draft and I only need to adjust the audience, one paragraph or a test case, two extra rounds are not a disaster. I can see what changed. The task remains cheap to check. That is where Luna earns its keep.
If every round is a full rerun, the model keeps bringing back the same defect or the output touches money, customers or production data, I would move up sooner. Saving on the model and spending the afternoon debugging its leftovers is not really saving. It just hides the invoice in your calendar.
| Luna is still attractive when… | I would move up when… |
|---|---|
| The correction is narrow and visible. | The same error survives two focused corrections. |
| A test, checklist or quick review catches mistakes. | The result carries legal, financial, safety or production consequences. |
| The brief is stable and has examples. | The requirements are moving while the model works. |
| You need room for many tasks in one window. | The task is rare and expensive to redo. |
I would measure cost per finished task, not cost per first answer. Include the review, the retries and the time spent explaining the same thing again. Your result may match the Reddit chart. It may not. Either outcome is useful because it belongs to your work.
11 / ARTIFICIAL INTELLIGENCE
Why Reddit Users Disagree
One commenter says Sol completed a hard project in one shot where several Luna sessions failed. Another says Sol overthinks and makes simple work harder. Someone prefers Terra xhigh for speed. Another uses Sol for planning and Terra for implementation. Those experiences can all be honest at the same time.
A writer cares about tone and whether the model obeys a small instruction. A developer cares about tests and whether the change breaks something three files away. A researcher may care about source tracking. Same model name, different pain point.
| What changes the result | Why it matters |
|---|---|
| The task | Coding, writing, research and implementation reward different behaviour. |
| The prompt | A clear acceptance test leaves less room for the model to invent its own target. |
| The product | Chat, Codex, Work and API can have different tools, limits and routing. |
| The setting | Low, medium, high, xhigh and max are different runs. |
| The reviewer | A fast, careful reviewer can safely use a cheaper model more often. |
| The date | Prices, limits and rollout behaviour move while screenshots stay still. |
That last point is worth remembering. A chart is a snapshot. Keep it for the pattern, check the current OpenAI page for the product details and run a few of your own tasks before changing a workflow. Otherwise you are asking an old screenshot to explain a living product. It will not answer.
12 / ARTIFICIAL INTELLIGENCE
My Pick: Luna Max
So, which one should you pick? I would begin with Luna Max. Give it a proper brief, examples and a way to check the result. If it needs two or three alterations, that is not automatically a failure. Watch what those alterations cost you.
Move to Terra when the job is broader but already planned. Move to Sol when the work is genuinely ambiguous, difficult or expensive to get wrong. Sol is useful. I just do not think it should be running every ordinary task in the background like a sports car being used to buy groceries.
| Question | My answer |
|---|---|
| Best daily balance? | GPT-5.6 Luna Max, based on my own few months of use. |
| Where does Terra fit? | Planned, broader implementation work. |
| When is Sol worth it? | When judgement, depth or the cost of failure justifies the extra usage. |
| Does the Reddit chart settle it? | No. It gives you a useful starting picture; your tasks finish the comparison. |
| Do two or three Luna revisions ruin the value? | Not if each revision is cheap to check and the final task still costs less overall. |
The test I would run
Pick five real tasks from one working week. Use the same files and prompt where possible. Record the first-pass result, number of revisions, waiting time, usage and whether you would actually send or ship the final output. That small log is more useful than arguing over a model leaderboard.
There is no prize for choosing the most expensive model. There is no prize for choosing the cheapest one either. Pay for the extra reasoning when you can point to what it bought you. Otherwise, keep the lighter model moving.
Sources
Sources, dates and image credits
Image credits and licence notes
- photo-two-laptops.jpg — A software developer working with two laptops, photographed by Tsinkala. Model choice is a work decision, not a leaderboard hobby. Photographer: Tsinkala; licence: CC BY-SA 4.0. direct image.
- reddit-gpt56-cost-speed-chart.png — The cost, output-token and step comparison chart embedded in the Reddit discussion. It is a community comparison, not an official OpenAI price sheet or a universal benchmark. Hosted at Reddit post author; Reddit-hosted community chart; confirm reuse permission before publication. direct image.
- photo-programming-code.jpg — Programming code photographed by Martin Vorel. More reasoning steps can help, but they also have a bill attached. Photographer: Martin Vorel; licence: CC BY-SA 4.0. direct image.
- photo-stopwatch.jpg — A stopwatch photographed by Jeremyida002. Speed, waiting time and quota burn are three different clocks. Photographer: Jeremyida002; licence: CC BY-SA 4.0. direct image.
- photo-software-developer-africa.jpg — A software developer at work, photographed by Daudi mukiibi. The useful question is how many corrections it takes to reach a usable result. Photographer: Daudi mukiibi; licence: CC BY-SA 4.0. direct image.
- photo-code-review-meeting.jpg — A Wikimedia code-review meeting photographed by Matthew. A second pair of eyes is still part of the model workflow. Photographer: Matthew (Wikimedia Foundation); licence: CC BY-SA 3.0. direct image.
Research current through 2026-09-28 UTC. The OpenAI page supplies the official family descriptions, pricing context and published benchmark framing. The Reddit chart and comments are community evidence, not a controlled or universal evaluation. The author’s Luna Max experience is personal testing, not a claim that every account will see the same result. Confirm current access, pricing and reuse terms before publication.