Does switching models save money?
One switch nearly wiped out the savings.
Jevex finished 6% cheaper, saving just 0.50¢ across four tasks. That's too little evidence to promise reliable savings.
| Estimated cost in US cents | Jevex | Always Sol |
|---|---|---|
| First three tasks | 0.34¢ | 6.58¢ |
| Last task, after the switch | 7.24¢ | 1.50¢ |
| All four tasks | 7.58¢ | 8.08¢ |
Both passed 24/24 checks. Jevex used Luna → Luna → Luna → Sol. The other run used Sol for every task.
We estimated costs from token counts and public API prices. These aren't actual billed charges.
Why did the last task cost so much?
Codex can reuse parts of the conversation it has already processed. Reused input is cheaper. This is the prompt cache.
On the last task, Jevex reused 19% of its input. Always Sol reused 92%. Jevex also used more input and output tokens. Together, those differences ate up 92% of its earlier savings.
We can't tell how many cache misses the switch itself caused from this one test.
Luna's cache doesn't carry over to Sol. Sol may already have matching cached input from earlier Sol requests. Our switch still reported 19% cached input, but we haven't established where those hits came from.
How much does the cache discount matter?
For planning, assume 0% cached input on the first task after switching. A switch can still find some cache, but we shouldn't count on it. The slider starts at this conservative assumption.
Move the slider to change the percentage getting that discount on the last task only. The bars show the cost of all four tasks. Always Sol stays at 8.08¢.
Jevex would cost 0.86¢ more overall, 10.6% more expensive.
Higher percentages are hypothetical, not promised after a switch. The conversation text stays the same. Only its price changes. This calculation doesn't run Codex or predict how much cache it will find.
If less than 12.2% of the last task gets the discount, Jevex costs more overall. With no cache discount on that task, the total would be 8.94¢, or 11% more than Always Sol.
So, is Jevex worth it?
We don't know yet. The saving in this test was only 0.50¢. Reprocessing about 2,642 more input tokens on Sol would erase it.
With our conservative 0% cache assumption on the switched task, Jevex would cost 11% more overall. We need to test whether switches still save money under that assumption.
The comparison also started without any cached input. If it had reused 90% of its input on all four tasks, its estimated cost would be 4.56¢. The recorded Jevex run would cost 66% more. That's an assumption, not a result we measured.
Next, we need repeated tests on real coding work, with both runs starting from the same conversation and cache conditions. We should include switches back to Sol and compare the full cost of getting the work done.
See the task results and test limits
| Task | Run | Model | Checks passed | Input reused | Cost |
|---|---|---|---|---|---|
| Slug cleanup | Jevex | gpt-6-luna | 4/4 | 38% | 0.14¢ |
| Slug cleanup | Fixed Sol | gpt-6.1-sol | 4/4 | 0% | 4.72¢ |
| CSV field parsing | Jevex | gpt-6-luna | 5/5 | 70% | 0.13¢ |
| CSV field parsing | Fixed Sol | gpt-6.1-sol | 5/5 | 93% | 0.90¢ |
| Deterministic task scheduling | Jevex | gpt-6-luna | 5/5 | 92% | 0.07¢ |
| Deterministic task scheduling | Fixed Sol | gpt-6.1-sol | 5/5 | 93% | 0.96¢ |
| Untrusted expression parser | Jevex | gpt-6.1-sol | 10/10 | 19% | 7.24¢ |
| Untrusted expression parser | Fixed Sol | gpt-6.1-sol | 10/10 | 92% | 1.50¢ |
- We ran four small Python tasks once per approach. Passing these checks doesn't prove the models can complete larger coding jobs.
- We added the last task specifically to get a model switch. These tasks aren't a sample of everyday coding work.
- Jevex ran first. Its first task reused some input; the other run's first task reused none. The starting cache conditions weren't equal.
- We resumed both conversations in a new temporary directory for the last task. That changed the request context too.
- Codex reported running token totals. We subtracted the earlier totals to get each task's usage. Jev only saw the current prompt when choosing a model.
For a stronger test, use the same starting conversation, repeat in both orders, try longer contexts and tool use, and compare Always Sol with routing that switches less often. Count retries and failed work too.
See how we estimated cache costs
Why the last task cost 5.74¢ more
This splits the price difference by cache reuse and token counts. It doesn't prove the switch caused every cache miss.
| Difference | Extra cost |
|---|---|
| Lower cache share | 5.155¢ |
| More input tokens | 0.224¢ |
| More output tokens | 0.358¢ |
| Jev request | 0.002¢ |
What happens when previously reused text needs processing again?
Each row below shows the added cost of processing that many tokens at the regular input rate instead of the cheaper cached rate. It doesn't include output or retries.
| Input processed again | Extra on Luna | Extra on Sol |
|---|---|---|
| 10,000 tokens | 0.09¢ | 1.90¢ |
| 30,000 tokens | 0.27¢ | 5.70¢ |
| 100,000 tokens | 0.90¢ | 19.00¢ |
At the rates used here, new Luna input costs the same per token as cached Sol input. Luna's cheaper output can still help. Switching back to Sol is much more expensive if it needs to process a long conversation again.
What about charges for writing a new cache?
The API has separate prices for reading cached input, processing new input, and saving new input into the cache. Our logs listed zero cache writes, so we priced other input at the regular rate. We haven't checked those counts against an invoice.
If Jevex reused none of the last task's input and all of that input was charged at the cache-write rate, its total would be 10.80¢. That's 34% more than the recorded comparison. This is a separate what-if example; the slider uses regular input prices.
A model change can prevent reuse. Changes to the input or settings, an expired cache, or server placement can also cause misses. See OpenAI's cache guide and cache diagnostics.
Prices used for this test are from 2026-10-02: OpenAI API prices and Jev prices. We included Jev's routing charge. These figures don't measure ChatGPT billing or usage limits.