
If you are in a hurry: when my Claude quota runs out, I use GPT-6 Luna by default (half a cent per task), GPT-6.1 Sol for the delicate work, and I wait for Claude for incidents. But the part worth reading is why, and a single task tells it: fixing a date bug. Ten of thirteen models got it wrong, and six of them returned bad data without raising any error.
The quota runs out on Thursday
I use Claude Code every day, and with heavy use the subscription quota runs out before the week does. When that happens there are two options: wait for the reset, or keep going with another pay-as-you-go model.
I asked another AI which model to use, and the answer was easy: the cheapest "Flash" models, with a table of prices per million tokens. It sounds reasonable, but I had been there before.
A few months ago I tried MiniMax for the same job and it was not up to it. Nothing blew up: the problem was that the work looked finished and it was not. That is the expensive kind of failure, because you find out late.
So this time I measured before paying. Thirteen models, eight tasks taken from my own work, tests the model never saw, and less than ten dollars in total. The result changed what I have configured, and not for the reason I expected.
What I was really looking for
For heavy use, a subscription like Claude's is much cheaper than paying for the API token by token, and no pay-as-you-go option comes close. So the question is not "what replaces the subscription?" but "what covers the days without quota, without lowering the quality?".
And one tip before switching models: a coding agent re-reads the whole session context on every turn, so one long session burns through the quota much faster than several short ones. The cheapest way to stretch it is not to switch models, it is to use /clear between tasks.
How I measured it
I used opencode, a terminal coding agent similar to Claude Code, but open source and able to work with any provider. That way the thirteen alternative models worked with the same tools and the same loop; the only thing that changed was the model. With Claude I did something different, and I explain it below. The models came from two providers, OpenRouter and DeepInfra.
The eight tasks come from what I actually do, with synthetic data:
| Task | What it asks for |
|---|---|
| t1 | PowerShell script that exports the permissions of Microsoft 365 shared mailboxes, with Pester tests |
| t2 | Analyze a Unified Audit Log export to find suspicious inbox rules (external forwarding, hidden folders, payment keywords), typical of a BEC |
| t3 | Newsletter shortcode for a WordPress plugin, with nonce and escaping |
| t4 | Flask endpoint to re-run a scan, following the project's conventions |
| t5 | Fix a bug from a traceback |
| t6 | Remotion component in strict TypeScript (an animated lower third for video) |
| t7 | A bulk scan that sometimes hangs forever, with no error and no log |
| t8 | Refactor a service to separate the engine from the API without changing its behavior |
The rules of the test bench:
- Hidden tests. The model gets the task and the repository. When it finishes, I copy in tests it never saw and run them. I validated them first against a reference solution, so that a failure is the model's fault and not the test's.
- A clean copy per attempt, with a starting commit, to keep exactly what each model changed.
- Auto mode and a 20-minute limit. No help from me: if it asks, nobody answers.
- Blind review of tasks t4, t7 and t8: a reviewer scored the diffs, shuffled and unlabeled, without knowing which model wrote each one.
- Cost, checked against the invoice. For seven of the thirteen models, what opencode calculates matches what was billed or is within two cents. For four it does not: opencode used catalog prices higher than what DeepInfra and OpenRouter actually charged, up to double for GLM-5.3 Flash. The table uses the billed cost.
Claude did not run in opencode, and that needs saying. The reference is Claude Opus 5.5 in Claude Code, on a subscription: each task was done by an isolated subagent, with the same prompt, the same copy of the repository and the same hidden tests, but with Claude Code's tools, not opencode's. To have a like-for-like comparison, I also ran Claude Opus 5.5 through the API inside opencode on four tasks (t1, t4, t5 and t7): it passed all four, t5 included, at about 0.19 dollars per task.
And the blind reviewer was also Claude. A Claude Code subagent running Opus 5.5, the same model as the reference. It did not know who wrote each diff, but a model may prefer the style it would write itself. That is why I use the review scores only as support; the conclusion comes from the tests.
Each attempt was launched like this:
opencode run --dir ./trabajo -m openrouter/openai/gpt-6-luna --auto --format json "$(cat prompt.md)"
--auto approves everything that is not explicitly forbidden. Since I do not like letting a model run commands unattended, each copy carried an opencode.json with what it could not do:
{
"permission": {
"edit": "allow",
"webfetch": "deny",
"external_directory": "deny",
"bash": {
"*": "allow",
"git push*": "deny",
"curl *": "deny",
"Invoke-WebRequest*": "deny",
"ssh *": "deny",
"Remove-Item C:\\*": "deny"
}
}
}
A warning, since this is a security blog: this reduces risk, it does not contain anything. The rules compare text, so a python -c with urllib gets past curl * with no trouble. For code you do not know, use a virtual machine or a container with no network.
Results
This is how the eight tasks turned out, with Claude as the reference. Cost is what I paid per task on average, and the last column is task t5, which I explain in the next section:
| Model | Tests | $/task | Median | t5 |
|---|---|---|---|---|
| Claude Opus 5.5 (Claude Code, reference) | 8/8 | quota | – | ✅ |
| Claude Opus 5.5 (API in opencode, control) | 4/4 | 0.19 | 90 s | ✅ |
| GPT-6 Luna | 8/8 and 7/8 (two runs) | 0.005 | 103 s | ✅ both times |
| GPT-6.1 Sol | 8/8 | 0.077 | 121 s | ✅ |
| GPT-6 Sol | 8/8 | 0.082 | 187 s | ✅ |
| Qwen3.8 Max | 7/8 | 0.099 | 282 s | ⚠️ detail |
| Kimi K3 | 7/8 | 0.195 | 283 s | ⚠️ detail |
| Grok Build | 7/8 | 0.070 | 258 s | ⚠️ detail |
| DeepSeek V4 Pro | 7/8 | 0.044 | 221 s | ❌ wrong dates |
| DeepSeek V4.1 Flash | 7/8 and 7/8 (two runs) | 0.004 | 63–82 s | ❌ wrong dates, both times |
| Jev Router (TypeSafe on OpenRouter) | 7/8 | not logged | 99 s | ❌ wrong dates |
| GLM-5.3 | 6/8 | 0.09 | 146 s | ⚠️ detail |
| GLM-5.3 Flash | 6/8 | 0.004 | 341 s | ❌ wrong dates |
| MiniMax M2.7 | 5/8 | 0.055 | 120 s | ❌ wrong dates |
| MiniMax M3 | 3/8 | 0.019 | 741 s | ❌ wrong dates |

Looking only at the tests column, it seems almost everyone passes: 7 out of 8 here, 7 out of 8 there. The 7/8 of almost all of them comes from the same task, t5.
The exception is Luna's 7/8 in its second run, and since it is my default model, I will explain it. It failed t6, the animated lower third: the prompt asked for the opacity to go from 1 to 0 over the last 15 frames, and Luna made it reach 0 already on the last visible frame, so the lower third disappears one frame earlier than the test expected. In a 30 frames per second video that is a thirtieth of a second, and the prompt allows both readings. But the test flags it, and I count it as a failure.
The task that decided everything: a date
t5 is short. A script normalizes the data breaches returned by different threat intelligence providers, and when importing those from a new provider this error shows up:
breaches.py:33: in parse_leak_date
return date.fromisoformat(s)
E ValueError: Invalid isoformat string: '2024-03'
The prompt was the traceback and one sentence: "The function must do what its docstring says." And the docstring says this (translated from the Spanish original):
def parse_leak_date(value: str | None) -> date | None:
"""Converts a breach date to `date`.
Formats the providers send:
- Full ISO: "2024-03-15" -> 2024-03-15
- Year-month only: "2024-03" -> 2024-03-01
- European: "15/03/2024" -> 2024-03-15 (day/month/year)
- English month: "March 2024" -> 2024-03-01
...
Any other value raises ValueError with the original value in the message.
"""
The traceback bug is easy: the branch that handles dates with a dash swallows 2024-03 before it reaches the year-month branch. But there is a second bug the traceback does not show. Whoever reads the contract and compares it with the code sees it:
if "/" in s:
month, day, year = (int(p) for p in s.split("/")) # month/day: US format
return date(year, month, day)
The docstring says day/month/year and the code reads month/day/year. No existing test covers it.
The models fell into three groups.
The ones that only fixed the symptom. DeepSeek V4 Pro, DeepSeek V4.1 Flash, GLM-5.3 Flash, the Jev Router and both MiniMax models did practically the same thing: move two lines.
if "T" in s:
return datetime.fromisoformat(s.replace("Z", "+00:00")).date()
+ if len(s) == 7 and s[4] == "-":
+ return date(int(s[:4]), int(s[5:]), 1)
+
if "-" in s:
return date.fromisoformat(s)
The traceback goes away and the repository's tests pass. It is a fix anyone would approve in a quick review. But 05/03/2024, March 5, now comes out as May 3, with no error at all. With 15/03/2024 it at least fails, because there is no month 15. With any day from 1 to 12 you get a valid, wrong date.
Think about where that ends up: a breach report with the dates swapped, an incident timeline out of order, a "last 30 days" filter that leaves out what it should not. Nobody notices until someone compares against the original.
DeepSeek V4.1 Flash did exactly the same on the second run. It was not bad luck: it is how it behaves on this kind of task.
The ones that fixed the dates but not the message. Qwen3.8 Max, Kimi K3, Grok Build and GLM-5.3 read the docstring and fixed the European format. They failed on the last line: for an invalid value such as 13/13/2024, the error said month must be in 1..12, not 13 instead of including the original value. It is a detail, and I mark it as a detail. But it was also written in the contract.
The ones that met the whole contract. Claude, GPT-6 Sol, GPT-6.1 Sol and GPT-6 Luna. Luna, on top of that, in both runs. Its fix checks each format with an exact regular expression instead of looking for a single character, reads day/month/year, and turns any failure into the ValueError with the original value.

What surprised me most is that price predicted nothing. Kimi K3 was the most expensive model in the test, at 0.20 dollars per task, and it missed the message detail. GPT-6 Luna cost 0.005 dollars, almost forty times less, and met everything. The price table the other AI gave me sorted the models by the one thing that did not matter.
What the tests do not see
Hidden tests catch a lot, but not everything. That is why there was a blind review.
The race condition in t4. The endpoint had to return 409 if a scan was already running. Almost everyone did it like this: read the status, check it, write it. If two requests arrive at once, both read "free" and two scans start. Neither the prompt nor the tests mentioned it. Only Claude and the two GPT Sol models protected the check and the status change with a lock, so they could not happen at the same time. GPT-6.1 Sol even added a test that fires two requests in parallel. Luna did not. That is why I do not give Luna anything involving concurrency.
More is not better. In t7, the scan that hung, all six models in the blind review found the cause: the thread only caught DNS errors, any other error killed it and the queue waited forever. Claude fixed it by turning a 43-line module into a 128-line one, with its own polling loop and per-domain timeouts. The blind reviewer called it "out of proportion" and gave it 1 out of 5 for faithfulness to the request. GPT-6 Sol solved it with a change of about 40 lines, tests included, and got the top score. Luna, which came after the blind review, also found the complete solution in its first run; I checked this myself by reading its diff: catch any exception, guarantee with a finally that every domain ends up in results or in errors, and validate the number of threads.
The overall scores from the first-round blind review, out of 5:
| Model | t4 | t7 | t8 |
|---|---|---|---|
| GPT-6 Sol | 5 | 5 | 4 |
| Qwen3.8 Max | 3 | 4 | 5 |
| Kimi K3 | 3 | 4 | 4 |
| Grok Build | 3 | 3 | 3 |
| DeepSeek V4 Pro | 2 | 2 | 3 |
| Claude Opus 5.5 | 2* | 2* | 3* |
*All three of Claude's solutions deleted opencode.json and added a copy of the hidden tests. Neither came from the model: Claude ran in Claude Code, where that file is not used, and the copy of the tests was left by my grading process. The reviewer could not know that and penalized them, rightly: as they stood, they could not be merged. Looking only at code correctness, it gave them 5, 4 and 5.
Quirks worth knowing
GLM-5.3 thought for 32,000 tokens and wrote nothing. On the first task it reasoned up to the output limit, created no files and cost 0.08 dollars. I repeated it with a different reasoning setting and exactly the same thing happened. If a model can spend its entire budget thinking, you want to know before leaving it alone.
MiniMax came last, and that validates the test. MiniMax M3 passed 3 out of 8 and hit the 20-minute limit on several tasks; M2.7 passed 5 out of 8. That is exactly what I lived using it. If my test bench had scored well the model that failed me in practice, the bench would be the thing that was wrong. That it reproduced it gave me confidence in the rest of the results.
The router does not tell you what it did. The Jev Router, TypeSafe's router on OpenRouter, picks a model for each request. It failed t5 like the cheap models, and opencode's log does not show which model it picked or how much it cost: it records 0 dollars for all eight tasks. So you do not know who did what or what you are paying per task. I still think routing is a good idea, but with my own criteria and between models I have already measured, not blindly.
What I use now
A three-step ladder:
| Step | Model | For what | $/task |
|---|---|---|---|
| 1 (default) | GPT-6 Luna | Almost everything: scripts, components, plugins, refactors with tests | ~0.005 |
| 2 (precise) | GPT-6.1 Sol, with GPT-6 Sol as backup | Bugs, business logic, data parsing, concurrency, code without tests | ~0.08 |
| 3 | Wait for Claude | Tenant incidents, production, architecture, security | quota |
To decide which step a task goes to, I ask myself one question: can I objectively check that it came out right? If there are tests or a clear check, step 1. If not, step 2 at least. And for a real incident I do not trust two passed complex tasks: there I wait for Claude or pay for it through the API.
In opencode I have it as two agents, and you switch between them with Tab. This is the file, in ~/.config/opencode/opencode.jsonc, with no keys: it reads them from the OPENROUTER_API_KEY environment variable.
{
"$schema": "https://opencode.ai/config.json",
"enabled_providers": ["openrouter"],
"provider": {
"openrouter": {
"whitelist": ["openai/gpt-6-luna", "openai/gpt-6.1-sol", "openai/gpt-6-sol"],
"models": {
// GPT-6.1 Sol is very new and may not be in opencode's catalog yet
"openai/gpt-6.1-sol": {
"name": "GPT-6.1 Sol",
"tool_call": true,
"reasoning": true,
"cost": { "input": 2, "output": 10, "cache_read": 0.1 },
"limit": { "context": 1050000, "output": 128000 }
}
}
}
},
"model": "openrouter/openai/gpt-6-luna",
"small_model": "openrouter/openai/gpt-6-luna",
"agent": {
"build": { "model": "openrouter/openai/gpt-6-luna" },
"plan": { "model": "openrouter/openai/gpt-6.1-sol" },
"preciso": {
"mode": "primary",
"model": "openrouter/openai/gpt-6.1-sol",
"description": "Bugs, business logic, data parsing, concurrency and code without tests."
}
}
}
If you use OpenRouter, set a spending limit on the key. It is a checkbox in the dashboard, and an agent in auto mode with a loop that does not stop can spend a lot.
Two pricing details
For daily use, a subscription is still the cheapest option. What changes is that covering the days without quota with Luna costs cents per task. Two details from OpenAI's pricing page worth knowing before using these models in long sessions:
- The long context surcharge. With long prompts (from about 272,000 tokens, according to OpenRouter), input costs twice as much and output 1.5 times as much. In a long agent session it is easy to go past that. The fix is the same as before: a new session per task (
/newin opencode). - The Flex tier, at half price in exchange for waiting. It is in OpenAI's direct API and also on OpenRouter, which offers it as one more provider (
openai/flex) for GPT-6 Luna and GPT-6.1 Sol. I have not tried it with opencode, and the day I looked, GPT-6.1 Sol's Flex was degraded, with 89% availability over the last half hour.
Run your own test in an afternoon
My results apply to my tasks. If your work is different, the test is quick to set up:
- Pick 5 to 8 real tasks you do every week, with made-up data. They should cover your languages, and at least one should be hard.
- Write the tests first and hide them. Validate them with your own solution. The model does not see them until it finishes.
- Include a contract task: a bug with a traceback where the obvious fix leaves another bug that only shows up when you read the documentation. It is the one that separates models the most, and it tells you whether the model reads or only reacts.
- A clean copy and auto mode per attempt, with a time limit and dangerous commands forbidden.
- Repeat at least once with the models you care about. A single run cannot tell luck from habit.
- Compare the cost with the real invoice, not with what the tool says. opencode overestimated four models for me, GLM-5.3 Flash at double the billed amount, because it used a catalog price and not the price of the provider that actually served the requests.
- Include a model you already know from practice. If your test does not reproduce what you lived with it, the test is wrong.
What I did not measure
- Eight tasks, one or two runs. It is a strong signal, not a guarantee.
- Short tasks. Real work sessions are much longer, with a different mix of cache and context, and the cost per task may change.
- I designed t5, and a single task weighs a lot in the conclusion. That it repeated the same way across two runs gives it weight, but it is still one task.
- Only with opencode. With another agent, with other tools and instructions, the results may change. Claude's reference ran in Claude Code, not in opencode; the API control in opencode was only four tasks.
- An in-house judge. The blind reviewer was Claude Opus 5.5, the same model as the reference. It did not know who wrote each diff, but it is not a neutral judge.
- Models and prices from September 27 to 30, 2026. They change every few weeks. I plan to repeat this every three months.
Sources
- OpenAI: API pricing, checked on September 30, 2026.
- opencode: documentation and permissions.
- OpenRouter and DeepInfra: prices for each model in their catalogs, September 27 to 30, 2026, and their usage invoices for September.
- MiniMax: Token Plan pricing.
Keep reading on IT Rafa
- Jev picks the Claude Code model: I measured it before recommending it, the other half of this story: choosing a model inside Claude.
- AI-generated code: the real security cost.
- I rebuilt my blog with an AI: here is what broke.