· 17 min read

The Claude sphere, switched off with an empty battery, sends the work to a verdict panel. Lines run from it to seven models: OpenAI in green with a check mark and a $0.005 tag; Kimi, Grok, Qwen and Z.ai in orange with a warning sign; DeepSeek and MiniMax in red with a cross.

If you are in a hurry: when my Claude quota runs out, I use GPT-6 Luna by default (half a cent per task), GPT-6.1 Sol for the delicate work, and I wait for Claude for incidents. But the part worth reading is why, and a single task tells it: fixing a date bug. Ten of thirteen models got it wrong, and six of them returned bad data without raising any error.

The quota runs out on Thursday

I use Claude Code every day, and with heavy use the subscription quota runs out before the week does. When that happens there are two options: wait for the reset, or keep going with another pay-as-you-go model.

I asked another AI which model to use, and the answer was easy: the cheapest "Flash" models, with a table of prices per million tokens. It sounds reasonable, but I had been there before.

A few months ago I tried MiniMax for the same job and it was not up to it. Nothing blew up: the problem was that the work looked finished and it was not. That is the expensive kind of failure, because you find out late.

So this time I measured before paying. Thirteen models, eight tasks taken from my own work, tests the model never saw, and less than ten dollars in total. The result changed what I have configured, and not for the reason I expected.

What I was really looking for

For heavy use, a subscription like Claude's is much cheaper than paying for the API token by token, and no pay-as-you-go option comes close. So the question is not "what replaces the subscription?" but "what covers the days without quota, without lowering the quality?".

And one tip before switching models: a coding agent re-reads the whole session context on every turn, so one long session burns through the quota much faster than several short ones. The cheapest way to stretch it is not to switch models, it is to use /clear between tasks.

How I measured it

I used opencode, a terminal coding agent similar to Claude Code, but open source and able to work with any provider. That way the thirteen alternative models worked with the same tools and the same loop; the only thing that changed was the model. With Claude I did something different, and I explain it below. The models came from two providers, OpenRouter and DeepInfra.

The eight tasks come from what I actually do, with synthetic data:

TaskWhat it asks for
t1PowerShell script that exports the permissions of Microsoft 365 shared mailboxes, with Pester tests
t2Analyze a Unified Audit Log export to find suspicious inbox rules (external forwarding, hidden folders, payment keywords), typical of a BEC
t3Newsletter shortcode for a WordPress plugin, with nonce and escaping
t4Flask endpoint to re-run a scan, following the project's conventions
t5Fix a bug from a traceback
t6Remotion component in strict TypeScript (an animated lower third for video)
t7A bulk scan that sometimes hangs forever, with no error and no log
t8Refactor a service to separate the engine from the API without changing its behavior

The rules of the test bench:

Claude did not run in opencode, and that needs saying. The reference is Claude Opus 5.5 in Claude Code, on a subscription: each task was done by an isolated subagent, with the same prompt, the same copy of the repository and the same hidden tests, but with Claude Code's tools, not opencode's. To have a like-for-like comparison, I also ran Claude Opus 5.5 through the API inside opencode on four tasks (t1, t4, t5 and t7): it passed all four, t5 included, at about 0.19 dollars per task.

And the blind reviewer was also Claude. A Claude Code subagent running Opus 5.5, the same model as the reference. It did not know who wrote each diff, but a model may prefer the style it would write itself. That is why I use the review scores only as support; the conclusion comes from the tests.

Each attempt was launched like this:

opencode run --dir ./trabajo -m openrouter/openai/gpt-6-luna --auto --format json "$(cat prompt.md)"

--auto approves everything that is not explicitly forbidden. Since I do not like letting a model run commands unattended, each copy carried an opencode.json with what it could not do:

{
  "permission": {
    "edit": "allow",
    "webfetch": "deny",
    "external_directory": "deny",
    "bash": {
      "*": "allow",
      "git push*": "deny",
      "curl *": "deny",
      "Invoke-WebRequest*": "deny",
      "ssh *": "deny",
      "Remove-Item C:\\*": "deny"
    }
  }
}

A warning, since this is a security blog: this reduces risk, it does not contain anything. The rules compare text, so a python -c with urllib gets past curl * with no trouble. For code you do not know, use a virtual machine or a container with no network.

Results

This is how the eight tasks turned out, with Claude as the reference. Cost is what I paid per task on average, and the last column is task t5, which I explain in the next section:

ModelTests$/taskMediant5
Claude Opus 5.5 (Claude Code, reference)8/8quota–✅
Claude Opus 5.5 (API in opencode, control)4/40.1990 s✅
GPT-6 Luna8/8 and 7/8 (two runs)0.005103 s✅ both times
GPT-6.1 Sol8/80.077121 s✅
GPT-6 Sol8/80.082187 s✅
Qwen3.8 Max7/80.099282 s⚠️ detail
Kimi K37/80.195283 s⚠️ detail
Grok Build7/80.070258 s⚠️ detail
DeepSeek V4 Pro7/80.044221 s❌ wrong dates
DeepSeek V4.1 Flash7/8 and 7/8 (two runs)0.00463–82 s❌ wrong dates, both times
Jev Router (TypeSafe on OpenRouter)7/8not logged99 s❌ wrong dates
GLM-5.36/80.09146 s⚠️ detail
GLM-5.3 Flash6/80.004341 s❌ wrong dates
MiniMax M2.75/80.055120 s❌ wrong dates
MiniMax M33/80.019741 s❌ wrong dates
Dot chart of 12 models sorted by billed cost per task, on a log scale. The two cheapest, DeepSeek V4.1 Flash and GLM-5.3 Flash ($0.004), return wrong dates. GPT-6 Luna, at $0.005, meets the contract. Next come MiniMax M3, DeepSeek V4 Pro and MiniMax M2.7, with wrong dates; Grok Build with the wrong error message; GPT-6.1 Sol and GPT-6 Sol, which meet it, around $0.08; and GLM-5.3, Qwen3.8 Max and Kimi K3, the most expensive at $0.20, with the wrong error message.

Looking only at the tests column, it seems almost everyone passes: 7 out of 8 here, 7 out of 8 there. The 7/8 of almost all of them comes from the same task, t5.

The exception is Luna's 7/8 in its second run, and since it is my default model, I will explain it. It failed t6, the animated lower third: the prompt asked for the opacity to go from 1 to 0 over the last 15 frames, and Luna made it reach 0 already on the last visible frame, so the lower third disappears one frame earlier than the test expected. In a 30 frames per second video that is a thirtieth of a second, and the prompt allows both readings. But the test flags it, and I count it as a failure.

The task that decided everything: a date

t5 is short. A script normalizes the data breaches returned by different threat intelligence providers, and when importing those from a new provider this error shows up:

breaches.py:33: in parse_leak_date
    return date.fromisoformat(s)
E   ValueError: Invalid isoformat string: '2024-03'

The prompt was the traceback and one sentence: "The function must do what its docstring says." And the docstring says this (translated from the Spanish original):

def parse_leak_date(value: str | None) -> date | None:
    """Converts a breach date to `date`.

    Formats the providers send:
      - Full ISO:          "2024-03-15"            -> 2024-03-15
      - Year-month only:   "2024-03"               -> 2024-03-01
      - European:          "15/03/2024"            -> 2024-03-15   (day/month/year)
      - English month:     "March 2024"            -> 2024-03-01
      ...
    Any other value raises ValueError with the original value in the message.
    """

The traceback bug is easy: the branch that handles dates with a dash swallows 2024-03 before it reaches the year-month branch. But there is a second bug the traceback does not show. Whoever reads the contract and compares it with the code sees it:

    if "/" in s:
        month, day, year = (int(p) for p in s.split("/"))   # month/day: US format
        return date(year, month, day)

The docstring says day/month/year and the code reads month/day/year. No existing test covers it.

The models fell into three groups.

The ones that only fixed the symptom. DeepSeek V4 Pro, DeepSeek V4.1 Flash, GLM-5.3 Flash, the Jev Router and both MiniMax models did practically the same thing: move two lines.

     if "T" in s:
         return datetime.fromisoformat(s.replace("Z", "+00:00")).date()

+    if len(s) == 7 and s[4] == "-":
+        return date(int(s[:4]), int(s[5:]), 1)
+
     if "-" in s:
         return date.fromisoformat(s)

The traceback goes away and the repository's tests pass. It is a fix anyone would approve in a quick review. But 05/03/2024, March 5, now comes out as May 3, with no error at all. With 15/03/2024 it at least fails, because there is no month 15. With any day from 1 to 12 you get a valid, wrong date.

Think about where that ends up: a breach report with the dates swapped, an incident timeline out of order, a "last 30 days" filter that leaves out what it should not. Nobody notices until someone compares against the original.

DeepSeek V4.1 Flash did exactly the same on the second run. It was not bad luck: it is how it behaves on this kind of task.

The ones that fixed the dates but not the message. Qwen3.8 Max, Kimi K3, Grok Build and GLM-5.3 read the docstring and fixed the European format. They failed on the last line: for an invalid value such as 13/13/2024, the error said month must be in 1..12, not 13 instead of including the original value. It is a detail, and I mark it as a detail. But it was also written in the contract.

The ones that met the whole contract. Claude, GPT-6 Sol, GPT-6.1 Sol and GPT-6 Luna. Luna, on top of that, in both runs. Its fix checks each format with an exact regular expression instead of looking for a single character, reads day/month/year, and turns any failure into the ValueError with the original value.

Three columns for parse_leak_date("05/03/2024"), meaning March 5, 2024. Symptom only: returns 2024-05-03 with no error, six models. Dates right, message wrong: returns 2024-03-05 but the error for 13/13/2024 does not quote the value, four models. Full contract: 2024-03-05 and the error quotes the original value, three models plus Claude.

What surprised me most is that price predicted nothing. Kimi K3 was the most expensive model in the test, at 0.20 dollars per task, and it missed the message detail. GPT-6 Luna cost 0.005 dollars, almost forty times less, and met everything. The price table the other AI gave me sorted the models by the one thing that did not matter.

What the tests do not see

Hidden tests catch a lot, but not everything. That is why there was a blind review.

The race condition in t4. The endpoint had to return 409 if a scan was already running. Almost everyone did it like this: read the status, check it, write it. If two requests arrive at once, both read "free" and two scans start. Neither the prompt nor the tests mentioned it. Only Claude and the two GPT Sol models protected the check and the status change with a lock, so they could not happen at the same time. GPT-6.1 Sol even added a test that fires two requests in parallel. Luna did not. That is why I do not give Luna anything involving concurrency.

More is not better. In t7, the scan that hung, all six models in the blind review found the cause: the thread only caught DNS errors, any other error killed it and the queue waited forever. Claude fixed it by turning a 43-line module into a 128-line one, with its own polling loop and per-domain timeouts. The blind reviewer called it "out of proportion" and gave it 1 out of 5 for faithfulness to the request. GPT-6 Sol solved it with a change of about 40 lines, tests included, and got the top score. Luna, which came after the blind review, also found the complete solution in its first run; I checked this myself by reading its diff: catch any exception, guarantee with a finally that every domain ends up in results or in errors, and validate the number of threads.

The overall scores from the first-round blind review, out of 5:

Modelt4t7t8
GPT-6 Sol554
Qwen3.8 Max345
Kimi K3344
Grok Build333
DeepSeek V4 Pro223
Claude Opus 5.52*2*3*

*All three of Claude's solutions deleted opencode.json and added a copy of the hidden tests. Neither came from the model: Claude ran in Claude Code, where that file is not used, and the copy of the tests was left by my grading process. The reviewer could not know that and penalized them, rightly: as they stood, they could not be merged. Looking only at code correctness, it gave them 5, 4 and 5.

Quirks worth knowing

GLM-5.3 thought for 32,000 tokens and wrote nothing. On the first task it reasoned up to the output limit, created no files and cost 0.08 dollars. I repeated it with a different reasoning setting and exactly the same thing happened. If a model can spend its entire budget thinking, you want to know before leaving it alone.

MiniMax came last, and that validates the test. MiniMax M3 passed 3 out of 8 and hit the 20-minute limit on several tasks; M2.7 passed 5 out of 8. That is exactly what I lived using it. If my test bench had scored well the model that failed me in practice, the bench would be the thing that was wrong. That it reproduced it gave me confidence in the rest of the results.

The router does not tell you what it did. The Jev Router, TypeSafe's router on OpenRouter, picks a model for each request. It failed t5 like the cheap models, and opencode's log does not show which model it picked or how much it cost: it records 0 dollars for all eight tasks. So you do not know who did what or what you are paying per task. I still think routing is a good idea, but with my own criteria and between models I have already measured, not blindly.

What I use now

A three-step ladder:

StepModelFor what$/task
1 (default)GPT-6 LunaAlmost everything: scripts, components, plugins, refactors with tests~0.005
2 (precise)GPT-6.1 Sol, with GPT-6 Sol as backupBugs, business logic, data parsing, concurrency, code without tests~0.08
3Wait for ClaudeTenant incidents, production, architecture, securityquota

To decide which step a task goes to, I ask myself one question: can I objectively check that it came out right? If there are tests or a clear check, step 1. If not, step 2 at least. And for a real incident I do not trust two passed complex tasks: there I wait for Claude or pay for it through the API.

In opencode I have it as two agents, and you switch between them with Tab. This is the file, in ~/.config/opencode/opencode.jsonc, with no keys: it reads them from the OPENROUTER_API_KEY environment variable.

{
  "$schema": "https://opencode.ai/config.json",
  "enabled_providers": ["openrouter"],
  "provider": {
    "openrouter": {
      "whitelist": ["openai/gpt-6-luna", "openai/gpt-6.1-sol", "openai/gpt-6-sol"],
      "models": {
        // GPT-6.1 Sol is very new and may not be in opencode's catalog yet
        "openai/gpt-6.1-sol": {
          "name": "GPT-6.1 Sol",
          "tool_call": true,
          "reasoning": true,
          "cost": { "input": 2, "output": 10, "cache_read": 0.1 },
          "limit": { "context": 1050000, "output": 128000 }
        }
      }
    }
  },
  "model": "openrouter/openai/gpt-6-luna",
  "small_model": "openrouter/openai/gpt-6-luna",
  "agent": {
    "build":   { "model": "openrouter/openai/gpt-6-luna" },
    "plan":    { "model": "openrouter/openai/gpt-6.1-sol" },
    "preciso": {
      "mode": "primary",
      "model": "openrouter/openai/gpt-6.1-sol",
      "description": "Bugs, business logic, data parsing, concurrency and code without tests."
    }
  }
}

If you use OpenRouter, set a spending limit on the key. It is a checkbox in the dashboard, and an agent in auto mode with a loop that does not stop can spend a lot.

Two pricing details

For daily use, a subscription is still the cheapest option. What changes is that covering the days without quota with Luna costs cents per task. Two details from OpenAI's pricing page worth knowing before using these models in long sessions:

Run your own test in an afternoon

My results apply to my tasks. If your work is different, the test is quick to set up:

  1. Pick 5 to 8 real tasks you do every week, with made-up data. They should cover your languages, and at least one should be hard.
  2. Write the tests first and hide them. Validate them with your own solution. The model does not see them until it finishes.
  3. Include a contract task: a bug with a traceback where the obvious fix leaves another bug that only shows up when you read the documentation. It is the one that separates models the most, and it tells you whether the model reads or only reacts.
  4. A clean copy and auto mode per attempt, with a time limit and dangerous commands forbidden.
  5. Repeat at least once with the models you care about. A single run cannot tell luck from habit.
  6. Compare the cost with the real invoice, not with what the tool says. opencode overestimated four models for me, GLM-5.3 Flash at double the billed amount, because it used a catalog price and not the price of the provider that actually served the requests.
  7. Include a model you already know from practice. If your test does not reproduce what you lived with it, the test is wrong.

What I did not measure

Sources

Keep reading on IT Rafa

Leave a Reply

Your email address will not be published. Required fields are marked *