A Free Frontier Model Appeared With No Owner. The 80% Score Was Really 58%.
At 20:04 UTC on 20 August 2026, a model called stealth/ox-alpha appeared in OpenRouter's catalogue with no company name attached. It was free. It read just over a million tokens of context. And within four days it had processed more tokens than most commercial models see in a month — from developers who could not name the company receiving their code.
Four days later, OpenCode's public telemetry showed roughly 26 trillion tokens, 327,000 unique users and 8,328,244 completed sessions on that one client alone, putting Ox Alpha second on OpenCode's own usage leaderboard behind DeepSeek V4 Flash and ahead of Xiaomi's MiMo-V2.5, as AI Insiders reported on 26 August. Not one of those sessions came with a vendor.
Two things about this story got much less attention than they deserved: the benchmark number that made it go viral was wrong, and the data terms that govern where your code goes contradict each other depending on which client you used. Here is what the record actually says.
What Ox Alpha is
A stealth model is a pre-release frontier system published anonymously through an API platform so the lab behind it can collect real usage before an official launch. OpenRouter has run fourteen of these since April 2025. Ox Alpha is the fourteenth.
The listing describes “a reasoning model designed for coding, sustained agentic work, and production workloads.” The published specs:
| Spec | Ox Alpha |
|---|---|
| Model ID | stealth/ox-alpha |
| Listed | 20 Aug 2026, 20:04:55 UTC |
| Price | $0.00 input / $0.00 output |
| Context window | 1,048,576 tokens |
| Max output | 131,072 tokens |
| Input types | Text, image, video (audio rejected) |
| Reasoning | Always on, cannot be disabled |
| Tool calling | Yes — but JSON output is not schema-enforced |
| Developer | Anonymous; unclaimed as of 26 Aug 2026 |
That schema-enforcement gap is the practical catch for anyone building agents. If your pipeline expects structured responses, a free frontier endpoint that returns almost-valid JSON is not a drop-in replacement for a paid one.
Distribution was unusually wide for day one: OpenRouter, OpenCode (advertising capacity of 100 trillion tokens a day, with rate limits described as near-unlimited), Cline, and Nous Research's portal, which claimed capacity for a quadrillion tokens.
View Nous Research on appz.com
The 80% that was 58%
The number that made Ox Alpha a phenomenon came from developer Ben Davis, who ran ten tasks sampled from DeepSWE — a 113-task long-horizon software-engineering benchmark drawn from 91 repositories — and scored Ox Alpha at roughly 80% against Claude Fable 5 at 65% and GPT-5.6 Sol at 52%. The chart went around X within a day and is still circulating as evidence that an unknown model beat both frontier labs.
Davis then ran the full set. A reproducible 113-task run published on GitHub resolved 66 tasks, or 58.4%, over about 20 hours of agentic work. A separate full-set run landed near 63%. A community LiveCodeBench run with no tools and no agent harness reported 28% Pass@1. The New Stack and Decrypt both traced this arc independently.
| Run | Setup | Ox Alpha score |
|---|---|---|
| Ben Davis, 10-task DeepSWE sample | Agent harness, hand-selected subset | ~80% (8/10, one near-miss counted) |
| Full 113-task DeepSWE, published on GitHub | Agent harness, ~20h | 58.4% (66/113) |
| Second full-set run (Wenqi & Kevin) | Agent harness | ~63% |
| Community LiveCodeBench | No tools, no harness | 28% Pass@1 |
| Official DeepSWE leaderboard | — | Not listed |
Eighty, sixty-three, fifty-eight, twenty-eight. Those are not four measurements of the same thing. They are four different benchmarks-plus-harnesses, and the spread between them is wider than the gap between any two frontier models released this year. As of 22 August, Ox Alpha had no entry on Artificial Analysis or Arena at all.
The honest read: Ox Alpha performs somewhere around GPT-5.6 Sol on long-horizon agentic coding. That is genuinely impressive for a free endpoint. It is not the “beats Claude Fable and GPT-5.6” story that travelled, and the difference between those two claims is entirely a matter of sample size and harness.
The transferable lesson for anyone evaluating models: a ten-task slice of a long-horizon agentic suite is not a benchmark, it is an anecdote with a percentage sign. When you see a model chart on X, check the denominator before you switch your stack.
View Claude Fable 5 on appz.com
The identity hunt has been more rigorous than the benchmarking
Ironically, the amateur detective work on who built it has been the most methodologically careful part of this story. A developer known as unclecode — author of the Crawl4AI open-source crawler — built a tool called modelprint that fires nine infrastructure probes at an anonymous API endpoint and matches the responses against known models.
Ox Alpha matched GLM-5.3 on six of nine probes, including all four normalised tokenizer counts. No other lab's best candidate cleared more than two of those four. The circumstantial case is consistent: Z.ai previewed GLM-5 on OpenRouter under the codename Pony Alpha, and GLM-5.3 shipped on 14 August, six days before Ox Alpha appeared.
unclecode is careful about what that proves, and so are we: matching fingerprints establish shared serving infrastructure, not model identity. One lab can serve two different checkpoints from the same stack. A minority reading of the tokenizer behaviour points at cl100k_base, an OpenAI encoding that sits oddly on a Chinese model, and Xiaomi's MiMo team keeps coming up as an alternative. No statement exists from Zhipu, Z.ai, Xiaomi or OpenRouter. Everything in this section is community inference.
Precedent is on the side of eventual disclosure. Of the thirteen historical stealth slots, seven were officially revealed — four by OpenRouter and three by the lab itself.
| Stealth codename | Revealed as | When |
|---|---|---|
| Quasar Alpha, Optimus Alpha | GPT-4.1 | April 2025 |
| Sonoma Sky Alpha, Sonoma Dusk Alpha | Grok 4 Fast (community inference) | September 2025 |
| Pony Alpha | GLM-5 | February 2026 |
| Hunter Alpha, Healer Alpha | Xiaomi MiMo-V2 | March 2026 |
| Elephant Alpha | Ling-2.6-flash | April 2026 |
| Owl Alpha | LongCat-2.0 (Meituan) | June 2026 |
| Ox Alpha | Unrevealed | Listed 20 Aug 2026 |
The part nobody read: whose terms apply
This is the question that actually matters if you pointed a coding agent at Ox Alpha and let it read your repository.
OpenRouter's Stealth Program has default terms that grant the operator a perpetual, sublicensable licence over what you send, for the stated purpose of training the stealth model. Thirteen of the fourteen stealth listings to date carried a variant of “logged by the provider and may be used to improve the model.”
Ox Alpha's model page is the exception: it says prompts and completions are retained by the provider and are not used for training — and OpenRouter's own announcement framed that as unusual, with the phrase “this time.” Meanwhile OpenCode's launch post advertised the route as Zero Data Retention, from a provider it does not name.
| Route | What it says about your data | Who is accountable |
|---|---|---|
| OpenRouter model page | Retained by the provider; not used for training | Anonymous third party |
| OpenRouter Stealth Program defaults | Perpetual, sublicensable licence for training, evaluation and improvement | Anonymous third party |
| OpenCode route | “Zero Data Retention” | Unnamed provider |
“Retained but not trained on” and “zero retention” are not the same promise. When a reseller's marketing copy and the platform's own record disagree, the record wins — and OpenRouter is explicit that it is not the developer, owner or operator of Ox Alpha and only routes requests to it. So the entity making the retention promise, the entity holding your code, and the entity you could complain to if it were breached are all the same unnamed party.
There is no accusation here. The likeliest explanation remains the boring one: a lab running a legitimate pre-release preview with better-than-usual terms. But “we promise not to train on it” is only worth what the promiser is worth, and the promiser is currently a blank field.
What to do about it
The free window is closing. OpenRouter emailed users that free access on its route ran through Monday 24 August; OpenCode's launch note promised a week from 20 August, which points at roughly 27 August. If you want to test it, you have days, not weeks.
A reasonable posture:
- Test it on code you would publish. Open-source repos, scratch projects, throwaway benchmarks. Not your production monorepo, not customer data, not anything under an NDA.
- Read the route, not the client. The terms that bind you are the ones on the endpoint your request actually hits, which is rarely the ones in the launch tweet.
- Benchmark it yourself, on your tasks. The 22-point spread between published runs is the strongest argument that generic scores will not tell you whether it works for your codebase.
- Do not migrate a schema-dependent agent to it. Non-enforced JSON will bite you in production long before the free window closes.
The wider pattern is worth naming. Free frontier-class inference is now a customer-acquisition channel, and the currency is not money — it is evaluation signal, usage traces and, in thirteen of fourteen historical cases, training data. That is a real trade, and it can be a good one. It stops being a good one when you cannot name the counterparty.
Ox Alpha will almost certainly be revealed. Every stealth slot so far has been, eventually. Until then, the most useful thing about it is not the score — it is the reminder that in 2026, the two hardest questions about any model are which benchmark produced that number and which legal entity received your request. Neither one is on the model card.