
What Is Ox Alpha? The 1M-Context Stealth Model With Video Input
Ox Alpha was the anonymous 1M-context model on OpenRouter in August 2026. Z.ai later revealed it as GLM-5.3-Flash. Specs, pricing, benchmarks and API access.
Ox Alpha was an anonymous "stealth" AI model that appeared on OpenRouter on August 20, 2026, with a 1,048,576-token context window and text, image and video input.[1] For six days nobody would say who built it. On August 26, Z.ai ended the guessing: Ox Alpha was an unreleased build of GLM-5.3-Flash, tested in public before launch.[3]
So the honest answer to "what is Ox Alpha" has changed since most explainers were written. It is no longer a free mystery model. It is a released, open-weight model with a price list, a model card and benchmark numbers. This guide covers what Ox Alpha was during the preview, what it turned out to be, and what you actually call today if you want the same model.
TL;DR
- Launch: Ox Alpha went live on OpenRouter on August 20, 2026 as
stealth/ox-alpha, listed as a reasoning model for coding, agentic work and production workloads.[1] - Identity: Z.ai confirmed on August 26 that it had tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter.[3]
- Size: 320B total parameters, 18B active, with a hybrid of sparse and linear attention.[3]
- Inputs: video, image, text and file in; text out; 1M context and 128K maximum output on Z.ai's API.[4]
- Price now: Z.ai lists GLM-5.3-Flash at $0.15 input, $0.03 cached input and $0.50 output per million tokens.[5]
- Weights: published on Hugging Face under the MIT license.[6]
What Ox Alpha was during the stealth preview
OpenRouter described Ox Alpha as "a reasoning model designed for coding, sustained agentic work, and production workloads," suited to long-horizon software engineering and to workflows that mix text with visual context.[1] The provider field read "stealth." The listing said the model was run by a third party that had chosen to stay anonymous, and that OpenRouter only routed the requests.
The spec sheet was unusual for a free listing. OpenRouter's own FAQ for the page gives a 1,048,576-token context window and says Ox Alpha "accepts text, images, and video as input and returns text."[1] Video input is still rare among coding-focused models, and a 1M window on a free endpoint got attention fast.
TechCrunch covered the frenzy on August 23. Its brief notes that the model was free on OpenRouter, that Stripe CEO Patrick Collison called it "very impressive," and that the guesses were all over the place.[2] Early speculation pointed at Z.ai's GLM family. A Wccftech update floated an unreleased Microsoft MAI model instead, and Reddit threads argued both ways. In other words, the community fingerprinting was pointing in the right direction, but nobody could prove it.
One detail from the preview still matters. OpenRouter's notice says prompts and completions sent to Ox Alpha "were retained by the provider and are not used for training."[1] If you pasted private code into Ox Alpha in August, that data went to Z.ai.
Ox Alpha was GLM-5.3-Flash
The reveal came in Z.ai's launch post for GLM-5.3-Flash, dated August 26, 2026: "Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips."[3]
OpenRouter updated the stealth page the same way. It now reads: "This stealth
model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash,"
and points users to the z-ai/glm-5.3-flash listing.[1]
That listing shows a release date of August 26, 2026.[7]
GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series. Z.ai says it starts from a newly trained base model rather than a fine-tune of an earlier GLM, and was pre-trained on a 30T-token multimodal corpus.[3]
Ox Alpha specs at a glance
| Attribute | Ox Alpha / GLM-5.3-Flash |
|---|---|
| Developer | Z.ai (confirmed August 26, 2026) |
| Stealth ID | stealth/ox-alpha (retired) |
| API model code | glm-5.3-flash, plus the faster glm-5.3-flashx |
| Parameters | 320B total, 18B active |
| Layers | 45 |
| Input | Video, image, text, file |
| Output | Text |
| Context window | 1M tokens on Z.ai's API |
| Maximum output | 128K tokens |
| Thinking | Always on; reasoning_effort of low, high or max |
| License | MIT, weights on Hugging Face |
Sources: Z.ai launch post and model guide, plus the Hugging Face model card.[3][4][6] OpenRouter's GLM-5.3-Flash listing and Z.ai's own guide agree on the context window: OpenRouter shows 1,048,576 tokens, which is the 1M Z.ai publishes.[7][4] If you plan to fill the window, budget for 1M.
How Ox Alpha got so much context for so little compute
The architecture explains both the 1M window and the low price. Compared with GLM-4.5, GLM-5.3-Flash has a similar total size (320B versus 355B) but nearly half the active parameters (18B versus 32B) and under half the layers (45 versus 92).[3]
The bigger change is attention. Z.ai combines linear attention, which tracks local dependencies through a running state, with sparse attention, which uses a lightweight indexer to pull in relevant tokens from far back in the context. A technique Z.ai calls IndexPool compresses four indexer key vectors into one to keep that indexer cheap at 1M tokens. Against GLM-5.3, Z.ai reports 3.0× less attention compute and a 4.4× smaller KV cache.[3]
That is also why the preview could run on Chinese AI chips. Z.ai says those chips are mainly limited by memory capacity and bandwidth, and that it reached a 3× improvement in end-to-end serving performance over its first baseline on the same hardware.[3] A smaller KV cache is exactly what a memory-bound chip needs.
Ox Alpha benchmarks, now that the numbers are official
During the preview, every Ox Alpha benchmark was a screenshot from a stranger. Now there is an official table. These are Z.ai's own numbers, not independent audits, so read them as the vendor's claim.[3]
| Benchmark | GLM-5.3-Flash | GLM-5.2 | Opus 4.8 |
|---|---|---|---|
| Terminal Bench 2.1 | 84.3 | 81.0 | 85.0 |
| DeepSWE v1.1 | 63.4 | 46.2 | 58.0 |
| NL2Repo | 56.3 | 48.9 | 69.7 |
| Toolathlon Verified | 78.4 | 59.9 | 76.2 |
| AutomationBench v1.0.6 | 48.8 | 26.2 | 41.0 |
| HLE w/ Tools | 55.3 | 54.7 | 57.9 |
| OSWorld 2.0 | 59.1 | n/a | 54.8 |
Two things stand out. First, the jump over GLM-5.2 on agentic work is large: AutomationBench nearly doubles. Second, the model is not uniformly ahead of Claude Opus 4.8. It trails clearly on NL2Repo, which asks for a whole repository from a natural-language spec. Z.ai's summary, "approaching Claude Opus 4.8," is fair. "Beats Opus" would not be.
Z.ai also reports a score of 57 on the Artificial Analysis Intelligence Index v4.1.1 at $0.045 per task (discounted).[3]
Ox Alpha pricing after the free preview
The free window is over. Here is what the same model costs on Z.ai's API, per million tokens:[5]
| Model | Input | Cached input | Output |
|---|---|---|---|
| GLM-5.3-Flash | $0.15 | $0.03 | $0.50 |
| GLM-5.3-FlashX | $0.37 | $0.075 | $1.25 |
| GLM-5.3 | $1.40 | $0.26 | $4.40 |
| GLM-5.2 | $1.40 | $0.26 | $4.40 |
That is where the "one-tenth the price" line in Z.ai's launch post comes from: $0.50 output against $4.40 for GLM-5.2 is roughly a ninth.[3] FlashX is the same model served for speed. Z.ai quotes up to 200 tokens per second for it.[4]
Because the weights are open, other hosts serve it too. OpenRouter lists more
than twenty providers for z-ai/glm-5.3-flash, and Z.ai's own endpoint there
matches the $0.15 / $0.50 list price.[7]
Calling the model Ox Alpha became
stealth/ox-alpha is retired, so do not hard-code it. Call glm-5.3-flash
on Z.ai's Chat Completions API, or z-ai/glm-5.3-flash on OpenRouter.[4][7]
A few behaviors trip people up when they move over from a text-only model:
- Thinking cannot be turned off.
thinking.typeonly acceptsenabled. Control cost withreasoning_effort, which takeslow,highormaxand defaults tomaxwhen omitted.[4][6] - Recommended sampling:
temperature: 1andtop_p: 0.95. For streaming, Z.ai suggests enabling bothstreamandtool_stream.[4] - Images go in
image_urlcontent blocks. Add several blocks for several images. Z.ai recommends passing a URL.[4] - Self-hosted chat apps should set
clear_thinking. In the open-weight chat template it defaults tofalse, and the model card says to passtruefor chat use.[6]
A minimal request looks like this:
curl https://api.z.ai/api/paas/v4/chat/completions \
-H "Authorization: Bearer $ZAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.3-flash",
"reasoning_effort": "high",
"messages": [{
"role": "user",
"content": [
{"type": "image_url", "image_url": {"url": "https://example.com/ui.png"}},
{"type": "text", "text": "Find the layout bugs in this screenshot."}
]
}]
}'If you self-host, the model card lists SGLang, vLLM, Transformers, KTransformers and Unsloth among the supported paths.[6]
FAQ
What is Ox Alpha?
Ox Alpha was the anonymous name for Z.ai's GLM-5.3-Flash during a public preview on OpenRouter and OpenCode that began on August 20, 2026. It is a multimodal reasoning model for coding and agent work with a 1M-token context window.
Who made Ox Alpha?
Z.ai. The company confirmed it in the GLM-5.3-Flash launch post on August 26, 2026, and OpenRouter's stealth page now says the model was "developed and operated by ZAI."
Is Ox Alpha still free?
No. The stealth preview has ended. The model now ships as GLM-5.3-Flash at $0.15 input and $0.50 output per million tokens on Z.ai's API.
Is Ox Alpha open source?
Yes, as GLM-5.3-Flash. The weights are on Hugging Face under the MIT license. During the preview itself, no weights were available.
What is Ox Alpha's context window?
OpenRouter listed 1,048,576 tokens for Ox Alpha. Z.ai's guide for GLM-5.3-Flash gives a 1M context and 128K maximum output.
Can Ox Alpha read video?
Yes. Z.ai lists video, image, text and file as input modalities for GLM-5.3-Flash. Output is text only.
Was Ox Alpha a Gemini or Microsoft model?
No. Those were community guesses reported during the preview. Z.ai and OpenRouter both identify the model as GLM-5.3-Flash.
Is Ox Alpha better than GLM-5.2?
On Z.ai's published benchmarks, yes, by a wide margin on agentic tasks. It also costs about a ninth as much on output. The one reason to stay on GLM-5.2 is if you have already tuned prompts around its behavior and do not need vision.
Where Ox Alpha fits in a GLM stack
Ox Alpha was a marketing experiment that worked: an anonymous model became the most popular listing of its week, and the reveal landed with a price list already attached. For builders, the practical takeaway is simpler. The model behind Ox Alpha is a cheap, open-weight, 1M-context coder that can look at screenshots and video, and the free stealth endpoint is gone.
If you already run GLM-5.2 through reAPI, our GLM-5.2 API guide covers the request shape, reasoning controls and limits, and the GLM-5.2 model page has current rates. For other long-context options, see our Kimi K3 guide and DeepSeek V4 1M-context guide. Just remember that searching for Ox Alpha today means searching for GLM-5.3-Flash.
References
- OpenRouter. Ox Alpha (stealth/ox-alpha): listing, reveal notice and FAQ. Retrieved September 2026 from openrouter.ai/stealth/ox-alpha
- TechCrunch. Who's behind the new 'stealth model' Ox Alpha? August 23, 2026. Retrieved September 2026 from techcrunch.com/2026/08/23/whos-behind-the-new-stealth-model-ox-alpha
- Z.ai. GLM-5.3-Flash: Frontier Intelligence, Flash Cost. August 26, 2026. Retrieved September 2026 from z.ai/blog/glm-5.3-flash
- Z.ai. GLM-5.3-Flash/FlashX model guide. Retrieved September 2026 from docs.z.ai/guides/llm/glm-5.3-flash
- Z.ai. Pricing. Retrieved September 2026 from docs.z.ai/guides/overview/pricing
- Z.ai. zai-org/GLM-5.3-Flash model card. Hugging Face. Retrieved September 2026 from huggingface.co/zai-org/GLM-5.3-Flash
- OpenRouter. Z.ai: GLM 5.3 Flash (z-ai/glm-5.3-flash). Retrieved September 2026 from openrouter.ai/z-ai/glm-5.3-flash
Further reading
- reAPI. GLM-5.2 API guide. reapi.ai/blog/glm-5-2-api-guide
- reAPI. GLM-5.2 model page. reapi.ai/models/glm-5-2
Author

Categories
More Posts

Sora Alternatives After Shutdown: Tools and APIs Compared
Choose a Sora replacement for browser creation or API workflows. Compare ClipDance, Runway, Seedance, Kling, Veo and Wan by task, duration and cost.


GPT Image 2.5 Model Not Found: Fix Codex and API Access
Fix GPT Image 2.5 model not found errors by checking model IDs, Image API versus Responses fields, Codex tools and account access before retrying requests.


AI Image API Content Filters: How Refusals Happen
Image API filtering runs in several layers. The same prompt can pass one host and fail another, but certain boundaries never change regardless of configuration.
