GPT Image 2.5 is live — OpenAI's newest image model, targeted edits that leave the rest of the frame alone
What Is Ox Alpha? The 1M-Context Stealth Model With Video Input
2026/09/29

What Is Ox Alpha? The 1M-Context Stealth Model With Video Input

Ox Alpha was the anonymous 1M-context model on OpenRouter in August 2026. Z.ai later revealed it as GLM-5.3-Flash. Specs, pricing, benchmarks and API access.

Ox Alpha was an anonymous "stealth" AI model that appeared on OpenRouter on August 20, 2026, with a 1,048,576-token context window and text, image and video input.[1] For six days nobody would say who built it. On August 26, Z.ai ended the guessing: Ox Alpha was an unreleased build of GLM-5.3-Flash, tested in public before launch.[3]

So the honest answer to "what is Ox Alpha" has changed since most explainers were written. It is no longer a free mystery model. It is a released, open-weight model with a price list, a model card and benchmark numbers. This guide covers what Ox Alpha was during the preview, what it turned out to be, and what you actually call today if you want the same model.

TL;DR

  • Launch: Ox Alpha went live on OpenRouter on August 20, 2026 as stealth/ox-alpha, listed as a reasoning model for coding, agentic work and production workloads.[1]
  • Identity: Z.ai confirmed on August 26 that it had tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter.[3]
  • Size: 320B total parameters, 18B active, with a hybrid of sparse and linear attention.[3]
  • Inputs: video, image, text and file in; text out; 1M context and 128K maximum output on Z.ai's API.[4]
  • Price now: Z.ai lists GLM-5.3-Flash at $0.15 input, $0.03 cached input and $0.50 output per million tokens.[5]
  • Weights: published on Hugging Face under the MIT license.[6]

What Ox Alpha was during the stealth preview

OpenRouter described Ox Alpha as "a reasoning model designed for coding, sustained agentic work, and production workloads," suited to long-horizon software engineering and to workflows that mix text with visual context.[1] The provider field read "stealth." The listing said the model was run by a third party that had chosen to stay anonymous, and that OpenRouter only routed the requests.

The spec sheet was unusual for a free listing. OpenRouter's own FAQ for the page gives a 1,048,576-token context window and says Ox Alpha "accepts text, images, and video as input and returns text."[1] Video input is still rare among coding-focused models, and a 1M window on a free endpoint got attention fast.

TechCrunch covered the frenzy on August 23. Its brief notes that the model was free on OpenRouter, that Stripe CEO Patrick Collison called it "very impressive," and that the guesses were all over the place.[2] Early speculation pointed at Z.ai's GLM family. A Wccftech update floated an unreleased Microsoft MAI model instead, and Reddit threads argued both ways. In other words, the community fingerprinting was pointing in the right direction, but nobody could prove it.

One detail from the preview still matters. OpenRouter's notice says prompts and completions sent to Ox Alpha "were retained by the provider and are not used for training."[1] If you pasted private code into Ox Alpha in August, that data went to Z.ai.

Ox Alpha was GLM-5.3-Flash

The reveal came in Z.ai's launch post for GLM-5.3-Flash, dated August 26, 2026: "Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips."[3]

OpenRouter updated the stealth page the same way. It now reads: "This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash," and points users to the z-ai/glm-5.3-flash listing.[1] That listing shows a release date of August 26, 2026.[7]

GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series. Z.ai says it starts from a newly trained base model rather than a fine-tune of an earlier GLM, and was pre-trained on a 30T-token multimodal corpus.[3]

Ox Alpha specs at a glance

AttributeOx Alpha / GLM-5.3-Flash
DeveloperZ.ai (confirmed August 26, 2026)
Stealth IDstealth/ox-alpha (retired)
API model codeglm-5.3-flash, plus the faster glm-5.3-flashx
Parameters320B total, 18B active
Layers45
InputVideo, image, text, file
OutputText
Context window1M tokens on Z.ai's API
Maximum output128K tokens
ThinkingAlways on; reasoning_effort of low, high or max
LicenseMIT, weights on Hugging Face

Sources: Z.ai launch post and model guide, plus the Hugging Face model card.[3][4][6] OpenRouter's GLM-5.3-Flash listing and Z.ai's own guide agree on the context window: OpenRouter shows 1,048,576 tokens, which is the 1M Z.ai publishes.[7][4] If you plan to fill the window, budget for 1M.

How Ox Alpha got so much context for so little compute

The architecture explains both the 1M window and the low price. Compared with GLM-4.5, GLM-5.3-Flash has a similar total size (320B versus 355B) but nearly half the active parameters (18B versus 32B) and under half the layers (45 versus 92).[3]

The bigger change is attention. Z.ai combines linear attention, which tracks local dependencies through a running state, with sparse attention, which uses a lightweight indexer to pull in relevant tokens from far back in the context. A technique Z.ai calls IndexPool compresses four indexer key vectors into one to keep that indexer cheap at 1M tokens. Against GLM-5.3, Z.ai reports 3.0× less attention compute and a 4.4× smaller KV cache.[3]

That is also why the preview could run on Chinese AI chips. Z.ai says those chips are mainly limited by memory capacity and bandwidth, and that it reached a 3× improvement in end-to-end serving performance over its first baseline on the same hardware.[3] A smaller KV cache is exactly what a memory-bound chip needs.

Ox Alpha benchmarks, now that the numbers are official

During the preview, every Ox Alpha benchmark was a screenshot from a stranger. Now there is an official table. These are Z.ai's own numbers, not independent audits, so read them as the vendor's claim.[3]

BenchmarkGLM-5.3-FlashGLM-5.2Opus 4.8
Terminal Bench 2.184.381.085.0
DeepSWE v1.163.446.258.0
NL2Repo56.348.969.7
Toolathlon Verified78.459.976.2
AutomationBench v1.0.648.826.241.0
HLE w/ Tools55.354.757.9
OSWorld 2.059.1n/a54.8

Two things stand out. First, the jump over GLM-5.2 on agentic work is large: AutomationBench nearly doubles. Second, the model is not uniformly ahead of Claude Opus 4.8. It trails clearly on NL2Repo, which asks for a whole repository from a natural-language spec. Z.ai's summary, "approaching Claude Opus 4.8," is fair. "Beats Opus" would not be.

Z.ai also reports a score of 57 on the Artificial Analysis Intelligence Index v4.1.1 at $0.045 per task (discounted).[3]

Ox Alpha pricing after the free preview

The free window is over. Here is what the same model costs on Z.ai's API, per million tokens:[5]

ModelInputCached inputOutput
GLM-5.3-Flash$0.15$0.03$0.50
GLM-5.3-FlashX$0.37$0.075$1.25
GLM-5.3$1.40$0.26$4.40
GLM-5.2$1.40$0.26$4.40

That is where the "one-tenth the price" line in Z.ai's launch post comes from: $0.50 output against $4.40 for GLM-5.2 is roughly a ninth.[3] FlashX is the same model served for speed. Z.ai quotes up to 200 tokens per second for it.[4]

Because the weights are open, other hosts serve it too. OpenRouter lists more than twenty providers for z-ai/glm-5.3-flash, and Z.ai's own endpoint there matches the $0.15 / $0.50 list price.[7]

Calling the model Ox Alpha became

stealth/ox-alpha is retired, so do not hard-code it. Call glm-5.3-flash on Z.ai's Chat Completions API, or z-ai/glm-5.3-flash on OpenRouter.[4][7] A few behaviors trip people up when they move over from a text-only model:

  • Thinking cannot be turned off. thinking.type only accepts enabled. Control cost with reasoning_effort, which takes low, high or max and defaults to max when omitted.[4][6]
  • Recommended sampling: temperature: 1 and top_p: 0.95. For streaming, Z.ai suggests enabling both stream and tool_stream.[4]
  • Images go in image_url content blocks. Add several blocks for several images. Z.ai recommends passing a URL.[4]
  • Self-hosted chat apps should set clear_thinking. In the open-weight chat template it defaults to false, and the model card says to pass true for chat use.[6]

A minimal request looks like this:

curl https://api.z.ai/api/paas/v4/chat/completions \
  -H "Authorization: Bearer $ZAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "glm-5.3-flash",
    "reasoning_effort": "high",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "image_url", "image_url": {"url": "https://example.com/ui.png"}},
        {"type": "text", "text": "Find the layout bugs in this screenshot."}
      ]
    }]
  }'

If you self-host, the model card lists SGLang, vLLM, Transformers, KTransformers and Unsloth among the supported paths.[6]

FAQ

What is Ox Alpha?

Ox Alpha was the anonymous name for Z.ai's GLM-5.3-Flash during a public preview on OpenRouter and OpenCode that began on August 20, 2026. It is a multimodal reasoning model for coding and agent work with a 1M-token context window.

Who made Ox Alpha?

Z.ai. The company confirmed it in the GLM-5.3-Flash launch post on August 26, 2026, and OpenRouter's stealth page now says the model was "developed and operated by ZAI."

Is Ox Alpha still free?

No. The stealth preview has ended. The model now ships as GLM-5.3-Flash at $0.15 input and $0.50 output per million tokens on Z.ai's API.

Is Ox Alpha open source?

Yes, as GLM-5.3-Flash. The weights are on Hugging Face under the MIT license. During the preview itself, no weights were available.

What is Ox Alpha's context window?

OpenRouter listed 1,048,576 tokens for Ox Alpha. Z.ai's guide for GLM-5.3-Flash gives a 1M context and 128K maximum output.

Can Ox Alpha read video?

Yes. Z.ai lists video, image, text and file as input modalities for GLM-5.3-Flash. Output is text only.

Was Ox Alpha a Gemini or Microsoft model?

No. Those were community guesses reported during the preview. Z.ai and OpenRouter both identify the model as GLM-5.3-Flash.

Is Ox Alpha better than GLM-5.2?

On Z.ai's published benchmarks, yes, by a wide margin on agentic tasks. It also costs about a ninth as much on output. The one reason to stay on GLM-5.2 is if you have already tuned prompts around its behavior and do not need vision.

Where Ox Alpha fits in a GLM stack

Ox Alpha was a marketing experiment that worked: an anonymous model became the most popular listing of its week, and the reveal landed with a price list already attached. For builders, the practical takeaway is simpler. The model behind Ox Alpha is a cheap, open-weight, 1M-context coder that can look at screenshots and video, and the free stealth endpoint is gone.

If you already run GLM-5.2 through reAPI, our GLM-5.2 API guide covers the request shape, reasoning controls and limits, and the GLM-5.2 model page has current rates. For other long-context options, see our Kimi K3 guide and DeepSeek V4 1M-context guide. Just remember that searching for Ox Alpha today means searching for GLM-5.3-Flash.

References

  1. OpenRouter. Ox Alpha (stealth/ox-alpha): listing, reveal notice and FAQ. Retrieved September 2026 from openrouter.ai/stealth/ox-alpha
  2. TechCrunch. Who's behind the new 'stealth model' Ox Alpha? August 23, 2026. Retrieved September 2026 from techcrunch.com/2026/08/23/whos-behind-the-new-stealth-model-ox-alpha
  3. Z.ai. GLM-5.3-Flash: Frontier Intelligence, Flash Cost. August 26, 2026. Retrieved September 2026 from z.ai/blog/glm-5.3-flash
  4. Z.ai. GLM-5.3-Flash/FlashX model guide. Retrieved September 2026 from docs.z.ai/guides/llm/glm-5.3-flash
  5. Z.ai. Pricing. Retrieved September 2026 from docs.z.ai/guides/overview/pricing
  6. Z.ai. zai-org/GLM-5.3-Flash model card. Hugging Face. Retrieved September 2026 from huggingface.co/zai-org/GLM-5.3-Flash
  7. OpenRouter. Z.ai: GLM 5.3 Flash (z-ai/glm-5.3-flash). Retrieved September 2026 from openrouter.ai/z-ai/glm-5.3-flash

Further reading