GLM-5.3-Flash doubles on September 9
1 hour ago
The launch discount ends on September ninth. Then the bill doubles. GLM-5.3-Flash is the model on that clock. GLM-5.3-Flash is Z.ai’s cheaper, faster sibling of GLM-5.3: three hundred and twenty billion in the file, eighteen billion awake per word, a million-word window. Z.ai’s live grant is seven and a half cents per million input tokens, the text you send, one and a half cents per million cached input tokens, and twenty-five cents per million output tokens, the text the model writes back. The crossed-out number is the real menu price. The live number is a fifty percent off grant with an end date. List is fifteen cents in, three cents cached, and fifty cents out, per million tokens. Z.ai times that grant at midnight on September ninth, Singapore time. OpenRouter times the same instant at four in the afternoon UTC. Your API bill can double on a calendar date.
Ask
Ask about this presentation
Answers are generated from this presentation.
Chapters
Show transcriptHide transcript
GLM-5.3-Flash doubles on September 9
The launch discount ends on September ninth. Then the bill doubles. GLM-5.3-Flash is the model on that clock. GLM-5.3-Flash is Z.ai’s cheaper, faster sibling of GLM-5.3: three hundred and twenty billion in the file, eighteen billion awake per word, a million-word window. Z.ai’s live grant is seven and a half cents per million input tokens, the text you send, one and a half cents per million cached input tokens, and twenty-five cents per million output tokens, the text the model writes back. The crossed-out number is the real menu price. The live number is a fifty percent off grant with an end date. List is fifteen cents in, three cents cached, and fifty cents out, per million tokens. Z.ai times that grant at midnight on September ninth, Singapore time. OpenRouter times the same instant at four in the afternoon UTC. Your API bill can double on a calendar date.
GLM-5.3-Flash lists at $0.15 in and $0.50 out
You can hand GLM-5.3-Flash about a million tokens at once. Z.ai caps the reply at a hundred and twenty-eight thousand tokens. Z.ai lists GLM-5.3-Flash at fifteen cents per million input tokens and fifty cents per million output tokens. Artificial Analysis clocks GLM-5.3-Flash on Z.ai’s API at forty-seven point six tokens a second. That is how fast the words come out. The wait before the first word is one point six four seconds. The LICENSE file is MIT, copyright two thousand twenty-six, Z.AI. MIT lets you use the weights, including commercially, with almost no extra gates.
Intelligence Index v4.2 scores 46 at $0.18 per task
Artificial Analysis scores GLM-5.3-Flash forty-six on Intelligence Index version 4.2. The Intelligence Index is a single score Artificial Analysis builds by running the same tests on every model itself. Cost per Index task is what Artificial Analysis billed to finish one of those tests. Artificial Analysis billed eighteen cents per Intelligence Index task. On August twenty-sixth, launch day, Artificial Analysis scored GLM-5.3-Flash fifty-seven on Intelligence Index version 4.1.1, at nine cents per Index task. Version 4.2 is a different mix of tests than version 4.1.1.
Artificial Analysis’s $0.10 blend is built at list
Z.ai is selling the same eighteen-billion-awake model at half list until a printed instant. After midnight Singapore time on September ninth, Z.ai’s list is fifteen cents in, three cents cached, and fifty cents out, per million tokens. Artificial Analysis’s ten cents blended per million tokens is built at that list. The task rate is a different number: eighteen cents per Index task, also built at list. A blended price is one dollar figure standing in for a mixed bill: seven parts prompt text the model has already seen, two parts new prompt, one part what it writes back. The mix uses seven parts cached at three cents, two parts new input at fifteen cents, and one part output at fifty cents. That mix is ten point one cents. Artificial Analysis rounds the mix to ten cents per million tokens. At the live grant the same mix uses one and a half cents per million cached tokens, seven and a half cents per million new input tokens, and twenty-five cents per million output tokens. That mix is about five cents per million tokens. A budget that survives past September ninth has to plan the ten cents per million tokens. Z.ai still prints cached input storage as Limited-time Free on the same page. That storage grant has no September ninth date.
Four OpenRouter hosts still show 50% off
Z.ai wrote the clock. OpenRouter’s banner matches four in the afternoon UTC on September ninth. OpenRouter lists twenty-four hosts for GLM-5.3-Flash today. Four hosts still show fifty percent off: Z.ai, Novita, DeepInfra, and GMICloud, at seven and a half cents in and twenty-five cents out, per million tokens. Sixteen hosts already bill list: fifteen cents in and fifty cents out, per million tokens. CoreWeave, Fireworks, and Together sit in that already-list group. CoreWeave’s cached input is five cents per million tokens. Z.ai’s cached list is three cents per million tokens. Wafer prints a different sticker: ten cents in and thirty-five cents out, per million tokens, with no promo badge. Modal takes sixty-seven percent off a forty-five cent in and a dollar fifty out list, and lands at fifteen cents in and fifty cents out. A budget that routes through Z.ai’s API pays double after that clock. A budget that routes through CoreWeave, Fireworks, or Together already pays list today.
GLM-5.3-Flash at 46 index points and $0.10 blended
GLM-5.3-Flash sits at forty-six index points, ten cents blended per million tokens, and eighteen cents per Index task, all at list. Qwen3.8-Flash-Next sits at the same forty-six index points, marked as an Artificial Analysis estimate, at nine cents blended per million tokens. Gemini 3.8 Flash on high scores forty-seven index points at fifty-eight cents blended per million tokens. The task rate for Gemini 3.8 Flash on high is a different number: seventy-four cents per Index task. GLM-5.3 on max scores forty-nine index points at ninety cents blended per million tokens, and a dollar twenty-six per Index task.
Past September 9 on Z.ai, plan $0.15 in and $0.50 out
If the bill has to survive past September ninth on Z.ai, plan fifteen cents in and fifty cents out per million tokens, Index forty-six, eighteen cents per Index task, forty-seven point six tokens a second. If you are still on the promo this week, the live invoice is seven and a half cents in and twenty-five cents out, per million tokens, and the clock is midnight Singapore time on September ninth. If you pick a host by sticker, OpenRouter already prints the split. Wafer bills ten cents in and thirty-five cents out, per million tokens. The MIT weights stay MIT if you run the files yourself. The doubling is an API event.
At midnight Singapore time the strikethrough becomes the list
The promo bars from the open become Z.ai’s list at midnight Singapore time on September ninth. The model stays the same. The invoice doubles. After four in the afternoon UTC that day, the next reading is whether Novita, DeepInfra, and GMICloud follow Z.ai to list.
Up next
5:48Gemini 3.8 Flash costs about forty percent more per measured task at the same token price
@model-watch · 1 day ago
5:41The cheap tier caught the frontier: GLM-5.3-Flash scores 57 at $0.10
@model-watch · 3 days ago
2:04Anthropic cancels Sonnet 5's price increase — and the week's biggest release has no spec sheet
@model-watch · 1 week ago
2:04Anthropic cancels Sonnet 5's price increase — and the week's biggest release has no spec sheet
@model-watch · 1 week ago