The cheap tier caught the frontier: GLM-5.3-Flash scores 57 at $0.10
16 hours ago
Two models, one lab, eight days apart. One costs nine times as much to run. The cheaper one is called GLM-5.3-Flash, and it arrived on August twenty-sixth. Artificial Analysis scores it fifty-seven on their Intelligence Index, a score they build by running the same tests on every model themselves. The larger GLM-5.3 scores sixty on that same index. On this chart the two sit at ten cents and ninety cents per million tokens. Those are blended prices: one dollar figure standing in for a mixed bill of the text you send and the text the model writes back. That is three points of score for about nine times the price. Three releases landed on the board in fourteen days.
Ask
Ask about this presentation
Answers are generated from this presentation.
Chapters
Show transcriptHide transcript
Three points, nine times the price
Two models, one lab, eight days apart. One costs nine times as much to run. The cheaper one is called GLM-5.3-Flash, and it arrived on August twenty-sixth. Artificial Analysis scores it fifty-seven on their Intelligence Index, a score they build by running the same tests on every model themselves. The larger GLM-5.3 scores sixty on that same index. On this chart the two sit at ten cents and ninety cents per million tokens. Those are blended prices: one dollar figure standing in for a mixed bill of the text you send and the text the model writes back. That is three points of score for about nine times the price. Three releases landed on the board in fourteen days.
GLM-5.3-Flash spec card
Z.ai’s own list prices GLM-5.3-Flash in each direction: fifteen cents per million input tokens, the text you send it, and fifty cents per million output tokens, the text it writes back. You can hand it about one point three million tokens in one go. Artificial Analysis clocks it writing about forty-two tokens a second, which is how fast the words come out once it starts, and waiting about one point seven seconds before that first word appears. Its license field reads MIT: use it commercially, no permission needed.
Qwen3.8 spec card
Qwen3.8 landed on August twelfth. Alibaba lists it at two dollars per million input tokens and six dollars per million output tokens. Its license field reads qwen3.8-max, a custom license Alibaba wrote itself. The smaller model in the same family reads apache-2.0, a standard open license.
Gemini 3.7 Flash spec card
Gemini 3.7 Flash arrived the next day, on August thirteenth. Google lists it at seventy-five cents per million input tokens and three dollars seventy-five per million output tokens. Google does not publish the model itself, so Gemini has no license field to read. And that seventy-five cent price has a date attached.
Gemini’s 2027 price doubles
Google publishes two prices for Gemini 3.7 Flash side by side. Through December thirty-first, twenty twenty-six, input is seventy-five cents per million tokens and output is three dollars seventy-five. From January first, twenty twenty-seven, input becomes a dollar fifty and output becomes seven dollars fifty. If you are modelling a Gemini Flash bill for next January, model twice the invoice you pay now.
The half price ends September 9
GLM-5.3-Flash carries the other dated number. Its seven and a half cents per million input tokens and twenty-five cents per million output tokens is a launch promotion, and Z.ai’s page says the promotion ends at midnight on September ninth, eight days from today. Budget on the list price behind it: fifteen cents in, fifty cents out, per million tokens.
18 billion of 320 billion
Two things make GLM-5.3-Flash cheap. The first one is what wakes up inside the model. GLM-5.3-Flash holds three hundred and twenty billion adjustable numbers. For any single word it reads, it switches on eighteen billion of those numbers and leaves the rest asleep. Model makers call the ones that switch on the active parameters. Qwen3.8 holds two point four trillion adjustable numbers and switches on ninety-five billion of them per word, a little over five times as many awake as GLM-5.3-Flash uses. And Qwen3.8 lists at two dollars per million input tokens against GLM-5.3-Flash’s fifteen cents, which is more than thirteen times the input price. The serving cost follows the numbers that actually run each word. Hardware, batching, margin and a live promotion all sit on that price too, so the parameter count points the direction and stops there. It does mean a model with a bigger headline number can be the cheaper one to rent. Google publishes no parameter counts for Gemini 3.7 Flash, so Gemini stays off this comparison.
Seven of every ten tokens
The second thing that makes GLM-5.3-Flash cheap is what repeats. Take that ten cents again. Artificial Analysis mixes it at seven parts text the model has already read, two parts new prompt, and one part what the model writes back, and mixes it that way for every model on their board. Text you already sent is called cached input, and Z.ai prices cached input in the same row as the rest: three cents per million tokens, a fifth of the fifteen cents Z.ai charges for new input, so eighty percent off. Put those three published rates through the seven, two and one mix and you get a dollar and one cent across ten units. That is ten point one cents a unit, against the ten cents Artificial Analysis publishes. The seller’s own table rebuilds the board’s figure. So the shape of your prompt moves your invoice about as much as which model you picked, and a job that never sends the same instructions twice misses that discount entirely.
The whole board, six models
Six models sit on this board. GLM-5.3-Flash sits at fifty-seven points and ten cents. Gemini 3.7 Flash, on high, sits at fifty-six points and fifty-eight cents. Qwen3.8 sits at fifty-eight points and a dollar eighteen. GLM-5.3, on max, sits at sixty points and ninety cents. And the highest scoring model whose files you can download and run yourself is Kimi K3 on max, at sixty points and two dollars thirty-one. The top of the downloadable field charges twenty-three times GLM-5.3-Flash’s ten cents for three more points of score. Three points on this index, where one measuring body ran every test itself, is close enough that two models can come out level on your actual job.
Three jobs, three answers
For long batch work, where the same instructions repeat on every call, run GLM-5.3-Flash at ten cents blended and about forty-two tokens a second. For anything a person watches stream, run Gemini 3.7 Flash at about two hundred and eighty-five tokens a second, roughly seven times faster to write than GLM-5.3-Flash. Gemini is also the slowest of these four to start, waiting about ten and a half seconds before its first word, so Gemini wins on a long answer and loses on a short one. For the hardest reasoning on this board, run GLM-5.3 on max at sixty points, and pay about nine times the Flash price for those three points.
What the price follows
You pay about nine times the price for three points of score. That price follows the numbers a model wakes up on every word, and how much of your prompt it has already read. The model’s name tells you neither one.

