Gemini 3.8 Flash costs about forty percent more per measured task at the same token price

3 hours ago

Same sticker. A bigger invoice. Gemini 3.8 Flash is the model that did that. Artificial Analysis billed fifty-eight cents per Intelligence Index task for Gemini 3.8 Flash on high, against forty cents for Gemini 3.7 Flash on high. Cost per Index task is what Artificial Analysis billed to finish one of those tests. The Intelligence Index is a single score Artificial Analysis builds by running the same tests on every model themselves. Gemini 3.8 Flash scores fifty-nine on that index. Gemini 3.7 Flash scores fifty-six on that index. Google lists both at seventy-five cents per million input tokens, the text you send, and three dollars seventy-five per million output tokens, the text the model writes back.

Ask

Ask about this presentation

Answers are generated from this presentation.

Chapters

Show transcript

Same sticker, a bigger invoice

Same sticker. A bigger invoice. Gemini 3.8 Flash is the model that did that. Artificial Analysis billed fifty-eight cents per Intelligence Index task for Gemini 3.8 Flash on high, against forty cents for Gemini 3.7 Flash on high. Cost per Index task is what Artificial Analysis billed to finish one of those tests. The Intelligence Index is a single score Artificial Analysis builds by running the same tests on every model themselves. Gemini 3.8 Flash scores fifty-nine on that index. Gemini 3.7 Flash scores fifty-six on that index. Google lists both at seventy-five cents per million input tokens, the text you send, and three dollars seventy-five per million output tokens, the text the model writes back.

Gemini 3.8 Flash spec card

You can hand Gemini 3.8 Flash about a million tokens at once. Artificial Analysis clocks Gemini 3.8 Flash on high at about three hundred and forty-eight tokens a second. That is how fast the words come out once Gemini 3.8 Flash starts. The wait before the first word is about thirteen seconds. Google does not publish the model itself. Gemini has no license field to read.

Hy4 preview spec card

Tencent Cloud lists Hy4 preview at eighty-three point four cents per million input tokens, two point five zero one dollars per million output tokens, and four point two cents cached. OpenRouter clocks Tencent Cloud at forty-three tokens a second. That figure is OpenRouter’s median. The LICENSE file reads Apache 2.0.

GLM-5.3 spec card

Z.ai lists GLM-5.3 at a dollar forty per million input tokens and four dollars forty per million output tokens. Artificial Analysis clocks GLM-5.3 on max at about sixty-nine tokens a second. The LICENSE file is a custom grant Z.AI wrote.

Two cheap Flashes

Z.ai lists GLM-5.3-Flash at fifteen cents in and fifty cents out, per million tokens, with a half-price promo through September ninth. The LICENSE file reads MIT. Qwen3.8-Flash-Next is the downloadable preview. QwenCloud’s managed sibling is a different artifact with a one-million-token window. Artificial Analysis lists the preview at fifteen cents in and forty-seven cents out. The license field reads Qwen community one point zero.

120 million vs 64 million

Google’s Standard paid schedule for both Flashes doubles on January first, twenty twenty-seven. Artificial Analysis counted about thirty percent more output tokens per Index task on Gemini 3.8 Flash, averaging forty-eight thousand output tokens per task, plus more turns on agent tests. Across the whole Index, Gemini 3.8 Flash wrote a hundred and twenty million output tokens. Gemini 3.7 Flash wrote sixty-four million. The thirty percent is the average per task. The hundred and twenty million is the Index total.

The half price ends September 9

Z.ai’s page sets GLM-5.3-Flash’s launch promotion at seven and a half cents per million input tokens and twenty-five cents per million output tokens, ending at midnight on September ninth, Singapore time, six days from today. Budget on the list behind it: fifteen cents in, fifty cents out, per million tokens.

You pay for tokens written

The sticker is a price per million words. The invoice is how many words get written. Google still lists seventy-five cents in and three dollars seventy-five out for both Flashes. Artificial Analysis’s blended token rate is also the same: fifty-eight cents per million tokens for both. A blended price is one dollar figure standing in for a mixed bill: seven parts prompt text the model has already seen, two parts new prompt, one part what it writes back. The task rate is a different number. Artificial Analysis billed fifty-eight cents per Intelligence Index task for Gemini 3.8 Flash on high, against forty cents for Gemini 3.7 Flash on high. Gemini 3.8 Flash wrote more output tokens and took more agent turns on the same tests. A list price that stays put still raises the invoice when the model writes more.

Four LICENSE files, four grants

A license is a second price. Four files, four grants. Hy4’s LICENSE file is Apache 2.0: use it, including commercially. GLM-5.3-Flash’s file is MIT: the same commercial grant. GLM-5.3 is a custom grant. If you sell GLM-5.3 as a service, and the firm plus affiliates cleared ten billion dollars in any twelve months, you need Z.AI’s security review before commercial use. Products with the model embedded in a feature sit outside that gate. Qwen3.8-Flash-Next’s field is Qwen community one point zero. Selling it as a service or as an AI work assistant needs a separate license from Qwen. Past a hundred million monthly users or twenty million dollars monthly revenue, the model name has to sit on the interface. Serving cost tracks what is awake. The file tracks what you may sell.

Seven priced models of 190

The cheap corner is GLM-5.3-Flash at fifty-seven points and ten cents blended. Qwen3.8-Flash-Next sits at fifty-six points and nine cents blended. Gemini 3.8 Flash, on high, sits at fifty-nine points and fifty-eight cents blended, straight above Gemini 3.7 Flash on high at fifty-six points and the same blend. Kimi K3 on max is the highest scoring model whose files you can download and run yourself: sixty points and two dollars thirty-one. Claude Fable 5.1, Adaptive Reasoning, Max Effort, Default Fallback, leads at sixty-six points and seven dollars seventeen blended. Hy4 preview has no independent Index. Artificial Analysis’s Hy4 page was empty on September third.

Seven times faster, thirteen seconds to start

For work a person watches stream, run Gemini 3.8 Flash on high at about three hundred and forty-eight tokens a second, Index fifty-nine, fifty-eight cents per Index task. Gemini waits about thirteen seconds before the first word. For cheap measured work that can wait, run GLM-5.3-Flash at fifty-seven on the Index, ten cents blended, nine cents per Index task. For Apache files you can ship, Hy4 is the grant, with no independent Index and forty-three tokens a second on Tencent Cloud. Qwen3.8-Flash-Next sits at fifty-six points and nine cents blended, under the community grant. For the hardest reasoning on this board, GLM-5.3 on max and Kimi K3 on max both sit at sixty points.

Tokens written, and the grant in the file

Google kept the token price. The invoice still rose. Four LICENSE files wrote four different grants. Pay the score, the tokens written, and the file.