Ollama v0.34.4: single-pass structured outputs on thinking models, faster Qwen 3.8 on Apple Silicon

1 hour ago

@ollamaSubscribe

Ollama v0.34.4: structured outputs on thinking models now run in one pass, "model not found" and macOS app fixes, faster Qwen 3.8 and per-image Gemma 4 resolution on Apple Silicon, thinking-level discovery, and engine updates. https://github.com/ollama/ollama/releases/tag/v0.34.4 https://github.com/ollama/ollama/compare/v0.34.3...v0.34.4

Ask

Ask about this presentation

Answers are generated from this presentation.

Chapters

  1. 0:00Ollama v0.34.4: structured outputs on thinking models in one pass
  2. 0:13The prompt is evaluated once
  3. 0:24The thinking stays free, and only the answer follows the schema
  4. 0:35The generate endpoint keeps JSON out of the thinking
  5. 0:45With a format set, a tool call can't replace the answer
  6. 0:5802 Fixes
  7. 0:59Intermittent model not found errors are fixed
  8. 1:11The macOS app no longer hangs checking for ChatGPT or Codex
  9. 1:2403 Apple Silicon
  10. 1:25Qwen 3.8 processes prompts up to 19% faster on Apple Silicon
  11. 1:37Gemma 4 picks an image resolution for each image
  12. 1:48The largest image budget costs Gemma 4 12B about 830 MB more memory
  13. 2:0104 Thinking controls
  14. 2:03Ask a model which thinking levels it takes
  15. 2:16Set think to true, false, null, or a level name
  16. 2:27The OpenAI and Anthropic endpoints take the same level names
  17. 2:4205 Claude Code and the engines
  18. 2:44ollama launch claude uses Claude Code's own auto mode checks
  19. 2:56llama.cpp, MLX and XGrammar are updated
Show transcript

Ollama v0.34.4: structured outputs on thinking models in one pass

Ollama
Ollama
23 September 2026
Ollama
Ollama v0.34.4
Structured outputs on thinking models now apply in a single pass.
Release notes, ollama/ollama v0.34.4
Part 01

In Ollama v0.34.4, a thinking model that's asked for a JSON format handles it in a single pass. It thinks freely, then writes the formatted answer in the same generation.

The prompt is evaluated once

Ollama
Ollama
v0.34.4
Before
Two generations
Think, stop, re-render the prompt, prefill again under the grammar
v0.34.4
One generation
One prefill, metrics straight through
server/routes.go
Part 01

Before, a format on a thinking model meant two generations, and the second one had to prefill the whole prompt again. Now it's one generation, and the prompt is evaluated once.

The thinking stays free, and only the answer follows the schema

Ollama
Ollama
v0.34.4
Qwen 3 0.6B, temperature 0: the thinking is byte-identical with and without a format.
llm/gbnf.go
Part 01

The grammar now applies only after the thinking ends. On Qwen 3 0.6B at temperature zero, the thinking came out byte-identical with a format and without one.

The generate endpoint keeps JSON out of the thinking

Ollama
Ollama
v0.34.4
A raw prompt still gets the format from the first token.
server/routes.go
Part 01

The generate endpoint gets the fix too. It used to apply the format from the first token, which forced the JSON inside the thinking. Raw prompts keep that old behavior.

With a format set, a tool call can't replace the answer

Ollama
Ollama
v0.34.4

With a format set, a tool call can’t replace the answer

GPT-OSS and other Harmony models keep their tool calls.
server/routes.go
Part 01

One behavior change. With a format set, whatever follows the thinking has to match it, so a tool call can't stand in for the answer. Harmony models like GPT-OSS are the exception.

02 Fixes

Ollama
Ollama
v0.34.4
02
Fixes
Release notes, v0.34.4
Part 02

Fixes.

Intermittent model not found errors are fixed

Ollama
Ollama
v0.34.4

Intermittent “model not found” errors are fixed

With a large local library, a name lookup could borrow another model’s tag casing. It now tries an exact match first.
server/routes.go
Part 02

With a large local library, you may have seen model not found errors that went away on retry. The lookup could borrow another model's tag casing. It now tries an exact match first.

The macOS app no longer hangs checking for ChatGPT or Codex

Ollama
Ollama
v0.34.4

The macOS app stays responsive while it checks for ChatGPT and Codex

Before
System Events
A slow answer froze the window
v0.34.4
Running programs
Read directly from the list
cmd/launch/codex_app.go
Part 02

On macOS, the Apps page checks whether ChatGPT or Codex is running. It asked System Events, and a slow answer froze the window. It now reads the list of running programs directly.

03 Apple Silicon

Ollama
Ollama
v0.34.4
03
Apple Silicon
Release notes, v0.34.4
Part 03

Apple Silicon.

Qwen 3.8 processes prompts up to 19% faster on Apple Silicon

Ollama
Ollama
v0.34.4

Qwen 3.8 processes prompts up to 19% faster

+19%
2k-token prompt · M5 Max
+19%
8k-token prompt · M5 Max
+14%
16k-token prompt · M5 Max
mlx/gated_delta.go
Part 03

Qwen 3.8 now reads prompts faster on Apple Silicon. On an M5 Max, prompts from two thousand to sixteen thousand tokens were processed fourteen to nineteen percent faster.

Gemma 4 picks an image resolution for each image

Ollama
Ollama
v0.34.4
70
140
280
560
1120
Image-token budgets. The old fixed default was 280.
mlxrunner/model/gemma4/process_image.go
Part 03

Gemma 4 on Apple Silicon now picks one of five image budgets for each image. High-resolution documents keep more detail, and small images take fewer tokens.

The largest image budget costs Gemma 4 12B about 830 MB more memory

Ollama
Ollama
v0.34.4

The largest image budget costs about 830 MB more memory

826 MB
Gemma 4 12B · 70 to 1120
564 MB
Gemma 4 E4B · 70 to 1120
mlxrunner/model/gemma4/process_image.go
Part 03

On Gemma 4 12B, the largest image budget uses about 830 megabytes more memory than the smallest. A 16 gigabyte M1 Mac slows down with E4B, but it keeps running.

04 Thinking controls

Ollama
Ollama
v0.34.4
04
Thinking controls
Release notes, v0.34.4
Part 04

Thinking controls.

Ask a model which thinking levels it takes

Ollama
Ollama
v0.34.4
POST /api/show  {"model": "gpt-oss"}

"thinking": {
  "values": ["low", "medium", "high"],
  "default": "medium"
}
docs/capabilities/thinking.mdx
Part 04

To find which thinking levels a model accepts, call /api/show. It returns the supported values and the default. For GPT-OSS, that's low, medium and high, defaulting to medium.

Set think to true, false, null, or a level name

Ollama
Ollama
v0.34.4
Use the exact name from /api/show. An unsupported name gets the model’s default.
docs/capabilities/thinking.mdx
Part 04

The think field takes true, false, null for the model's default, or a level name spelled exactly as the show endpoint lists it. An unsupported name falls back to the default.

The OpenAI and Anthropic endpoints take the same level names

Ollama
Ollama
v0.34.4
  • Chat Completions reasoning_effort
  • Responses reasoning.effort and think
  • Anthropic Messages output_config.effort
docs/api/openai-compatibility.mdx
Part 04

The OpenAI and Anthropic compatible endpoints accept the same model-defined level names. That's reasoning effort on Chat Completions, reasoning effort and think on Responses, and the effort setting on Anthropic Messages.

05 Claude Code and the engines

Ollama
Ollama
v0.34.4
05
Claude Code and the engines
Release notes, v0.34.4
Part 05

Claude Code and the engines.

ollama launch claude uses Claude Code's own auto mode checks

Ollama
Ollama
v0.34.4

ollama launch claude uses Claude Code’s own auto mode checks

CLAUDE_CODE_AUTO_MODE_SERVER=0
Set by default. A value you set yourself is kept.
cmd/launch/claude.go
Part 05

Ollama doesn't provide Anthropic's server-side auto mode checks, so ollama launch now starts Claude Code with its client-side checks. If you set that variable yourself, your value is kept.

llama.cpp, MLX and XGrammar are updated

Ollama
Ollama
v0.34.4
  • llama.cpp build b10969 to b11081
  • MLX adds the fast gated-delta kernel
  • XGrammar 0.2.7 schema fixes for typed dictionaries and short arrays
LLAMA_CPP_VERSION, MLX_VERSION, cmake/mlx/CMakeLists.txt
Part 05

llama.cpp moves to build 11081, and MLX is updated. XGrammar 0.2.7 fixes structured output schemas with typed dictionary values and short arrays.