Ollama v0.34.4: structured outputs on thinking models in one pass
In Ollama v0.34.4, a thinking model that's asked for a JSON format handles it in a single pass. It thinks freely, then writes the formatted answer in the same generation.
The prompt is evaluated once
Before, a format on a thinking model meant two generations, and the second one had to prefill the whole prompt again. Now it's one generation, and the prompt is evaluated once.
The thinking stays free, and only the answer follows the schema
The grammar now applies only after the thinking ends. On Qwen 3 0.6B at temperature zero, the thinking came out byte-identical with a format and without one.
The generate endpoint keeps JSON out of the thinking
The generate endpoint gets the fix too. It used to apply the format from the first token, which forced the JSON inside the thinking. Raw prompts keep that old behavior.
With a format set, a tool call can't replace the answer
With a format set, a tool call can’t replace the answer
One behavior change. With a format set, whatever follows the thinking has to match it, so a tool call can't stand in for the answer. Harmony models like GPT-OSS are the exception.
02 Fixes
Fixes.
Intermittent model not found errors are fixed
Intermittent “model not found” errors are fixed
With a large local library, you may have seen model not found errors that went away on retry. The lookup could borrow another model's tag casing. It now tries an exact match first.
The macOS app no longer hangs checking for ChatGPT or Codex
The macOS app stays responsive while it checks for ChatGPT and Codex
On macOS, the Apps page checks whether ChatGPT or Codex is running. It asked System Events, and a slow answer froze the window. It now reads the list of running programs directly.
03 Apple Silicon
Apple Silicon.
Qwen 3.8 processes prompts up to 19% faster on Apple Silicon
Qwen 3.8 processes prompts up to 19% faster
Qwen 3.8 now reads prompts faster on Apple Silicon. On an M5 Max, prompts from two thousand to sixteen thousand tokens were processed fourteen to nineteen percent faster.
Gemma 4 picks an image resolution for each image
Gemma 4 on Apple Silicon now picks one of five image budgets for each image. High-resolution documents keep more detail, and small images take fewer tokens.
The largest image budget costs Gemma 4 12B about 830 MB more memory
The largest image budget costs about 830 MB more memory
On Gemma 4 12B, the largest image budget uses about 830 megabytes more memory than the smallest. A 16 gigabyte M1 Mac slows down with E4B, but it keeps running.
04 Thinking controls
Thinking controls.
Ask a model which thinking levels it takes
POST /api/show {"model": "gpt-oss"}
"thinking": {
"values": ["low", "medium", "high"],
"default": "medium"
}
To find which thinking levels a model accepts, call /api/show. It returns the supported values and the default. For GPT-OSS, that's low, medium and high, defaulting to medium.
Set think to true, false, null, or a level name
The think field takes true, false, null for the model's default, or a level name spelled exactly as the show endpoint lists it. An unsupported name falls back to the default.
The OpenAI and Anthropic endpoints take the same level names
- Chat Completions
reasoning_effort - Responses
reasoning.effortandthink - Anthropic Messages
output_config.effort
The OpenAI and Anthropic compatible endpoints accept the same model-defined level names. That's reasoning effort on Chat Completions, reasoning effort and think on Responses, and the effort setting on Anthropic Messages.
05 Claude Code and the engines
Claude Code and the engines.
ollama launch claude uses Claude Code's own auto mode checks
ollama launch claude uses Claude Code’s own auto mode checks
CLAUDE_CODE_AUTO_MODE_SERVER=0
Ollama doesn't provide Anthropic's server-side auto mode checks, so ollama launch now starts Claude Code with its client-side checks. If you set that variable yourself, your value is kept.
llama.cpp, MLX and XGrammar are updated
- llama.cpp build b10969 to b11081
- MLX adds the fast gated-delta kernel
- XGrammar 0.2.7 schema fixes for typed dictionaries and short arrays
llama.cpp moves to build 11081, and MLX is updated. XGrammar 0.2.7 fixes structured output schemas with typed dictionary values and short arrays.



















