What ollama/ollama shipped
Written by FoxPlug from public releases; not affiliated with Ollama. An automatic summary of the public release, pull request and commit data of github.com/ollama/ollama. Ollama did not write it and does not use or endorse FoxPlug. Every line links to the public change it describes.
Week of September 21, 2026
What shipped
- Models run on MLX on Apple Silicon by default in v0.40.0, with supported architectures automatically using the MLX runtime. Release
- v0.34.4 applies structured outputs on thinking models in a single pass, making them faster and more reliable. Release
- Gemma 4 now selects image resolution dynamically across 70-1120 pixel budgets based on input resolution and aspect ratio. Pull request #18603
- Settings page opens faster by deferring model discovery and showing saved settings first without blocking the UI. Pull request #18598
- Fixed intermittent 'model not found' errors by properly tracking canonicalized model name parts during lookups. Pull request #18438
- The typical_p parameter now logs a warning instead of hard failing when used in requests. Pull request #18627
- macOS update icon and menu items now stay in sync at startup to prevent showing conflicting update states. Pull request #18622
- XGrammar updated to 0.2.7 with schema fixes for typed dictionary values and short arrays. Pull request #18615
Changelog entry
- Models run on MLX on Apple Silicon by default (v0.40.0) Release
- Structured outputs on thinking models now apply in single pass (v0.34.4) Release
- Gemma 4 dynamically selects image resolution based on input Pull request #18603
- Settings page defers model discovery for faster loading Pull request #18598
- Fixed intermittent model not found errors in large local libraries Pull request #18438
- typical_p parameter now logs warning instead of failing Pull request #18627
- macOS update icon and menu stay synchronized at startup Pull request #18622
- XGrammar updated to 0.2.7 with schema fixes Pull request #18615
Week of September 14, 2026
What shipped
- Server-side MLX imports now support safetensors through the create pipeline with remote upload, draft layers, and transfer limits, while GGUF create is limited to wrapping existing inputs. Pull request #14969
- The API now exposes model thinking levels and defaults through /api/show and ollama show so clients can offer controls matching each model's capabilities. Pull request #18473
- Server allows registry cross-host redirects among allowlisted hosts to improve registry flexibility. Pull request #18533
- Registry and blob transfer redirects now validate target schemes and addresses, re-check DNS on each redirect, and prevent downgrading https to http. Pull request #18512
- The MLX runner now releases freed KV buffers during speculative decode to prevent buffer pool exhaustion. Pull request #18510
- On macOS, closing the window with the red button or Command–W now keeps it closed when switching back with Command–Tab. Pull request #18518
- First-run onboarding is now shown when running ollama without a subcommand, with desktop-aligned copy and optional browser sign-up, shared between CLI and desktop. Pull request #18495
- MLX and MLX-C dependencies updated with a fix for quantized matmul corruption affecting nvfp4 multimodal models. Pull request #18449
Why it matters
The MLX engine graduates from experimental to production status with improved safetensors import support and memory management. Registry security and redirect handling are tightened to prevent downgrade attacks. New API metadata exposes model capabilities so clients can offer appropriate controls.
Changelog entry
- mlxrunner: Move MLX engine from x/ to top-level production directories Pull request #18489
- server: Add server-side MLX safetensors imports with remote upload and draft layer support Pull request #14969
- api: Expose model thinking levels and defaults through /api/show Pull request #18473
- server: Allow registry cross-host redirects among allowlisted hosts Pull request #18533
- server: Tighten redirect validation for registry requests with scheme and DNS checks Pull request #18512
- mlxrunner: Release freed KV buffers during speculative decode to prevent exhaustion Pull request #18510
- app: Keep closed macOS windows from reopening on activation Pull request #18518
- cli: Add first-run onboarding shared with desktop app Pull request #18495
- deps: Bump MLX and MLX-C with quantized matmul corruption fix Pull request #18449
- api: Deprecate typical_p parameter for new models Pull request #18448
- tests: Fix metadata deletion test flake Pull request #18493
v0.34.3-rc1: MLX engine moves to production, safetensors imports gain server-side support, registry redirects get stricter security validation, and model thinking levels are exposed through the API.
v0.34.3-rc1 ships significant infrastructure improvements: the MLX engine graduates from experimental status to top-level production use, safetensors imports now work server-side with upload and draft layer support, registry redirects are hardened against downgrade attacks, and the API exposes model thinking levels so clients can offer matching controls. Speculative decode also gains memory efficiency improvements.