Zum Inhalt springen

Ollama Release Notes

28 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Ollama 0.35.1: Clef-Modelle, mehr Websuchen und CAPABILITY in Modelfiles

Ollama unterstützt die multimodalen Decision-Modelle Clef und Clef Flash über /v1/systemone, erlaubt bis zu zehn Websuchen pro Antwort, bietet CAPABILITY-Deklarationen in Modelfiles und meldet für Decision-Modelle nur noch „decision“ als Fähigkeit.

Clef decision models

Ollama now supports Clef and Clef Flash, Cloudflare's new open-source decision models, through /v1/systemone.

Clef (27B) and Clef Flash (9B) are multimodal: requests can now include images alongside the text state, shared by all questions and scored jointly with it.

curl http://localhost:11434/v1/systemone -d '{
  "model": "clef-flash",
  "state": "The user took this screenshot.",
  "images": ["<base64-encoded image>"],
  "questions": {
    "has_ollama": {"type": "noul", "instructions": "Does this image contain Ollama?"}
  }
}'
{
  "model": "clef-flash",
  "answers": {
    "has_ollama": {
      "type": "noul",
      "noul": 0.958
    }
  },
  "usage": {
    "input_tokens": 548,
    "output_tokens": 0
  }
}

What's Changed

  • Models using web search can now perform up to ten searches per response, up from three
  • Modelfiles now support CAPABILITY declarations, so model creators can explicitly declare what a model can do. Declarations are preserved when creating from GGUF or safetensors, through model inheritance, and on Modelfile export
  • ollama show and the model list now report only decision as the capability for decision models, so clients no longer offer them for general chat, tools, or thinking
  • Updated llama.cpp and the MLX engine

Full Changelog: https://github.com/ollama/ollama/compare/v0.35.0...v0.35.1

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Ollama 0.35.0: Unterstützung für Decision-Modelle über /v1/systemone

Ollama unterstützt jetzt Decision-Modelle wie Nimble und Tev1 über /v1/systemone, die statt Text Entscheidungen, Wahrscheinlichkeiten und Scores zurückgeben.

Decision models

Ollama now supports decision models through /v1/systemone, based on TypeSafe’s Jev API.

Decision models return choices, probabilities, and scores instead of text. Use them for tasks such as ticket triage, model routing, and content classification.

Available models:

ollama pull nimble

Send context and one or more questions:

curl http://localhost:11434/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "nimble",
    "state": "Our checkout has returned 500 errors since 9am.",
    "questions": {
      "label": {
        "type": "choice",
        "instructions": "Which label fits this ticket?",
        "criteria": {
          "billing": "Payments and refunds",
          "bug": "Software errors",
          "account": "Login and account access"
        }
      }
    }
  }'

Example response:

{
  "model": "nimble",
  "answers": {
    "label": {
      "type": "choice",
      "choice": "bug",
      "probabilities": {
        "billing": 0.0125,
        "bug": 0.9781,
        "account": 0.0093
      },
      "confidence": 0.8906
    }
  },
  "usage": {
    "input_tokens": 174,
    "output_tokens": 1
  }
}

The API supports three question types: …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Ollama 0.40.0: MLX auf Apple Silicon jetzt Standard

Auf Apple Silicon laufen Modellarchitekturen, die die MLX-Runtime unterstützt, jetzt standardmäßig auf MLX, darunter gemma4, qwen3.6, qwen3.5, mehrere Decision-Modelle sowie das Embedding-Modell embeddinggemma-2.

What's Changed

Models run on MLX on Apple Silicon by default

In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX.

ollama pull qwen3.8
ollama run qwen3.8

Additional models include gemma4, qwen3.6 and qwen3.5

Decision models are now available on MLX as well: Nimble tev1 clef clef-flash

MLX now has support for an embedding model: embeddinggemma-2

We will continue testing and enabling additional models.

Full Changelog: https://github.com/ollama/ollama/compare/v0.35.1...v0.40.0

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Ollama 0.34.4: Strukturierte Ausgaben in einem Durchgang, Fehlerbehebungen

Strukturierte Ausgaben bei Thinking-Modellen laufen in einem Durchgang, außerdem gibt es Fehlerbehebungen (Model-not-found, hängende macOS-App), schnellere Qwen-3.8-Verarbeitung und bessere Bildauflösung bei Gemma 4 auf Apple Silicon.

What's Changed

  • Structured outputs on thinking models now apply in a single pass, making them faster and more reliable.
  • Fixed intermittent "model not found" errors with a large local library
  • Fixed the macOS app becoming unresponsive when checking if ChatGPT or Codex is running.
  • Qwen 3.8 prompt processing is faster on Apple Silicon.
  • Gemma 4 on Apple Silicon now picks the best image resolution per image, keeping more detail in high-resolution images.
  • Updated llama.cpp, MLX, and XGrammar.

Full Changelog: https://github.com/ollama/ollama/compare/v0.34.3...v0.34.4

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Ollama 0.34.3: Thinking-Steuerung in show, Nemotron-H-Vision mit MLX

GET /api/show und ollama show zeigen nun die Thinking-Steuerung und den Standardwert jedes Modells an, Nemotron-H-Vision-Modelle laufen auf Apple Silicon mit MLX, und Modell-Pulls von HuggingFace sowie das Fensterverhalten der macOS-App wurden korrigiert.

What's Changed

GET /api/show now advertises each model's thinking controls and default:

Available in the CLI with:

ollama show gemma4
    thinking
        levels     false, true
        default    true

Available in the API with:

curl http://localhost:11434/api/show -d '{"model": "glm-5.3-flash:cloud"}'
{
  "thinking": {
    "values": ["low", "high", "max"],
    "default": "max"
  }
}

Also available on ollama.com directly for cloud models.

  • Nemotron H vision models are now supported on Apple Silicon with MLX
  • Ollama's macOS app will now no longer reopen windows you've closed when activating it
  • Fix for model pulls from HuggingFace

Full Changelog: https://github.com/ollama/ollama/compare/v0.34.2...v0.34.3

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.34.2: Setup beim ersten Start, ollama://apps, MLX-Speicherfix

Beim ersten Start von ollama gibt es ein Setup mit Anmelde- oder lokaler Option, ollama://apps öffnet die Apps-Seite der Desktop-App, und übermäßiges Speicherwachstum bei MLX Speculative Decoding wurde behoben.

What's Changed

  • Added first-run setup when running ollama, with options to sign in or continue locally. Setup completion is shared with the desktop app on macOS and Windows.
  • Added ollama://apps to open the desktop app’s Apps page directly on macOS and Windows.
  • Fixed excessive memory growth during long generations with MLX speculative decoding.
  • Updated llama.cpp.

Full Changelog: https://github.com/ollama/ollama/compare/v0.34.1...v0.34.2

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.34.1: MLX-Safetensors-Create stabil, schnellere /api/tags

ollama create mit MLX-Safetensors ist nicht mehr experimentell, /api/tags ist bei großen Modellbibliotheken deutlich schneller, die Erkennung wiederholter Tokens ist weniger fehleranfällig und typical_p ist veraltet.

What's Changed

  • MLX safetensors ollama create no longer experimental. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization.
  • Improved MLX memory handling on Apple Silicon
  • Runaway repeat token detection now requires 100 repeat tokens for reduced false positives (e.g. OCR)
  • /api/tags is much faster on large model libraries (3.1 s → 294 ms cold in testing), and model capabilities are now reported consistently.
  • Deprecated typical_p: it can no longer be set when creating new models, existing GGUF models retain support.
  • MLX and llama.cpp updates

Full Changelog: https://github.com/ollama/ollama/compare/v0.34.0...v0.34.1-rc1

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.34.0: Ollama-Modelle in ChatGPT Desktop, schnellere strukturierte Ausgaben

Ollama-Modelle lassen sich jetzt direkt in ChatGPT Desktop nutzen, zudem wurden strukturierte Ausgaben auf Apple Silicon beschleunigt und Tool-Suche sowie Response Compaction für OpenAI-kompatible Clients ergänzt.

Use Ollama models in ChatGPT Desktop

Ollama models can now be used directly in ChatGPT Desktop, so you can keep your existing workflow while running open models. Setup is available from the Ollama app on MacOS.

<img width="1374" height="1300" alt="CleanShot 2026-09-08 at 11 04 07 AM@2x" src="https://github.com/user-attachments/assets/e10e299d-c11f-447d-9234-afa855824efe" />

This release also improves structured output performance on Apple Silicon, adds support for OpenAI-compatible client tool search and response compaction.

Full Changelog: https://github.com/ollama/ollama/compare/v0.33.3...v0.34.0

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.33.3: gemma4 mit Bild und Audio auf MLX, gecachte Tokens

gemma4 unterstützt auf der MLX-Engine nun Bilder und Audio, gecachte Prompt-Tokens werden gemeldet und im GGUF-Modell definierte Standardparameter werden berücksichtigt.

What's Changed

  • gemma4 now supports images and audio on MLX engine
  • Report cached prompt tokens
  • Honor GGUF model defined default parameters
  • MLX, MLX-C, llama.cpp update

New Contributors

Full Changelog: https://github.com/ollama/ollama/compare/v0.33.2...v0.33.3

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.33.2: Systemdesign mit Dark Mode, macOS-Übergabe an laufende Instanz

Die Ollama-App folgt wieder dem Systemdesign inklusive Dark Mode, die macOS-App übergibt an eine bereits laufende Instanz, und der Claude-Desktop-Proxy unterbricht laufende Anfragen bei Katalog-Updates nicht mehr.

What's Changed

  • Ollama's app now follows the system appearance again, restoring dark mode support
  • Fixed the macOS app to properly hand off to an already-running instance instead of starting a second one
  • The Claude Desktop proxy no longer interrupts in-flight requests when the model catalog updates

Full Changelog: https://github.com/ollama/ollama/compare/v0.33.1...v0.33.2

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.33.1: Qwen3.8 Flash Next unter MLX, strukturierte Ausgaben

Unter MLX wird Qwen3.8 Flash Next unterstützt, der mlxrunner bietet strukturierte Ausgaben und vermeidet Metal-GPU-Timeouts beim Laden von langsamem Speicher.

What's Changed

  • MLX: Qwen3.8 Flash Next support
  • cmake: make external compat patches idempotent
  • MLX and llama.cpp update
  • mlxrunner: add structured output support
  • mlxrunner: avoid Metal GPU timeouts when loading models from slow storage

New Contributors

Full Changelog: https://github.com/ollama/ollama/compare/v0.33.0...v0.33.1

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.33.0: Claude Desktop mit Ollama-Gateway, besseres Prefill-Caching

Claude Desktop lässt sich nun mit Ollama als Drittanbieter-Gateway konfigurieren, das Prefill-Caching wurde verbessert (u. a. Hänger bei abgebrochenem Prefill behoben) und der DeepSeek-Harness-Launcher weicht bei Fehlern auf npx aus.

What's Changed

Claude Desktop

Developers can now easily configure Claude Desktop to seamlessly work with Ollama as a third-party gateway provider.

<img width="1298" height="780" alt="image" src="https://github.com/user-attachments/assets/c02fd93b-f599-40a9-83c3-d0555c0ac182" />

Improved caching

  • Fixed a hang where agent clients that cancel long prefills
  • Prefill restore points are now trustworthy by construction: a cancelled prefill keeps every restore point it crossed, so retries resume where they stopped instead of restarting from scratch
  • Resumed prefills no longer record restore points that fail to cover what they claim; on models with recurrent layers this previously forced a request matching 46k of 47k tokens to reprocess from zero
  • Disabled Claude Code's "tokens left" token-countdown system message, which Ollama moved to the front of the prompt and broke the KV cache on every request

Other improvements

  • DeepSeek Harness launcher now falls back to npx when the global npm install fails, with Windows command-shim support
  • MLX dependency update (#17886)
  • Fixed broken default packaging caused by macOS-specific assumptions affecting Linux/Windows builds

Full Changelog: https://github.com/ollama/ollama/compare/v0.32.15...v0.33.0

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.32.15: Onboarding, schnellerer erster Token, Parser-Hänger behoben

Ollama bietet einen neuen Onboarding-Ablauf beim ersten Start der Desktop-App, halbiert durch Metadaten-Caching die Zeit bis zum ersten Token (ca. 995 ms auf 524 ms) und behebt ein Hängen von chat und generate nach Parser-Fehlern.

What's Changed

  • New desktop onboarding flow on first launch
  • Caches resolved model metadata between requests, cutting time-to-first-token by roughly half (TTFT dropped from ~995 ms to ~524 ms in benchmarks)
  • Fixes a bug where chat and generate could wedge after a mid-stream parser error
  • Qwen 3.8 system messages are now normalized so non-leading system messages are handled consistently
  • MLX and llama.cpp dependency updates

New Contributors

Full Changelog: https://github.com/ollama/ollama/compare/v0.32.14...v0.32.15

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.32.14: WebP-Transkodierung für llama-server, Qwen-Renderer

WebP-Bilder werden für llama-server transkodiert, und der Qwen-Renderer toleriert nun System-Nachrichten, die nicht am Anfang stehen.

What's Changed

  • llm: transcode WebP images for llama-server
  • renderers/qwen: tolerate non-leading system messages

Full Changelog: https://github.com/ollama/ollama/compare/v0.32.13...v0.32.14

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.32.12: Unterstützung für Qwen 3.8 27B, MLX-Variante

Ollama unterstützt jetzt Qwen 3.8 27B, das für Coding und Agenten-Aufgaben ausgelegt ist, auf Apple Silicon zusätzlich mit optimierter MLX-Variante (qwen3.8:27b-mlx).

<img width="600" alt="ollama and qwen capybara" src="https://github.com/user-attachments/assets/5fa5e93b-d48b-4a6a-b72a-25379ab9429c" />

Qwen 3.8 27B

This release adds the support of Qwen 3.8 27B. Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks.

ollama run qwen3.8:27b

For Apple Silicon devices, Ollama has in particular optimized for maximum performance and output quality suitable for repeated tasks and coding agents.

ollama run qwen3.8:27b-mlx

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.32.11: DeepSeek Harness und Muse Code in ollama launch, Websuche

ollama launch unterstützt nun DeepSeek Harness (dsh) und Muse Code (muse), und die OpenAI-kompatible Responses API unterstützt Websuche.

What's Changed

  • ollama launch dsh now supports DeepSeek Harness, DeepSeek's open-source agent harness
  • ollama launch muse now supports Muse Code, Meta's agentic coding CLI
  • The OpenAI-compatible Responses API now supports web search
  • Muse Glimmer template updates

Full Changelog: https://github.com/ollama/ollama/compare/v0.32.10...v0.32.11

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.32.10: repeat_penalty-Standard 1.0, schnelleres NVFP4-Prefill

Modelle ohne gesetzten repeat_penalty nutzen nun standardmäßig 1.0 statt 1.1, NVFP4-MLX-Modelle haben ein um etwa 7–8 % schnelleres Prefill, und die Blob-Verifizierung bei gleichem Digest in OCI-Manifests wurde korrigiert.

What's Changed

  • Models that don't set a repeat_penalty now default to 1.0 (off) instead of 1.1, matching other engines and speeding up speculative decoding; set a per-model parameter if an older model repeats itself.
  • Faster prefill on NVFP4 MLX models with a global scale, about 7–8% on Qwen3.6 and Muse Glimmer.
  • Fixed blob verification being skipped when an OCI manifest's config and layer share a digest.

New Contributors

Full Changelog: https://github.com/ollama/ollama/compare/v0.32.8...v0.32.10-rc1

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.32.9: NVIDIA Nemotron 3.5 Lightning, Muse-Glimmer-Parserfix

Ollama unterstützt jetzt NVIDIA Nemotron 3.5 Lightning samt Nemotron-3-Architektur und korrigiert einen Grenzfall im Function-Calling-Parser von Muse Glimmer.

NVIDIA Nemotron 3.5 Lightning

NVIDIA Nemotron 3.5 Lightning is an open 30B mixture-of-experts (MoE) model with 3B active parameters built for that execution layer of always-on agents. It is designed for harnesses like OpenClaw and Hermes Agent – all supported by the NVIDIA NemoClaw open source security and management stack for running always-on AI agents.

ollama run nemotron-3.5-lightning

What's Changed

  • Added the Nemotron 3 architecture
  • Handle boundary condition in Muse Glimmer function calling parser

Full Changelog: https://github.com/ollama/ollama/compare/v0.32.8...v0.32.9

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.32.8: Muse Glimmer auf allen Plattformen verfügbar

Muse Glimmer ist jetzt auf allen Plattformen verfügbar, einschließlich NVIDIA, AMD und weiteren, und lässt sich per ollama launch mit Claude Code, Pi, OpenClaw und Hermes nutzen.

Muse Glimmer

Muse Glimmer is now available on all platforms. Muse Glimmer can power coding agent applications such as Claude Code, Codex, Pi and more, as well as long-running personal assistants such as OpenClaw and Hermes.

Ollama's MLX engine provides state-of-the-art performance on Apple Silicon for this model, with support for DFlash and image input as of Ollama 0.32.7.

To download and run Muse Glimmer locally:

ollama run muse-glimmer

To run Muse Glimmer with Claude Code, download Ollama and run:

ollama launch claude --model muse-glimmer

For a lighter-weight coding agent, try Pi:

ollama launch pi --model muse-glimmer

For personal assistant frameworks such as OpenClaw and Hermes, use:

ollama launch openclaw --model muse-glimmer
ollama launch hermes --model muse-glimmer

What's Changed

  • Add Muse Glimmer support for NVIDIA, AMD, and additional platforms

Full Changelog: https://github.com/ollama/ollama/compare/v0.32.7...v0.32.8

Originalquelle(öffnet in neuem Tab)Problem melden