Zum Inhalt springen

Ollama Release Notes

28 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.32.7: Muse Glimmer 30B von Meta auf Apple Silicon mit MLX

Ollama unterstützt ab Version 0.32.7 das multimodale 30B-Modell Muse Glimmer von Meta zunächst über die MLX-Engine auf Apple Silicon, inklusive DFlash und Bildeingabe, und es lässt sich per ollama run muse-glimmer:30b-mlx sowie mit ollama launch für Agenten wie Claude Code, Pi, OpenClaw und Hermes nutzen.

Muse Glimmer

Note: Muse Glimmer is currently available via initial support via Ollama's MLX engine on Apple Silicon. Additional support and optimizations for Apple Silicon, NVIDIA, AMD, and other platforms will be available in the coming days.

Muse Glimmer, Meta's newest open model and the first released by Meta Superintelligence Labs, is now available on Ollama. It's a 30B multimodal model purpose-built for agent workloads that run locally.

With Ollama, you can now use Muse Glimmer to power coding agent applications such as Claude Code, Codex, Pi and more, as well as long-running personal assistants such as OpenClaw and Hermes.

Ollama's MLX engine provides state-of-the-art performance on Apple Silicon for this model, with support for DFlash and image input as of Ollama 0.32.7.

To download and run Muse Glimmer locally:

ollama run muse-glimmer:30b-mlx

To run Muse Glimmer on Apple Silicon with Claude Code, download Ollama and run:

ollama launch claude --model muse-glimmer:30b-mlx

For a lighter-weight coding agent, try Pi:

ollama launch pi --model muse-glimmer:30b-mlx

For personal assistant frameworks such as OpenClaw and Hermes, use:

ollama launch openclaw --model muse-glimmer:30b-mlx
ollama launch hermes --model muse-glimmer:30b-mlx

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.32.6: Schnelleres Qwen3.5 auf Apple-GPUs, OpenAI-Streaming angepasst

Ollama 0.32.6 beschleunigt Qwen3.5 auf Apple-GPUs durch automatische Nutzung des MTP-Heads für Speculative Decoding, gleicht das Streaming von /v1/chat/completions an das OpenAI-Format an, meldet abgeschnittene Antworten mit finish_reason: "length", behebt mehrere TUI-Fehler und entfernt die experimentelle Bildgenerierung vorübergehend (dafür weiter 0.32.5 nutzen).

What's Changed

  • Qwen3.5 is faster on Apple GPUs: the MLX engine now uses the model's MTP head for speculative decoding automatically
  • /v1/chat/completions streaming now matches OpenAI's wire format: role only on the first chunk, finish_reason on its own chunk, and usage in a separate chunk with stream_options.include_usage.
  • Truncated OpenAI responses now report finish_reason: "length" instead of "tool_calls".
  • ollama run kimi-k3 now offers kimi-k3:cloud for cloud-only models that publish no default tag, instead of failing.
  • TUI fixes: pipe-delimited prose no longer renders as a table, Enter accepts the highlighted @ file completion, and /prompt scrolling is no longer laggy.
  • Experimental image generation has been temporarily removed. Continue using 0.32.5 for image generation support
  • Updated the MLX and llama.cpp engines.

Full Changelog: https://github.com/ollama/ollama/compare/v0.32.5...v0.32.6-rc0

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.32.5: MLX-Metal-Fehler bei NVFP4-Modellen behoben

Ollama 0.32.5 behebt einen MLX-Metal-Fehler, der die Ausgabequalität von NVFP4-Modellen, insbesondere Laguna, verschlechtern konnte.

What's Changed

  • Fixed an MLX Metal bug that could reduce output quality for NVFP4 models, particularly Laguna.

Full Changelog: https://github.com/ollama/ollama/compare/v0.32.4...v0.32.5

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.32.4: Laguna auf Apple-GPUs, Qwen3-MoE-Decoding-Fix

Ollama 0.32.4 unterstützt Laguna auf Apple-GPUs über die MLX-Engine, quantisiert die Ausgabeköpfe von Draft-Modellen beim Erstellen von Speculative-Decoding-Drafts im gewünschten Typ und behebt das Decoding von Qwen3 MoE bei unterschiedlich quantisierten Experten, mit schnellerer gepackter Gate/Up-Projektion (ca. 4–9 % auf M5 Max).

What's Changed

  • Support Laguna on Apple GPUs via the MLX engine
  • Quantize draft-model output heads at the requested type when creating speculative-decoding drafts.
  • Fixed Qwen3 MoE decoding for differently-quantized experts, plus faster packed gate/up projection (~4–9% on M5 Max).

Full Changelog: https://github.com/ollama/ollama/compare/v0.32.3...v0.32.4

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Version 0.32.3: Download- und GLM-Fixes, bessere GPU-Unterstützung, Laguna 2.1

Ollama 0.32.3 behebt hängende Modell-Downloads und GLM-Tool-Calls, die am Ende der Generierung verworfen wurden, verbessert Integrationen (Claude Code Channels, Anthropic-Thinking-Streams, Hermes Desktop mit --force-build), erweitert die GPU-Unterstützung (CUDA auf Windows ARM64, B200 über CUDA 12, geringerer Speicherbedarf bei Linux-CUDA/ROCm-iGPUs) und ergänzt Chat, Thinking und Tool Calling für Laguna 2.1.

What's Changed

  • Fixed model downloads that stall before sending data.
  • Improved integrations: restored Claude Code Channels, fixed Anthropic thinking streams, and made Hermes Desktop respect --force-build.
  • Expanded GPU support with CUDA on Windows ARM64, B200 support through CUDA 12, and lower memory use on Linux CUDA/ROCm iGPUs.
  • Added chat, thinking, and tool calling support for Laguna 2.1 models, including a Metal inference fix.
  • Fixed GLM tool calls being silently dropped at the end of generation.
  • Updated the MLX and llama.cpp engines.

Full Changelog: https://github.com/ollama/ollama/compare/v0.32.1...v0.32.3

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Ollama 0.32.1: Besseres Tool Calling bei Gemma 4 und MLX-Fixes

Ollama 0.32.1 verbessert Tool Calling und Multi-Turn-Reasoning bei Gemma 4, behebt ein Cache-Leck bei rekurrenten MLX-Modellen, lässt MLX-Textmodelle OLLAMA_LOAD_TIMEOUT beachten, weist bei Websuche und Fetch des Agenten auf ollama signin hin, übergibt dem interaktiven Agenten das aktuelle Arbeitsverzeichnis und behebt die Modellauswahl bei ollama launch mit veralteten Modellen.

What's Changed

  • Improved Gemma 4 tool calling and multi-turn reasoning, including more reliable tool-response continuations
  • Fixed a recurrent MLX model cache leak that could increase memory use across requests, and improved cache snapshot performance
  • MLX text model loading now respects OLLAMA_LOAD_TIMEOUT
  • Agent web search and fetch now tell users to run ollama signin when authentication is required
  • The interactive agent now receives the current working directory for better project context
  • Fixed ollama launch so choosing Pick another model for a deprecated model passed with --model opens the model picker
  • Updated VS Code setup documentation for the official Ollama extension

Full Changelog: https://github.com/ollama/ollama/compare/v0.32.0...v0.32.1-rc0

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Ollama 0.32.0: Neuer interaktiver Agent und ChatGPT-Integration

Ollama 0.32.0 startet mit ollama einen neuen interaktiven Agenten für Coding und delegierte Aufgaben, benennt die Codex-App-Integration in ChatGPT um (ollama launch chatgpt), zeigt im ollama launch-Menü nur noch die beliebtesten Integrationen und warnt vor dem Start älterer Agent-Modelle.

What's Changed

  • New interactive agent experience: running ollama now launches an agent to help you code and delegate work
❯ ollama
Ollama 0.32.0

▸ Chat, Code, & Work (glm-5.2:cloud)
    Chat with models, code, search the web, and delegate real work
  • Renamed the Codex App integration to ChatGPT: use ollama launch chatgpt (and --restore to return to your usual ChatGPT profile)
  • Simplified integration selection: the ollama launch menu now only offers the most popular integrations (other integrations can be accessed through ollama launch
  • Warns before launching older agent models: CodeLlama, Qwen2.5(-coder), Llama 3.x, Mistral, StarCoder, and the base DeepSeek-R1 tags now prompt a deprecation warning before ollama launch continues

Full Changelog: https://github.com/ollama/ollama/compare/v0.31.2...v0.32.0

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Ollama

Ollama 0.31.2: Flash Attention auf älteren NVIDIA-GPUs und Fixes

Ollama 0.31.2 aktiviert Flash Attention auf älteren NVIDIA-GPUs (Compute Capability 6.x), erlaubt iGPUs das Auslagern von Vision-Modellen mit Padding, behebt Structured Output bei Thinking-Modellen ohne Thinking sowie das Laden von Modellen auf Pfaden mit Nicht-UTF-8-Zeichen, härtet die GGUF-Modellerstellung ab und deaktiviert Telemetrie bei ollama launch für Claude Code standardmäßig.

What's Changed

  • Enabled flash attention on older NVIDIA GPUs (compute capability 6.x)
  • iGPU can now offload vision models with padding to fit available memory
  • Fixed structured output for thinking models when thinking is disabled
  • Hardened GGUF model creation
  • ollama launch for Claude Code now disables telemetry by default
  • Fixed loading models on paths with non-UTF-8 characters
  • Updated the MLX and llama.cpp engines

New Contributors

Full Changelog: https://github.com/ollama/ollama/compare/v0.31.1...v0.31.2

Originalquelle(öffnet in neuem Tab)Problem melden