Zum Inhalt springen

Rapid-MLX Release Notes

32 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Folge Rapid-MLX, um die Release Notes in deinen Feed zu holen.

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.15.7: Embeddings und zuverlässigerer Agent-Setup

Rapid-MLX 0.15.7 liefert über /v1/embeddings native Text- und Code-Embeddings mit EmbeddingGemma 2 (768 Dimensionen, 8192-Token-Limit), verbessert die Einrichtung von Continue, Cline und Qwen Code und bringt Fixes für Prompt-Wiederverwendung sowie MTP-Scheduling.

What's new in v0.15.7

Rapid-MLX 0.15.7 adds native text and code embeddings, improves local agent setup, and strengthens Desktop and inference reliability.

Highlights

  • EmbeddingGemma 2 text/code embeddings: serve normalized 768-dimensional vectors through /v1/embeddings, with standard 4-bit and BF16 aliases, optional smaller dimensions, and an explicit 8192-token limit. This release supports text/code inputs; images, audio and video are unsupported. (#4279)
  • More reliable agent setup: Continue and Cline configuration is written where their clients read it; keyed Continue connections preserve the configured server credential. Qwen Code previews redact retained secrets while keeping the saved configuration intact. (#4257, #4274, #4275)
  • Warmer coding-agent turns: prompt reuse avoids repeated full-prompt tokenization; MTP scheduling and reproducibility receive further fixes. Improvements depend on the model, hardware and workload. (#4214, #4221, #4238) …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.15.7: Embeddings und zuverlässigerer Agent-Setup

Rapid-MLX 0.15.7 liefert über /v1/embeddings native Text- und Code-Embeddings mit EmbeddingGemma 2 (768 Dimensionen, 8192-Token-Limit), verbessert die Einrichtung von Continue, Cline und Qwen Code und bringt Fixes für Prompt-Wiederverwendung sowie MTP-Scheduling.

[0.15.7] — 2026-10-07

Rapid-MLX 0.15.7 adds native text and code embeddings, improves local agent setup, and strengthens Desktop and inference reliability.

Highlights

  • EmbeddingGemma 2 text/code embeddings: serve normalized 768-dimensional vectors through /v1/embeddings, with standard 4-bit and BF16 aliases, optional smaller dimensions, and an explicit 8192-token limit. This release supports text/code inputs; images, audio and video are unsupported. (#4279)
  • More reliable agent setup: Continue and Cline configuration is written where their clients read it; keyed Continue connections preserve the configured server credential. Qwen Code previews redact retained secrets while keeping the saved configuration intact. (#4257, #4274, #4275)
  • Warmer coding-agent turns: prompt reuse avoids repeated full-prompt tokenization; MTP scheduling and reproducibility receive further fixes. Improvements depend on the model, hardware and workload. (#4214, #4221, #4238) …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.15.6: Interaktive Startoberfläche, PFlash standardmäßig aus

Rapid-MLX 0.15.6 öffnet beim Start von rapid-mlx im Terminal eine interaktive Startoberfläche, deaktiviert die PFlash-Prompt-Kompression standardmäßig für alle Aliase (nur per --pflash auto oder --pflash always aktivierbar) und verbessert Modell-Setup, Cache-Wiederverwendung, Speicherfreigabe und BYOM-Diagnose.

What's new in v0.15.6

Rapid-MLX 0.15.6 makes the first local-model session easier to start and long-running servers more predictable. It adds an interactive command-line front door, keeps prompt compression explicit, and strengthens model setup, cache reuse, memory recovery, and bring-your-own-model diagnostics.

Highlights

A clearer first step — Running bare rapid-mlx in a terminal now opens a compact interactive front door. It shows the active or last-used model when available, distinguishes the quick-start choice from the best model for the current Mac, and waits for an explicit keypress before chatting, starting a server, connecting a detected coding agent, or opening the model picker. rapid-mlx --help is grouped by task, while non-interactive use keeps deterministic copy-and-paste guidance. (#4120)

Lossless prompts by default — PFlash prompt compression is off for every alias unless you explicitly select --pflash auto or --pflash always. Existing opt-in acceleration remains available. When an opted-in request is compressed, supported APIs report the original and retained token counts through response metadata or a response header, so compression is visible. (#4104) …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.15.6: Neue Startoberfläche, gruppierte Hilfe, schnellere Speicherfreigabe

Rapid-MLX 0.15.6 bringt eine interaktive Startoberfläche für rapid-mlx, eine nach Aufgaben gruppierte --help, PFlash standardmäßig aus, System One mit Clef und Clef-Flash, optionale BYOM-Telemetrie sowie schnellere Freigabe von KV-Zuständen und Metal-Speicher.

[0.15.6] — 2026-10-04

Rapid-MLX 0.15.6 adds a simpler interactive first run and improves reliability for long prompts, local agents, custom models, shared caches, and long-running servers.

Added

  • Bare rapid-mlx now opens a compact interactive front door that shows the active or last-used model when available and waits for an explicit keypress before chatting, serving, connecting a detected coding agent, or choosing another model.
  • rapid-mlx --help is grouped by task, with concise deterministic guidance for non-interactive use.
  • System One supports Clef and Clef-Flash decision models for validated text and media inputs.
  • Opt-in BYOM telemetry records closed-category preflight, suggestion, support-request, and import outcomes without sending local paths or custom model names.

Changed

  • PFlash prompt compression is off by default for every alias. --pflash auto and --pflash always remain explicit opt-ins, and compressed responses expose retained and original token counts on supported APIs.
  • Generated Codex, OpenCode, Claude, and Pi configurations preserve user settings, use live server metadata, and avoid persisting local-server secrets.
  • Supported MTP models yield into ordinary batching when requests overlap while retaining the qualified speculative path for a single request.

Fixed

  • Completed single-request KV state is released promptly, and idle admission rechecks reclaimed Metal memory before refusing new work. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Version 0.15.5: Computer-Use im Hintergrund, schnelleres MTP, Modellprüfung vor Download

Rapid-MLX 0.15.5 ermöglicht experimentelle Computer-Use-Eingaben für geeignete Aktionen im Hintergrund, reduziert den Host-Synchronisierungsaufwand bei akzeptierten MTP-Drafts und prüft eigene Modelle schon vor dem Download.

What's new in v0.15.5

Rapid-MLX 0.15.5 makes experimental Computer Use less disruptive, speeds up qualified speculative decoding, and improves model import and server recovery.

Highlights

Computer Use gains background input for eligible actions. Eligible clicks, plain text entry, and non-Command keystrokes can target the exact app window approved for the task without bringing it forward. Window identity is checked immediately before input, and focus restoration is skipped if you switch apps during an action. Finder, Command shortcuts, and automatic fallback when background routing is unavailable still use the foreground; forced background mode refuses that fallback. This remains an experimental feature and still requires macOS Screen Recording and Accessibility permission. (#4052)

Less host overhead during accepted MTP drafts. Qualified continuous self-MTP decoding now performs one host synchronization per accepted draft cycle rather than one per proposed token. The output contract and existing profile eligibility checks are unchanged; the gain depends on draft acceptance and workload. (#4071)

Bring-your-own-model setup fails earlier and recovers better. Rapid checks uncataloged models before downloading them, suggests a runnable alternative when a model does not fit the current machine or runtime, and provides an …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.15.5 (Mac): Modell-Entladen im Chat, Background-Computer-Use, --context-length

Rapid-MLX 0.15.5 macht das Entladen von Modellen im Chat besser erreichbar, bietet Background-first Computer Use, schnelleres kontinuierliches MTP-Decoding für Qwen und GLM, sicherere Modellimporte und eine neue --context-length-Option.

[0.15.5] — 2026-10-03

Rapid-MLX 0.15.5 makes experimental Computer Use less disruptive, speeds up qualified speculative decoding, and improves model import and server recovery.

Changed

  • Model unload is now a prominent labelled action beside the active model in the chat composer, while the resident-memory footer remains available as a secondary entry point. Multi-model pools say Unload all, and active work retains the existing guarded/disabled behaviour.
  • Background-first Computer Use. Eligible clicks, plain text entry, and non-Command keystrokes can target the exact approved app window without bringing it forward. Finder, Command shortcuts, and automatic fallback when background routing is unavailable still use the foreground; forced background mode refuses that fallback.
  • Faster continuous MTP decoding. Qualified Qwen and GLM speculative paths reduce host synchronization inside an accepted draft cycle.
  • Safer model import. Uncataloged models are checked before download, runnable alternatives are suggested when needed, and local MLX import and quantization can be cancelled cleanly.

Added

  • A universal --context-length override for advanced server and CLI use.
  • Additional headless agent profiles and clearer caller attribution in local telemetry.

Fixed

  • Text-capable vision checkpoints can fall back to the text lane when optional vision dependencies are unavailable. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Version 0.15.4: Beschleunigungsprofile für Qwen3.8 27B und GLM-5.3 Flash

Rapid-MLX 0.15.4 führt experimentelle, eng qualifizierte Beschleunigungsprofile für Qwen3.8 27B (qwen3.8-27b-tensorfold) und GLM-5.3 Flash auf 256-GB-Macs ein und bringt Share Compute in die Desktop-App.

What's new in v0.15.4

Rapid-MLX 0.15.4 adds two narrowly qualified experimental acceleration paths for Qwen3.8 27B and GLM-5.3 Flash, plus a complete Share Compute experience in Desktop. The accelerated profiles pin their model and runtime inputs and fail closed when a request or machine is outside the measured contract.

Highlights

Experimental Qwen3.8 27B acceleration. The new qwen3.8-27b-tensorfold profile pairs one qualified 4-bit checkpoint with its DFlash2 drafter and exposes the active backend and limitations through the server status APIs and Desktop settings. It is an explicit opt-in, admits one text request at a time, and rejects tools, media, grammar constraints, and unqualified model or runtime revisions. The ordinary qwen3.8-27b-4bit path remains available as the feature-complete fallback. This release does not make a fixed speed claim for the experimental profile. (#3929, #3930, #3934)

Experimental GLM-5.3 Flash acceleration on 256 GB Macs. The dedicated glm5.3-flash-tensorfold alias uses the checkpoint's embedded MTP head and selects its qualified accelerated backend by default on eligible systems. Desktop labels it Experimental and provides an opt-out. Startup validates the exact runtime and checkpoint revisions; unsupported tools, media, grammar, and …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.15.4 (Mac): neue Profile, Share Compute und GLM-Chat-Fixes

Rapid-MLX 0.15.4 ergänzt experimentelle beschleunigte Profile für Qwen3.8 27B und GLM-5.3 Flash, Share Compute im Desktop sowie Kernel für hohe Parallelität und behebt Desktop-Chat-Probleme mit dem GLM-TensorFold-Profil.

[0.15.4] — 2026-10-01

Rapid-MLX 0.15.4 adds experimental accelerated profiles for Qwen3.8 27B and GLM-5.3 Flash, and brings Share Compute to the Desktop app.

Added

  • Qualified experimental acceleration. Dedicated Qwen3.8 27B and GLM-5.3 Flash profiles pin the supported runtime and model revisions, expose active status, and fail closed outside their text-only capability boundaries.
  • Share Compute. Desktop can join a live compute pool, show contribution state and settled API credit, stop sharing, and restore the previous model.
  • High-concurrency lane kernels. Qualified dense models can opt into a row-invariant matrix multiplication path for eight or more concurrent rows.

Changed

  • GLM acceleration is easy to try and leave. The dedicated experimental alias enables its accelerated backend on eligible 256 GB systems; Desktop labels it Experimental and offers an explicit opt-out and ordinary fallback.

Fixed

  • GLM TensorFold Desktop chat. The built-in GLM acceleration profile now accepts ordinary Desktop chat requests by omitting unsupported app-provided tools and applying neutral sampling defaults only when controls are untouched. Explicit unsupported settings remain visible and fail closed.
  • Accelerated serving contracts. Qwen admits one accelerated request at a time. Runtime source revisions are verified before load, and GLM streaming preserves code whitespace and private reasoning boundaries. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Version 0.15.3: Überwachtes Computer Use für lokale Mac-Aufgaben

Rapid-MLX 0.15.3 bringt experimentelles, überwachtes Computer Use für lokale Mac-Aufgaben mit Freigabe kritischer Aktionen und eine dokumentierte API für eigene Clients sowie verbesserte Ersteinrichtung und Fehlerbehebung beim Start.

What's new in v0.15.3

Rapid-MLX 0.15.3 adds supervised Computer Use for local Mac tasks and a documented API for clients that provide their own interface. It also improves first-run setup, model serving, and recovery from common startup errors.

Experimental Computer Use

Computer Use remains an opt-in experimental Desktop feature. It can run supervised tasks in a selected macOS app or browser window, shows consequential actions for approval, and stops when its target or approval context changes. Credentials, payment details, and commerce actions are blocked.

Expect incomplete tasks: local planners can stop early or loop on longer work, browser URL verification requires macOS Automation permission, and changing or obscured windows can end a run. Review every approval prompt and be prepared to finish the task manually. The local control API requires authentication; screenshots stay disabled unless both the server and the request opt in.

Highlights

Give Rapid a task and let it find the right app. The Desktop Computer Use panel starts with your task, then resolves a bounded scope from open apps and windows. When several choices fit, Rapid asks you to choose; the active scope stays visible while the task runs. Add or select a local or HTTPS OpenAI-compatible planning Model directly in the panel, then follow progress and approve consequential actions. The Computer Use service starts without loading a chat model. Screen Recording belongs to the Desktop app and …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX Desktop 0.15.3: Computer Use mit Planungsmodell und /v1/cua-API

Rapid-MLX Desktop 0.15.3 fügt aufgabenorientiertes Computer Use mit Auswahl eines Planungsmodells, Hinweisen zu Mac-Berechtigungen, der authentifizierten lokalen /v1/cua-API und anonymen Zählungen zur Ersteinrichtung hinzu.

[0.15.3] — 2026-09-29

Rapid-MLX Desktop 0.15.3 adds supervised Computer Use for general Mac tasks. Describe the task, choose a planning Model, and approve consequential actions; Rapid resolves the app scope and asks you to choose when several scopes fit.

Added

  • Task-first Computer Use. Describe the task first. Rapid resolves a bounded scope from open apps and windows, asks you to choose when several scopes fit, and limits each run to at most three apps or windows. Progress, results, approval, and cancellation stay in the same workspace. Add or select a local or HTTPS OpenAI-compatible planning Model directly in the panel. The Computer Use service starts without loading a chat model.
  • Mac permission guidance. The Desktop app requests Screen Recording while its bundled helper requests Accessibility. Browser tasks can ask macOS for Desktop-to-browser Automation access before app scope is resolved; if authorization is unavailable, the panel explains how to retry.
  • Client API. The authenticated local /v1/cua API supports target discovery, bounded observations, run and event polling, approval decisions, and cancellation for clients with their own interface. Screenshot responses require an explicit server and request opt-in.
  • Anonymous first-run funnel counts. Official Desktop builds now send identifier-free, once-per-install setup milestones for installs whose first …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Version 0.15.2: Automatische Portwahl, neue Modelle, klarere Startfehler

Rapid-MLX 0.15.2 zeigt im Desktop die Verbindungsdaten des lokalen Servers zuerst an, wählt bei belegtem Port 8000 automatisch einen freien Port, unterstützt weitere Modelle wie GLM-5.3 Flash und Laya und liefert klarere Ursachen bei Startfehlern.

What's new in v0.15.2

0.15.2 makes local model serving easier to connect, harder to misconfigure, and much more useful when startup fails.

Highlights

Connect an agent without hunting for the server details. Desktop now puts the local endpoint and copy-ready connection information first, followed by guided integration choices. The server's purpose is visible as soon as it is running.

More capable local model paths. GLM-5.3 Flash can participate in the QuickSilver pool; GLM also gains an explicit default reasoning-effort control. Qwen4 can load validated PLE rows from a bounded sidecar. Laya and CLM System One are supported, and LFM2.5-VL gains a qualified DSpark companion.

The default port no longer turns a healthy launch into a dead end. If port 8000 is occupied and you did not explicitly choose a port, rapid-mlx serve selects the next free port from a bounded range and prints the choice. An explicit --port remains strict, and inherited listeners remain validated.

Missing optional support is actionable. Vision, image, video, and audio startup failures identify the exact extra and can offer to install it. Use --yes for unattended setup.

Startup failures keep their real cause. Hugging Face access failures, invalid configs, tokenizer failures, incompatible weights, quantization mismatches, and insufficient memory now retain stable categories across CLI and Desktop. Fatal server tracebacks are persisted, and the next launch can …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.15.2 (Mac): Agent-Verbindung, GLM-5.3-Reasoning, Absturzdiagnosen

Rapid-MLX 0.15.2 ergänzt eine übersichtliche Agent-Verbindungseinrichtung, weitere lokale Modelle, Reasoning-Steuerung für GLM-5.3, dauerhafte Absturzdiagnosen, automatische Portsuche bei rapid-mlx serve und installierbare optionale Extras direkt aus der Fehlermeldung.

[0.15.2] — 2026-09-24

Rapid-MLX 0.15.2 makes local model serving easier to connect, harder to misconfigure, and much more useful when startup fails.

Added

  • Agent connection setup is visible at a glance. Desktop now leads with the local endpoint and copy-ready connection details, then guides users through popular agent integrations.
  • More capable local models. GLM-5.3 Flash can join the QuickSilver pool, Qwen4 can load validated PLE rows from a bounded sidecar, Laya and CLM System One are supported, and LFM2.5-VL gains a qualified DSpark companion.
  • Reasoning control for GLM-5.3. The server accepts an explicit default reasoning effort and correctly recognises GLM's coercion behavior.
  • Persistent crash diagnostics. Fatal server tracebacks survive process death, and a subsequent start can report that the previous startup ended before reaching a terminal state.

Changed

  • rapid-mlx serve finds a free default port. When no port is specified, the CLI scans a bounded local range instead of failing immediately on 8000; explicit ports still fail closed.
  • Optional features are installable from the failure itself. Missing vision, image, video, or audio support now offers the exact extra to install, with --yes available for unattended setup.

Fixed

  • Startup failures retain their real cause. Hugging Face access failures, invalid model configs, tokenizer assets, incompatible weights, …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Version 0.15.1: Bessere Diagnose bei fehlgeschlagenen Erststarts

Rapid-MLX 0.15.1 macht fehlgeschlagene Erststarts besser diagnostizierbar, indem fehlende optionale Extras mit exaktem pip install-Befehl genannt werden, Ready: erst nach gebundenem Port erscheint und zusätzliche anonyme Startstatus-Telemetrie erfasst wird.

What's new in v0.15.1

0.15.1 makes a failed first start diagnosable — for you and for us.

Highlights

A model that needs an optional extra now tells you which one. Vision, video, image, and audio serve paths now converge on the same actionable failure when their optional runtime is unavailable. The CLI and python -m rapid_mlx.server print the exact pip install 'rapid-mlx[<extra>]' command, while the Desktop startup-failure panel names the closed reason and extra and links to the Startup Log for the installation details.

Ready: means ready. The server prints its ready banner only after the port is actually bound, so a script that waits for the line can connect immediately.

Every serving lane reports success the same way. Image, video, audio, embedding, and specialised text servers now emit the same model_served event as the default lane.

Telemetry: two additions, still anonymous, still closed enums. server_start_state records attempted followed by ready or failed, with a failed stage of resolve, download, preflight, prepare, engine_start, or bind. A missing optional runtime records error_class=missing_extra on model_serve_failed, with extra limited to vision, video, audio, or image. There is no new free text or identifier. Full disclosure: https://rapidmlx.com/docs/telemetry

Caveat: the bundled privacy policy names the new optional-extra field, but the …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.15.1 (Mac): Startfehler sichtbar, Ready erst nach Port-Bindung

Rapid-MLX 0.15.1 zeigt im Desktop den tatsächlichen Startfehler bei fehlenden optionalen Laufzeiten an, gibt Ready: erst nach dem Binden des Ports aus, meldet model_served einheitlich in allen Serving-Lanes und ergänzt geschlossene Startstatus-Telemetrie.

[0.15.1] — 2026-09-23

Rapid-MLX 0.15.1 makes first-start failures actionable in Desktop and aligns readiness and success reporting across every serving lane.

Added

  • Closed startup-state telemetry. Accepted serve invocations now record an attempted state followed by ready or failed; failures carry only a closed startup stage. Missing optional runtimes report a closed extra name rather than free-form text.

Changed

  • Ready: now means the server is accepting connections. The banner is printed only after the listener binds, so scripts waiting for it can connect immediately.
  • Serving success is consistent across lanes. Image, video, audio, embedding, and specialised text servers now emit the same model_served event as the default lane.

Fixed

  • Desktop shows the real optional-runtime startup failure. The startup-failure panel names the closed reason and affected vision, video, audio, or image extra, then opens the Startup Log containing the exact pip install 'rapid-mlx[<extra>]' hint instead of showing only a generic engine-start failure.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Version 0.15.0: Telemetrie-Kontrolle, Qwen-Image 2.1, geringere Leerlauf-CPU-Last

Rapid-MLX 0.15.0 führt datenschutzfreundliche Telemetrie mit lokalen Steuerungsmöglichkeiten (rapid-mlx telemetry off), Bildgenerierung und -bearbeitung mit Qwen-Image 2.1 sowie ein neues High-End-Modell ein und senkt die Leerlauf-CPU-Last im Desktop.

What's new in v0.15.0

Rapid-MLX 0.15.0

Rapid-MLX 0.15.0 adds privacy-safe product telemetry with clear local controls, expands local image generation and editing with Qwen-Image 2.1, and brings a new high-end local model to Server and Desktop. It also reduces idle Desktop CPU use and makes common failures easier to understand and recover from.

Private-by-construction telemetry with visible controls

Official 0.15.0 builds use one versioned telemetry pipeline for coarse product signals such as model use, endpoint use, successful activity buckets, setup events, and explicitly enumerated capability rejections. Events pass a closed schema, use bounded local queues and counters, and are transmitted only from an official release build after the applicable disclosure has been delivered. Raw prompts, responses, file paths, request bodies, API keys, and arbitrary model identifiers are not telemetry properties.

The first Desktop launch shows a one-time notice rather than an interrupting consent dialog. Users can turn telemetry off at any time in Settings or with rapid-mlx telemetry off; status explains the current decision and upload gate, preview prints an example event without sending it, and reset-id rotates the local installation identity. The former v1 collector and its per-request wire are removed, so eligible installs do not double-report.

Qwen-Image 2.1 generation and editing

Qwen-Image 2.1 now has a dedicated Server and Rapid Mac path instead of being…

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.15.0 (Mac): Telemetrie v2, Qwen-Image 2.1, rapid-mlx feedback

Rapid-MLX 0.15.0 bringt Telemetrie v2 mit sichtbaren Kontrollen, einen eigenen Qwen-Image-2.1-Pfad für Text-zu-Bild und Img2img, explizites Gemma 4 Assistant-Sidecar-MTP, den Befehl rapid-mlx feedback und mimo-v2.6-flash-4bit in der Modellauswahl.

[0.15.0] — 2026-09-22

Rapid-MLX 0.15.0 adds privacy-safe product telemetry with clear local controls, expands local image generation and editing with Qwen-Image 2.1, and improves Desktop efficiency, reliability, and failure recovery.

Added

  • Privacy-safe telemetry v2 with visible controls. Official release builds report only validated metadata from a closed registry after the applicable notice has been delivered. The old v1 collector is retired. Users can inspect the live decision and build gate with rapid-mlx telemetry status, preview an exact event without sending it, turn reporting off, or reset the local installation identity. Desktop shows the one-time disclosure and exposes the same control in Settings.
  • Qwen-Image 2.1 generation and editing. The Server and Images workspace use a dedicated Qwen-Image 2.1 runtime for text-to-image and one-image img2img instead of misrouting the family through the incompatible 1.x path.
  • Explicit Gemma 4 assistant-sidecar MTP. Qualified Gemma 4 targets can be paired with a named assistant sidecar without changing ordinary Gemma 4 inference or enabling speculation automatically.
  • A direct feedback route. rapid-mlx feedback and Desktop's Help menu open the community feedback channel without reading telemetry state or attaching diagnostics.
  • Xiaomi MiMo-V2.6 Flash is in the model picker as mimo-v2.6-flash-4bit …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Version 0.14.3: Schnellere multimodale Folgefragen und Community Benchmark

Rapid-MLX 0.14.3 beschleunigt Folgefragen in multimodalen Unterhaltungen deutlich (z. B. 54,8 % niedrigere TTFT beim zweiten Durchgang), erweitert Personal Intelligence um begrenzte lokale Arbeit, bringt einen vollständigen Community Benchmark im Desktop und unterstützt Bonsai-2-Hadamard-Modellpakete.

🖥️ Desktop app (.dmg): https://rapidmlx.com/download

What's new in v0.14.3

Rapid-MLX 0.14.3

Rapid-MLX 0.14.3 makes repeat multimodal conversations materially faster, extends Personal Intelligence from answers into bounded local work, and gives Community Benchmark a complete, trustworthy Desktop experience. It also fixes the downloaded DMG install window on affected macOS releases and adds native support for Bonsai 2 Hadamard model packs.

Faster multimodal conversations

Two default-on, fail-closed optimizations remove work from the serialized multimodal lane without changing outputs outside their qualified boundary.

Qualified Qwen3.6 35B media workload Before After
Second-turn media-prefix TTFT baseline 54.8% lower
Second-turn elapsed time baseline 22.4% lower
Singleton-lane generation throughput baseline 35.4% higher

The media-aware prefix cache retains the exact prior-turn media boundary, so a follow-up question about the same image need not encode the image and prefill the unchanged conversation again. Its one-time storage cost was 10–25 ms on short eligible first turns and cost-neutral on long prompts. The singleton fast path separately skips merging and repacking a cache whose shape is already correct; the measured 21-case suite kept identical response hashes and peak memory. Both paths retain explicit rollback settings. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.14.3 (Mac): Personal Intelligence mit lokalem Dateizugriff

Rapid-MLX 0.14.3 erlaubt Personal Intelligence den begrenzten Zugriff auf lokale Dateien und Code mit Freigabe pro Aktion, bietet einen vollständigen Community Benchmark im Desktop, startet Folgefragen zum selben Bild schneller und behebt das DMG-Installationsfenster unter betroffenen macOS-Versionen.

[0.14.3] — 2026-09-18

Rapid-MLX 0.14.3 makes multimodal conversations materially faster, gives Personal Intelligence bounded access to local files and code, and rebuilds the Community Benchmark experience around live, trustworthy evidence. It also fixes the downloaded DMG install window on affected macOS versions.

Added

  • Personal Intelligence can work with local files and code. From Desktop chat, qualified local models can search and read text, write reviewed files, move a file to recoverable Trash, and compile or run a bounded development command. Every local-data action shows its exact path and arguments for approval; mutation permission is never remembered, networking is denied, and hidden or credential-bearing locations remain unavailable.
  • Community Benchmark is now a complete Desktop experience. Runs expose live stages, pass counts, ETA, and the latest measurement; My Results and Community views preserve contributor identity across restarts; comparison copy distinguishes bounded recent evidence from complete totals; and publication now verifies immutable build provenance before accepting a run.
  • Bonsai 2 Hadamard model packs now run natively. Rapid-owned loading for packed Hadamard projections and inverse embeddings covers the qualified text, streaming, and image-serving paths without changing existing MLLM behavior.

Changed

  • Follow-up questions about the same image start much sooner. The serialized …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Version 0.14.2: Deutlich schnellere Modelle und Agent Mode im Desktop

Rapid-MLX 0.14.2 beschleunigt lokale Modelle deutlich (z. B. DeepSeek V4.1 Flash 2,02×, Qwen3.6 35B mit nativem MTP +57,1 %), führt einen freigabebasierten Agent Mode im Desktop ein und unterstützt experimentell K2 Horizon 7B.

What's new in v0.14.2

Rapid-MLX 0.14.2

Rapid-MLX 0.14.2 makes local models materially faster and introduces a bounded, approval-first Agent Mode in Desktop. It also adds an experimental native K2 Horizon 7B path, productizes the qualified 256 GB DeepSeek V4.1 lane, and fixes costly prompt-cache and tool-loop failures found in real Desktop use.

Faster local inference

The largest measured improvements in this release are summarized below. Each path is shape- and runtime-gated and retains its existing fallback outside the qualified boundary.

Model / workload Before After Change
DeepSeek V4.1 Flash, mixed DSpark K4 9.58 tok/s 19.39 tok/s 2.02×
Qwen3.6 35B, native MTP fixed suite 83.32 tok/s 130.93 tok/s +57.1%
GLM-5.3 Flash, six real tasks 26.58 tok/s 35.54 tok/s +33.7%
Qwen3.6 35B, compiled no-MTP HTTP decode 101.7 tok/s 121.6 tok/s +19.6%
Qwen3.6 35B, fused GDN HTTP path 62.3 tok/s 79.1 tok/s +27.0%

DeepSeek V4.1 remains an explicitly experimental, text-only lane for a 256 GB Mac Studio. Its pinned 2-bit target plus mixed-precision sidecar peaked at 218.23 GB; it is not a recommendation for ordinary Macs. GLM-5.3 preserved all 12 reasoning traces and final answers byte-for-byte in its paired qualification. Qwen3.6 kernels and compiled replay fail closed to the stock route when the …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.14.2 (Mac): Agent Mode, K2 Horizon 7B, DeepSeek V4.1 Flash

Rapid-MLX 0.14.2 ergänzt den experimentellen Agent Mode in Desktop und Server, experimentelle Unterstützung für K2 Horizon 7B und DeepSeek V4.1 Flash auf 256-GB-Macs, einen Hinweis auf Änderungen nach einem Update und beseitigt mehrere Re-Prefill- und Tool-Loop-Probleme im Desktop-Chat.

[0.14.2] — 2026-09-14

Rapid-MLX 0.14.2 focuses on faster local inference and a safer first Agent Mode. It adds an experimental native K2 Horizon runtime, productizes the qualified DeepSeek V4.1 lane for 256 GB Macs, and removes several expensive re-prefill and tool-loop failure modes from Desktop chat.

Added

  • Experimental Agent Mode in Desktop and Server. A bounded local agent loop can use configured MCP tools, pause before consequential actions, show a redacted action summary, and resume only after approval. Runs are authenticated, process-local, resource-bounded, and fail closed if the tool registry changes underneath them.
  • Experimental K2 Horizon 7B support. Server, Desktop, model discovery, reasoning parsing, and tool calls now share a Rapid-owned native adapter. On an M4 Pro 48 GB Mac it measured 49–51 tok/s through shipped paths, roughly matching Qwen3.5 9B decode speed while using slightly less peak memory.
  • Experimental DeepSeek V4.1 Flash support for 256 GB Macs. The pinned 2-bit target and mixed-precision DSpark sidecar now have a bounded, serial, deterministic K4 serving path. It measured 19.39 tok/s versus 9.58 tok/s autoregressive (2.02×) with a 218.23 GB peak; the alias remains text-only and explicitly experimental.
  • What changed after an update. The first launch on a new version shows a one-line "Updated to vX.Y.Z" notice with a link to that release's notes. …

Originalquelle(öffnet in neuem Tab)Problem melden