Zum Inhalt springen

Rapid-MLX Release Notes

32 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Folge Rapid-MLX, um die Release Notes in deinen Feed zu holen.

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Version 0.14.1: Analyse großer PDFs und zuverlässigere multimodale Chats

Rapid-MLX 0.14.1 ermöglicht dem Desktop die Analyse großer und gescannter PDFs mit begrenztem Cache und gezieltem Abruf von Abschnitten, macht multimodale Chats zuverlässiger und verstärkt die Regressionstests für sampled MTP.

What's new in v0.14.1

Rapid-MLX 0.14.1

Rapid-MLX 0.14.1 makes long-document work and multimodal chat more dependable, while adding stronger regression protection for sampled MTP decoding. It also records an intentionally negative large-model qualification so users can see why an expensive checkpoint was not added to the product catalog.

Large and scanned PDFs

Desktop can now analyze large selectable PDFs, image-only scans, and documents that mix selectable and scanned pages. Extraction runs behind a bounded cache, document reads use hard deadlines, and follow-up questions can retrieve specific outlines, pages, or sections instead of forcing the entire file into one prompt. (#3292)

The product-path dogfood covered both ends of the workflow:

Document Verified result
300-page selectable Chinese PDF 125,833 characters cached; all 300 pages completed; 30 chapter rows resolved to real page and character offsets
6-page image-only PDF OCR text reached the model and multi-turn synthesis decoded at 45 tok/s
Mixed selectable + scanned PDF Scanned pages were retained and the extraction reached an honest complete state

The same change closes several failure boundaries: OCR/render failures no longer masquerade as complete extraction, progress heartbeats cannot extend a read forever, removal races cannot publish an orphaned document, invalid IDs …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.14.1 (Mac): PDF-Analyse und Fix für HTTP-500 bei Wiederholungsschleifen

Rapid-MLX 0.14.1 bringt die Analyse großer und gescannter PDFs im Desktop, stärkere Regressionstests für sampled MTP und behebt, dass Wiederholungsschleifen in multimodalen Antworten den Metal-Speicher erschöpfen und mit leerem HTTP 500 enden.

[0.14.1] — 2026-09-10

Rapid-MLX 0.14.1 is a focused reliability update for document analysis, multimodal chat, and sampled speculative decoding.

Added

  • Large and scanned PDF analysis. Desktop can analyze long selectable PDFs, image-only scans, and documents that mix both forms. Extraction continues in a bounded background cache, document reads have hard deadlines, and follow-up questions can retrieve the relevant page ranges instead of placing an entire large document in one prompt.
  • Stronger sampled-MTP regression coverage. Distribution-level and real-weight checks now protect independent acceptance draws and rejection- residual sampling, including bugs that the prior 246-test MTP suite could miss.

Fixed

  • Multimodal repetition no longer exhausts Metal. Exact token loops are stopped at the scheduler boundary, the completed row is retired immediately, and the valid partial response finishes normally instead of ending in a delayed empty HTTP 500.
  • PDF extraction now reports incomplete OCR honestly, respects cancellation and removal races, bounds malformed document identifiers, and forces synthesis when a model repeatedly exceeds the document/tool budget.

Qualification

  • A 199 GiB experimental DeepSeek V4.1 Flash REAP 2-bit checkpoint loaded in 239.44 seconds with 213.51 GB peak MLX memory on a 256 GiB M3 Ultra, but decoded at only 7.31–7.92 tok/s. It remains outside the model catalog, …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Version 0.14.0: Schnelleres Decoding, Computer Use, Share Compute, größerer Katalog

Rapid-MLX 0.14.0 steigert auf einem M3 Ultra gemessen die Decode-Geschwindigkeit (z. B. +14,0 % bei Qwen3.8 27B MTP GDN, 2,00× bei Muse-Glimmer 30B DFlash für Code) und senkt die TTFT einer kurzen Anfrage hinter drei langen Prompts von 3,067 auf 1,206 s, außerdem kommen experimentelle Computer-Use- und Share-Compute-Workspaces, ein klarerer Community Benchmark, One-Command- und Always-on-Serving, One-Click-Modell-Unload und ein größerer Text- und Bildkatalog hinzu.

What's new in v0.14.0

Rapid-MLX 0.14.0

Rapid-MLX 0.14.0 makes a Mac faster and more useful as a private local-AI system. It ships measured decode and scheduling improvements, experimental Computer Use and Share Compute workspaces, a clearer Community Benchmark, one-command and always-on serving, one-click model unload, and a substantially larger text and image catalog.

Measured performance improvements

These are product-path measurements, not synthetic peak-kernel claims. Results identify the tested Mac and preserve opt-in/default boundaries; other hardware, prompts, and sampling settings can differ.

Path Before 0.14.0 result Measured change
Qwen3.8 27B MTP GDN verification, M3 Ultra 45.87 tok/s 52.28 tok/s +14.0% decode; verification synchronization 4.954 → 4.644 s
Muse-Glimmer 30B 8-bit DFlash, M3 Ultra 22.1–22.6 tok/s 41.3–46.3 tok/s 2.00× code median, 1.83× chat, 1.94× overall
Qwen3.8 Flash-Next 4-bit fused GDN, M3 Ultra 25.43 tok/s 27.04 tok/s +6.35% end-to-end decode; isolated GDN layer +26.28%
Short request behind three long prompts, M3 Ultra 3.067 s TTFT 1.206 s TTFT 60.7% lower TTFT / 2.54× faster

The Qwen3.8 MTP result used the same checkpoint, prompt, seed, K=3 draft depth, and 256-token output in three alternating runs; output hashes, acceptance rate, and verify-round count matched. The Muse-Glimmer acceleration is enabled only …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.14.0 (Mac): Computer Use, Share Compute, erweiterter Modellkatalog

Rapid-MLX 0.14.0 (Mac) bringt experimentelles lokales Computer Use und Share-Compute-Workspaces, erweitert den Modell- und Bildgenerierungs-Katalog (u. a. Granite 4.2, FLUX.1 schnell, Stable Diffusion 3.5 Large), macht den Community Benchmark verständlicher und verbessert den Dauerbetrieb des Dienstes sowie die Desktop-Steuerung.

[0.14.0] — 2026-09-09

Rapid-MLX 0.14.0 adds experimental local Computer Use and shared-compute workspaces, expands the model and image-generation catalog, makes Community Benchmark easier to understand and share, and improves long-running service operation and Desktop control.

Measured on the documented M3 Ultra workloads, Qwen3.8 27B MTP GDN verification improves decode from 45.87 to 52.28 tok/s (+14.0%), the qualified Muse-Glimmer 30B 8-bit DFlash pairing reaches 1.83× chat and 2.00× code throughput, and the opt-in short-request policy lowers TTFT behind three long prompts from 3.067 s to 1.206 s. These paths retain their documented hardware and opt-in boundaries.

Added

  • Experimental Computer Use in Desktop. Starter workflows can carry out a bounded task across supported Mac apps, with local planning, scoped control, visual recovery, and explicit review before consequential actions.
  • Experimental Share Compute workspace. A Mac can opt in to serve work from a compatible compute pool, with fit checks, loopback-safe configuration, and clearer diagnostics. It remains off until the user enables it.
  • Broader model support. This release adds NeoHorse 1 9B, MiniCPM5 2B, G9v3-39A5B, Granite 4.2 30B/8B/3B, an experimental Qwen3.8 27B Abliterated 4-bit variant, and a Qwen3.8 27B MTP companion, alongside new local image-generation choices including FLUX.1 schnell, Stable Diffusion 3.5 Large, SDXL Base, Bonsai Image 4B 2-bit, HiDream O1 Dev, and…

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.13.4: MTP-Preset für Qwen, Video-Workspace, privater Benchmark

Rapid-MLX 0.13.4 beschleunigt qualifizierte Qwen-Modelle bei paralleler Last standardmäßig per MTP-Preset (+14,1 % bis +30,8 % Durchsatz, abschaltbar mit --no-spec-decode), ergänzt einen experimentellen Video-Workspace und macht Community Benchmark standardmäßig privat mit explizitem Teilen.

What's new in v0.13.4

Rapid-MLX 0.13.4

Rapid-MLX 0.13.4 makes qualified local models faster under concurrent work, adds an experimental end-to-end video workspace, and turns Community Benchmark into a private-by-default local workspace with explicit sharing. It also improves memory fitting, model discovery, Markdown rendering, and headless operation.

Highlights

Faster qualified Qwen models by default

The four exact qualified Qwen3.5/Qwen3.6/Qwen3.8 artifacts now select their validated MTP preset and continuous scheduler automatically when no speculative-decoding option is supplied. Across mixed four-request cohorts, aggregate throughput improved by 14.1%–30.8%, with no ordinary-pass case becoming a continuous-mode failure. CLI and Server users can restore ordinary decoding with --no-spec-decode; Desktop exposes the same persistent off switch.

Qualified artifact Aggregate throughput change
Qwen3.5 4B 4-bit +30.8%
Qwen3.5 9B 4-bit +23.9%
Qwen3.6 27B 4-bit +14.1%
Qwen3.8 27B 4-bit +25.9%

For Qwen3.8 27B, the qualified single-request MTP path also scales with context: measured decode throughput was 1.43× the 0.13.3 ordinary path at 128 tokens of context and 2.34× at 32K. The new continuous scheduler keeps request/cache ownership transactional across admission, cancellation, and dynamic joins.

Local video generation in Desktop …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.13.4 (Mac): Videogenerierung und schnellere MTP-Modelle

Rapid-MLX 0.13.4 (Mac) bringt lokale Videogenerierung mit LTX 2.5, Wan 2.1 und CogVideoX-Fun im Desktop, einen standardmäßig privaten Community Benchmark, eine einheitliche Modellerkennung sowie standardmäßig schnellere qualifizierte MTP-Modelle mit 14,1–30,8 % mehr Durchsatz.

[0.13.4] — 2026-09-02

Rapid-MLX 0.13.4 makes qualified local models faster under concurrent work, adds an experimental end-to-end video workspace, and turns Community Benchmark into a private-by-default local workspace with explicit sharing.

Added

  • Local video generation in Desktop. An opt-in Video workspace supports LTX 2.5, Wan 2.1, and CogVideoX-Fun with queued jobs, progress, cancellation, restart-safe results, and early memory validation.
  • Private-by-default Community Benchmark. CLI and Desktop users can run, inspect, and archive reproducible measurements locally; sharing requires an explicit preview and consent, with duplicate submission protection.
  • Atomic model discovery and recommendations. CLI, Server, and Desktop now consume one validated model registry and recommendation policy.
  • Opt-in persistent conversation memory, isolated Mermaid previews, per-model request metrics, and clearer in-app update discovery.

Changed

  • Qualified MTP models are fast by default. Qwen3.5 4B/9B, Qwen3.6 27B, and Qwen3.8 27B automatically use their validated continuous speculative scheduler in CLI, Server, and Desktop. Desktop shows the active setting and keeps a persistent off switch for users who prefer ordinary decoding.
  • Across mixed four-request cohorts, the qualified paths improved aggregate throughput by 14.1%–30.8%. Qwen3.8 27B decode measured 1.43× the 0.13.3 ordinary path at 128 tokens of context and 2.34× at 32K. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.13.3: GLM-5.3-Flash und Text-Prefixes bei Vision-Modellen

Rapid-MLX 0.13.3 unterstützt GLM-5.3-Flash (glm5.3-flash-4bit, benötigt die 192-GB-Speicherstufe) im Serving-Pfad, nutzt Text-Prefixes auf hybriden Vision-Modellen wieder und verbessert Desktop-Navigation, Video-Workflows, Streaming-Ausgabe und Fehlerbehebung.

What's new in v0.13.3

Rapid-MLX 0.13.3

Rapid-MLX 0.13.3 brings GLM-5.3-Flash to the production serving path, makes hybrid multimodal conversations reuse eligible text prefixes, and improves the Desktop workflows around navigation, video generation, streaming output, and failure recovery.

Highlights

GLM-5.3-Flash support

Use glm5.3-flash-4bit from the CLI, OpenAI-compatible server, or Desktop. The release includes the processor and runtime compatibility required by the checkpoint, preserves explicit image-channel layouts, rejects malformed media before model execution, and keeps speculative decoding disabled because its qualification run did not improve throughput.

On a 256 GB M3 Ultra, the qualified 4-bit checkpoint decoded a sustained 512-token response at a median 29.2 tok/s and used 165.4 GB of active MLX memory. The alias therefore requires the 192 GB memory tier. The README's benchmark link records the exact checkpoint revision, request payload, and reproduction command.

Faster repeated work on hybrid vision models

Compatible text-only phases on a hybrid multimodal lane can reuse their text prefix state instead of recomputing it. Cache identity includes the rendered request and model state, while media-bearing requests continue through the validated vision path. This improves repeated background and conversation work without treating an image request as a text-only cache hit.

Desktop workflow improvements …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.13.3 (Mac): GLM-5.3-Flash lokal und Befehlspalette

Rapid-MLX 0.13.3 (Mac) führt GLM-5.3-Flash lokal über den normalen Serving-Pfad ein, ergänzt im Desktop eine native Befehlspalette samt Feedback-Eintrag und Diagnose-Shortcut und lässt hybride Vision-Konversationen passende Text-Prefixes wiederverwenden.

[0.13.3] — 2026-08-31

Rapid-MLX 0.13.3 adds production-safe GLM-5.3-Flash serving, reuses text prefixes on hybrid vision models, and improves Desktop navigation, video-job recovery, streaming presentation, and failure diagnosis.

Added

  • GLM-5.3-Flash runs locally through the normal serving path. The new glm5.3-flash-4bit alias includes the processor, quantization, image-layout, and runtime compatibility needed for text and multimodal requests. The qualified 4-bit checkpoint decoded a sustained 512-token response at a median 29.2 tok/s on an M3 Ultra and requires the 192 GB memory tier.
  • Desktop commands are easier to reach. A native command palette exposes common actions from the keyboard, completed interactions can offer a lightweight feedback entry, and failure diagnostics have a direct shortcut.

Changed

  • Hybrid vision conversations reuse eligible text prefixes. Repeated text-only phases on the multimodal lane retain compatible prefix state instead of recomputing it, while media-bearing requests continue through their validated vision path.
  • Large-model performance claims are reproducible. The README now links exact M3 Ultra workloads, checkpoint revisions, memory measurements, and ordinary-versus-MTP results for Qwen3.8 and GLM-5.3-Flash.
  • GUI regression coverage begins moving deterministic journeys into the Swift test process, shortening the expensive hosted-Mac merge lane without …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.13.2: Natives MTP, Prefix-Cache und mehr Desktop-Sicherheit

Rapid-MLX 0.13.2 bringt optionales natives MTP und schnelleren Long-Context-Prefill für Qwen3.8 Flash-Next, stellt den Prefix-Cache bei wiederholten Prompts wieder her, bündelt Laufzeitanforderungen bei Offline-Speech-Pulls und verschärft die Desktop-Sicherheit bei Anhängen, Zugangsdaten und Web-Browsing.

What's new in v0.13.2

Rapid-MLX 0.13.2 makes long-running local assistants faster and more dependable: Qwen3.8 Flash-Next gains opt-in native MTP and faster long-context prefill, repeated prompts recover their prefix cache, offline speech pulls include their runtime requirements, and Desktop tightens attachment, credential, and web-browsing safety. This stable release also includes the fixes validated after rc1 and replaces the stable Desktop updater feed.

Final release validation

  • The protected Desktop publication path now promotes the exact signed and notarized candidate bytes, including the canonical DMG and updater payloads, rather than rebuilding after the release tag. (#2775)
  • On memory-constrained Macs, choosing a photo with a model whose text lane is still usable now explains that text chat remains ready and recommends a lower-memory vision model. The notice clears after the user continues with a text turn, changes model capability, or chooses another attachment path. (#2778)

Highlights

Qwen3.8 Flash-Next gains native MTP — The engine can opt into the checkpoint's one-layer prediction head while the target model verifies every proposal and all recurrent, QSA, and KV state rolls back atomically. The measured fixed-K1 workload accepted 76.41% of proposals. (#2572, …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.13.2-rc1: Release Candidate mit nativem MTP für Qwen3.8

Rapid-MLX 0.13.2-rc1 ist ein erster Release Candidate zur Validierung mit optionalem nativem MTP für Qwen3.8 Flash-Next (+36 bis +42 % Decode), schnellerem Prefill und mehr Desktop-Sicherheit, der den stabilen Updater-Feed nicht ersetzt.

What's new in v0.13.2-rc1

Rapid-MLX 0.13.2-rc1 makes long-running local assistants faster and more dependable: Qwen3.8 Flash-Next gains opt-in native MTP and faster long-context prefill, repeated prompts recover their prefix cache, offline speech pulls include their runtime requirements, and Desktop tightens attachment, credential, and web-browsing safety. This first release candidate is published for validation and deliberately does not replace the stable updater feed.

Highlights

Qwen3.8 Flash-Next gains native MTP — The engine can opt into the checkpoint's one-layer prediction head while the target model verifies every proposal and all recurrent, QSA, and KV state rolls back atomically. The measured fixed-K1 workload accepted 76.41% of proposals. (#2572, #2655)

Context Serial decode Native MTP Change
128 25.17 tok/s 34.85 tok/s +38.5%
2K 23.64 tok/s 33.53 tok/s +41.8%
8K 22.82 tok/s 32.20 tok/s +41.1%
32K 21.16 tok/s 28.82 tok/s +36.2%

Desktop's normal temperature, top-p, top-k, min-p, and penalty settings can use the lane. Seeded requests and stateful grammar, tool, reasoning, or suppression processors continue on ordinary decoding rather than silently weakening their …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.13.2 (Mac): MTP, kleinerer Download, Folgevorschläge

Rapid-MLX 0.13.2 (Mac) ergänzt optionales natives MTP für Qwen3.8 Flash-Next (36–42 % mehr Decode-Durchsatz), Konversationstitel mit Folgevorschlägen und einwilligungsbasierte Aktivierungsmeilensteine, verkleinert den Desktop-Download um 43,53 % und wählt Erstchat-Empfehlungen nach verfügbarem Speicher.

[0.13.2] — 2026-08-30

Rapid-MLX 0.13.2 makes long-running local assistants faster and more reliable, improves offline speech and Desktop safety, and promotes the exact signed Desktop candidate that passed release validation.

Added

  • Opt-in native MTP for Qwen3.8 Flash-Next. Target verification and atomic recurrent-state rollback raise measured decode throughput by 36–42% across 128-token through 32K-context workloads while constrained requests retain ordinary decoding.
  • Conversation titles and follow-up suggestions. Completed local chats can derive a short title and offer three optional next steps without replacing a user rename or displaying malformed output.
  • Privacy-bounded activation milestones. Desktop records the first successful chat, dictation, and generated image only after explicit consent.

Changed

  • Flash-Next long-context prefill is faster and reusable. Batched QSA index-cache construction reduced measured 2K, 8K, and 32K time to first token by 28.9–32.5%, and semantic snapshots preserve reusable recurrent state through batching and persistence.
  • The Desktop download is smaller. LZMA packaging and dependency-proven pruning reduced the signed comparison DMG by 43.53% while retaining the release contract.
  • First-chat recommendations follow available memory. New installs choose a smaller default below 16 GB and prefer an eligible cached model in the same memory tier.

Fixed …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Rapid-MLX

Rapid-MLX 0.13.2-rc1 (Mac): Natives MTP und schnellere Prefills

Rapid-MLX 0.13.2-rc1 (Mac) bringt optionales natives MTP für Qwen3.8 Flash-Next (Decode von 21,16–25,17 auf 28,82–34,85 tok/s), einwilligungsbasierte Desktop-Aktivierungsmeilensteine sowie Konversationstitel und Folgevorschläge und schnellere Long-Context-Prefills.

[0.13.2-rc1] — 2026-08-30

Added

  • Opt-in native MTP for Qwen3.8 Flash-Next. The engine can use the checkpoint's one-layer prediction head with target verification and atomic rollback. On the measured 128-token through 32K-context workloads, decode rose from 21.16–25.17 tok/s to 28.82–34.85 tok/s with 76.41% proposal acceptance. Desktop's normal temperature and penalty settings are supported; seeded or stateful constrained requests continue on ordinary decoding. (#2572, #2655)
  • Consented Desktop activation milestones. Rapid can report the first successful chat reply, dictation, and generated image after explicit opt-in. Successful Messages and Completions requests also carry privacy-bounded surface and client attribution; raw user-agent text is never emitted. (#2428, #2436)
  • Conversation titles and next-step suggestions. After the first completed exchange, Desktop can derive one short local title without overwriting a user rename. Settled text answers can also offer three optional follow-up prompts; malformed, duplicate, or incomplete suggestions remain hidden. (#2698)

Changed

  • Long-context Flash-Next prefills are faster. Batched QSA index-cache …

Originalquelle(öffnet in neuem Tab)Problem melden