Zum Inhalt springen

MTPLX Release Notes

32 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Folge MTPLX, um die Release Notes in deinen Feed zu holen.

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.5.2: Hotfix gegen langsame lange Antworten

MTPLX 2.5.2 behebt als Hotfix die Verlangsamung langer Antworten aus 2.5.1, indem der spekulative Verifier sauber vom kompilierten in den Eager-Modus übergeht, bei identischer Ausgabe und mit MLX 0.32 als neuer Mindestversion.

4 August 2026 · build 25200

Hotfix for the 2.5.1 long-response slowdown. Long answers no longer start fast and decay: the speculative verifier now hands off cleanly when a response outgrows its compiled window.

  • Long responses hold their speed. The compiled-to-eager verifier handoff settles all state once at the ownership boundary instead of dragging unfinished GPU work through the rest of the answer.
  • Identical output. The fix changes execution order only. Tokens, acceptance, and peak memory are unchanged, and the handoff is visible in verifier stats so it cannot regress silently.
  • Faster MLX for existing installs. The minimum MLX version is now 0.32, converging older runtime environments to the stack fresh installs already run.

Download DMG · 56 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.5.1: Qwen 3.6 27B Speed V2 empfohlen

MTPLX 2.5.1 empfiehlt auf modernen Macs mit mindestens 32 GB Speicher als Erstes das Modell Qwen 3.6 27B Optimized Speed V2, das mit dynamischer 4-Bit-Hybrid-Quantisierung die Qualität beim Coding deutlich erhöht, und behebt eine Cold-Cache-Race-Condition.

3 August 2026 · build 25100

The coding-default release. Qwen 3.6 27B Optimized Speed V2 is now the first recommendation on modern Macs with at least 32 GB of unified memory. The original Optimized Speed model stays directly below it.

  • Much higher quality for coding. V2 uses dynamic 4-bit hybrid quantization with hand-tuned sensitive parts kept at up to 16-bit.
  • Built for longer agent work. V2 gets stronger as coding tasks become longer. It is slightly larger and can be a little slower for short chats.
  • First-class everywhere. Onboarding, the app picker, CLI defaults, quickstart, downloads, inspection, turbo profiles, and OpenCode all use one model identity.
  • Memory-aware recommendations. Modern 32 GB Macs get V2 first. Smaller Macs keep the existing 9B and 4B recommendations.
  • Focused and regression-gated. Runtime kernels, sampler defaults, and speculative depth are unchanged from 2.5.0. A cold-cache bookkeeping race is fixed, and the full Python, Swift, and signed-app gates passed.

Download DMG · 56 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.5.0: Qwen-3.8-Vorbereitung, HY3 und DeepSeek V4

MTPLX 2.5.0 bereitet die Unterstützung von Qwen 3.8 vor, macht Tool-Calls in Coding-Agent-Bridges bei mehrteiligen Änderungen zuverlässiger, bindet HY3 nativ ein und bringt einen experimentellen, optionalen schnellen Pfad für DeepSeek V4.

3 August 2026 · build 25000

The next-model readiness release. MTPLX is prepared for the shape of Qwen 3.8, coding-agent bridges survive real multi-file work, and new experimental performance lanes land without changing safe defaults.

  • Prepared for Qwen 3.8. Multi-layer MTP heads and checkpoint-declared architecture classes remove the known integration blockers; real-weight validation begins when the weights publish.
  • Tool calls survive real work. OpenCode CLI completed a ten-action multi-file change with 31 tests passing, OpenCode Desktop ran visibly, and Pi, Hermes, OpenAI, and Anthropic tool paths were exercised end to end.
  • HY3 is first-class. Native MTP loading, safe AR fallback, model discovery, think-tag handling, and OpenCode tool calls are wired together.
  • DeepSeek V4 gets an experimental fast path. David Tai's shape-specialized optimizations reached about 36 MTP tok/s on the tested 128 GB Mac. The lane stays opt-in while agent-quality calibration continues.
  • No measured Qwen V2 regression. Alternating baseline/candidate runs were flat within thermal/order noise for decode, prefill, and memory, followed by the full Python and Swift suites.

Download DMG · 56 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.4.2: Warme Sessions bleiben, Request-Log, DeepSeek-V4-Flash

MTPLX 2.4.2 verhindert unter Coding-Agenten das Verdrängen warmer Sessions und wirkungslose Hintergrund-Commits, ergänzt ein standardmäßig aktives Request-Log, ein experimentelles DeepSeek-V4-Flash-Backend und eine gegen den Code geprüfte Dokumentation.

2 August 2026 · build 24200

The agentic-cache release. If you run MTPLX under a coding agent, this is about the slowdowns you could feel but not see, each one a real, named mechanism, each one fixed or fenced.

  • Your warm session stops being evicted mid-run. Recently-active sessions are eviction-last under cross-session pressure, and divergent per-turn snapshots are bounded per session, so a long agent run keeps its warm state instead of paying a surprise full re-prefill.
  • Tool-turn commits stop being ghosts. A background commit that could burn a 26-second full-history re-forward without ever storing is fenced off, a bounded grace window lets nearly-finished commits land under fast agent loops, and OpenCode's per-request session headers are honored so consecutive turns stop being treated as strangers.
  • Every serve keeps a durable trail. A default-on, content-free request log (rotating, disable with one env) plus the 2.4.1 bit-exact capture make agent-session incidents diagnosable after the fact.
  • Experimental DeepSeek-V4-Flash backend. A from-scratch native port with an optional speculative lane, gated on committed-sequence identity, thanks to davidtai.
  • The documentation tells the truth. ~450 claims audited against the code; every confirmed drift fixed, from a phantom MLX fork to a wrong Anthropic base URL.

Download DMG · 56 MB Full release notes …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.4.1: Flüssigeres Streaming und besseres Rendering im Chat

MTPLX 2.4.1 macht das Streaming im Chatfenster auch bei langen Antworten flüssig, rendert Code, Tabellen und Formeln besser, behebt die Verlangsamung kurzer Anfragen aus 2.4.0 und zeigt das tatsächlich geladene Modell korrekt an.

1 August 2026 · build 24100

The smooth-streaming release. The chat window stops stuttering and bouncing, streamed markdown grows up, and the 2.4.0 short-turn regression is fixed.

  • Streaming stays smooth on long answers. Finalized lines fold into segments so the view count stays bounded, the scroll pin runs in the same display cycle as layout so the bubble can never visibly bounce, a hidden ~50 ms per-update sizing walk is gone, and new text reveals with typewriter pacing that always keeps up.
  • Code, tables, and math render properly. Live syntax coloring across twelve languages from a freeze-time lexer with O(new text) cost, streaming code cards that settle in place, real tables, and real math notation, Unicode superscripts, stacked matrices and fractions, no leaked dollar signs. Performance mode remains a true plain-text kill switch.
  • 2.4.0 short-turn regression fixed. The compiled-verify path could reserve KV budget above the configured ceiling, taxing short requests with setup work they never used. If short turns felt slower on 2.4.0, this was why. Warming prefills also now yield to real traffic within one small chunk.
  • The model chip tells the truth. A derivative artifact whose folder name extends a first-party model name is served under its own id, health payload, OpenAI model field, and app chip all report what is actually loaded. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.4.0: Schnelleres 35B-A3B, Lüfterfix, bessere Tool-Calls

MTPLX 2.4.0 beschleunigt das 35B-A3B-Modell durch einen kompilierten Decode-Stack, behebt dauerhaft auf Maximum festhängende Lüfter (#201) und verbessert Tool-Calls mit korrektem finish_reason sowie weiteren unterstützten Dialekten.

31 July 2026 · build 24000

The 35B speed release. The 35B-A3B MoE gets a compiled decode stack, the fan bug from 2.3.0 is dead, and tool calling gets another round of contract hardening.

  • 35B-A3B compiled decode stack. Target-prefix compiled route, whole-MoE fusion, GDN post-conv fusion, and a row-owned router, plus continuous batched serving with fixed-shape cohorts and ragged KV for the concurrent lane. Contributed by davidtai.
  • Fans no longer stay stuck at max (#201). A failed fan restore was silently treated as restored while the hardware stayed pinned. Restores now verify the fans are back on the Apple auto curve and retry with backoff until they are, and a watchdog drops any fan lease held while the engine sits idle.
  • Honest finish_reason on cut tool calls. A generation cut by max_tokens mid-tool-call reports length instead of tool_calls, so agent clients continue the turn instead of executing a truncated call. The think-splitter also stops leaking reasoning into content on bare function= strings.
  • More tool-call dialects. Bracket-style and Poolside arg_key/arg_value calls parse correctly, incomplete calls buffer instead of double-delivering, and calls to undeclared tools pass through per the OpenAI contract. Contributed by davidtai. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.3.0: Strukturierte JSON-Ausgabe und Tool-Call-Korrekturen

MTPLX 2.3.0 behebt beschädigte Tool-Call-Argumente, erzwingt strukturierte JSON-Schema-Ausgabe per Grammatikmasken bei voller Speculative-Decoding-Geschwindigkeit, erhält Sessions samt warmem Cache nach Verlaufsänderungen und fängt hängende Streams per Watchdog ab.

21 July 2026 · build 23000

The agent reliability release. Tool calls stop corrupting, structured output lands at full speculative speed, and agent sessions stop paying re-prefill taxes.

  • Tool-call arguments no longer collapse to {}. The intermittent argument corruption behind failed agent edits is root-caused and fixed in both parsing lanes, validated over 168 live agent turns with zero collapses. Nested edits arrays arrive intact.
  • Structured output at full speed. response_format with a JSON schema is enforced by grammar masks that compose with MTP speculative decoding: valid JSON, decode parity within noise, under five milliseconds of masking per request. Opt-in strict mode grammar-forces every tool call to a declared tool with schema-valid arguments. Contributed by PhilipJohnBasile.
  • Sessions survive history rewrites. When an agent client compacts its transcript, a common-prefix fallback keeps the session identity and its warm cache instead of forcing a cold re-prefill.
  • Leaner agent turns. The injected tool contract instructs whole-file reads and bans echoing file contents into visible text. Reasoning share of generated tokens dropped from 75 percent to under 30 in live sessions.
  • Silent hangs are contained. A heartbeat watchdog turns a frozen stream into a structured five-minute failure instead of an infinite hang. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.2.0: Context-Copy-Drafting, reparierte 4B-Modelle, FP16

MTPLX 2.2.0 aktiviert standardmäßig Context-Copy-Drafting, repariert und beschleunigt die 4B-Modelle, verhindert das Speichern unbrauchbarer Tune-Ergebnisse und ergänzt eine FP16-Option für eigene Modelle sowie weitere Korrekturen.

19 July 2026 · build 22000

The copy-drafting and small-Mac release. The model now drafts from your prompt as well as its MTP head, and the 4B pair is fixed and fast.

  • Context-copy drafting, on by default. When the model restates something already in your context (the function it is editing, a config block, a quote), whole spans are proposed as copy blocks and verified in one pass. Measured +53% on edit-heavy agent turns at temperature 0.6, parity on novel text. Exact at any temperature. Contributed by lBroth.
  • The 4B models actually work now. The old 4B drafted at zero acceptance for everyone. The engine heals existing downloads at load (no re-download), the Speed 4B is rebuilt at 227.8 tok/s, and a new Quality 4B ships at 191.7 tok/s with a 2.2x MTP multiplier. Macs under 16 GB finally get first-class recommendations.
  • Tune can no longer persist garbage. A zero-acceptance depth can never win or be saved, and poisoned records from earlier versions are quarantined automatically.
  • Forge your own FP16 models. New precision option with M1/M2 auto-select.
  • SSD session cache crash fixed. The writer thread no longer touches the GPU.
  • MoE launches at its measured depth. The 35B-A3B default corrected from the ceiling to measured D2 (by davidtai).

Download DMG · 54 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.1.0: Speicherlimit, warme Sessions, zwei neue Backends

MTPLX 2.1.0 begrenzt standardmäßig den Speicherverbrauch, hält Agent-Sessions warm, behebt Presence- und Frequency-Penalties im Quickstart-Chat sowie den Startup-Hang und bringt zwei neue native Backends und schnelleres Decoding unter Last.

17 July 2026 · build 21000

The community-fixes release. Memory is bounded by default, the agent session cache is fixed end to end, and the app startup hang is closed. Most of this release started as community reports.

  • Memory is bounded by default. The MLX allocator cache gets a RAM-tiered cap, the per-session admission gate is re-clamped on smaller machines, the paged KV pool stops growing past the context window, and a new --memory-budget flag fits MTPLX inside a declared RAM envelope.
  • Agent sessions stay warm. Prefix reuse survives every tool turn, follow-ups on hybrid models restore near the divergence point (0.4s instead of 33.8s on a 22k prompt), and warm state now survives daemon restarts. Cache hits show up in standard usage fields.
  • Penalties work in quickstart chat. Presence and frequency penalties were silent no-ops in the batched lane. Small models benefit most.
  • The startup hang is gone. Runtime installs run off the main thread and every subprocess wait has a deadline watchdog.
  • Hermes Desktop from the app. The Hermes tile launches Hermes Desktop when installed, pinned to the MTPLX backend.
  • Two new native backends. qwen3_5_mtp and hy_v3, both community contributions.
  • 8 to 10% faster decode under load. The model-owner thread is QoS pinned so background apps stop taxing generation.

Download DMG · 54 MB Full release notes …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.0.2: Weniger Wiederholungen, freier Host und Port

MTPLX 2.0.2 behebt Wiederholungsschleifen in langen Agent-Sessions, ermöglicht Serving auf beliebigem Host und Port aus der App, bietet warme Prefix-Wiederverwendung für alle Agent-Clients und respektiert die Einstellung „Off“ beim SSD-Session-Cache.

9 July 2026 · build 20200

The agent-reliability release. The repetition marathons in long agent sessions are fixed at the source, and the biggest quality-of-life issues since 2.0.1 are closed.

  • Repetition marathons fixed. Multi-turn reasoning history now renders on Qwen's trained contract (scoped to the active round). The captured looping session went from 3/4 marathons to 4/4 immediate healthy tool calls; root-caused to context construction, not quantization.
  • Serve on any host and port from the app. LAN serving (0.0.0.0) no longer misreports ports or kills healthy daemons, and the app explains the API-key requirement up front.
  • Warm prefix reuse for every agent client. Pi, Claude Code, Cline, and custom harnesses get the block-prefix warm restores that were gated to OpenCode.
  • Settings Off means Off. The SSD session cache setting is passed explicitly, so an explicit Off stays off and session-bank stops growing back. The runtime venv self-heals after updates.
  • Community fixes. Streaming reasoning-tag leak fix by @Osamaali313, quickstart host/port rendering by @hasegaw, and a bounded Ctrl-C shutdown under open streams.

Download DMG · 55 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.0.1: Turbo für alle dichten Modelle, schneller auf M1/M2

MTPLX 2.0.1 aktiviert Turbo standardmäßig für alle dichten Modelle auf jeder Apple-Silicon-Generation, beschleunigt M1/M2 und 6-Bit-Modelle, ergänzt das Modell Optimized Quality FP16 und prüft Kernel beim Laden gegen stock MLX.

7 July 2026 · build 20100

Turbo for every Mac. The v2 turbo default now covers every dense model on every Apple Silicon generation, with a load-time kernel safety net.

  • M1 and M2 get their speedup. The FP16 27B those Macs route to now defaults to turbo: decode up 19-31% across 0.5k-32k context, roughly 2x over plain autoregressive decode.
  • New 6-bit kernels. The 9B tier gains 33-62% decode and 43% faster 2k prefill under turbo.
  • New model: Optimized Quality FP16. The missing M1/M2 quality artifact, wired into the picker and the chip-aware routing.
  • Kernels prove themselves on your machine. Every turbo lane self-checks against stock MLX at load and falls back per lane if anything disagrees. Worst case is 2.0.0 speed, never wrong output.
  • Verified on real M1 hardware. Kernel exactness plus a live turbo smoke now run in CI on M1 runners.

Download DMG · 56 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.0.0: Session Cache v2 und Turbo-Decoding standardmäßig

MTPLX 2.0.0 bringt Session Cache v2 mit auf SSD erhaltenem KV-State, standardmäßig aktives Turbo-Decoding, schnelleres Long-Context-Decoding, mehr Stabilität sowie Verbesserungen bei Agent-Protokollen, Chat, Speicherbudgets, Vision und Lüftersteuerung.

6 July 2026 · build 20000

The coding-agent release. Long agent sessions in OpenCode, Pi, Hermes, and Claude Code that stay fast, stay warm, and do not fall over.

  • Session cache v2. KV state survives restarts on SSD; a 100k-token session restores in ~2s. Tool-call turns chain warm instead of re-prefilling minutes per turn.
  • Turbo decode, on by default. New verify kernels and compiled verify: 27B Optimized-Speed ~45 to 58-60 tok/s, Optimized-Quality 31-36 to 43-44 tok/s on M5 Max.
  • Long context. 64k decode +12%, 128k from 17 to 20+ tok/s, peak memory down 8-16 GB. Stock PyPI MLX, no fork, any Apple Silicon Mac.
  • Stability. The app no longer kills a healthy engine mid-session; fresh installs no longer crash at model load; SSD restores are corruption-free.
  • Agent protocol pass. OpenCode plan-to-build keeps its cache and tools; presence/frequency penalties end-to-end; honest model identity for third-party builds.
  • Chat. Markdown renders live while streaming; one compact activity strip per turn with grouped tool rounds and sources.
  • Memory that fits your Mac. Cache budgets scale to the machine, with explicit RAM and SSD limits in Settings.
  • Vision under MTP tells the truth. No more fabricated differences between similar screenshots.
  • Fans behave. Ramp on request arrival, RPM-verified, held through post-response cache work.

Download DMG · 55 MB Full release notes …

Originalquelle(öffnet in neuem Tab)Problem melden