Zum Inhalt springen

MTPLX Release Notes

32 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.12.2: Prompts auf Macs mit 8, 16 und 32 GB wieder akzeptiert

MTPLX 2.12.2 behebt, dass 2.12.1 auf Macs mit 8, 16 und 32 GB gewöhnliche Prompts mit dem vorgewählten Modell ablehnte, indem Prompts nun nach gemessenem Speicherbedarf bewertet und bei Bedarf in kleineren Chunks verarbeitet werden.

Latest 3 October 2026 · build 2012011

A fix for Macs with 8, 16 and 32 GB: 2.12.1 refused ordinary prompts there with the model the app picks for them, and 2.12.2 serves them.

  • Prompts priced at what they use. 2.12.1 charged every model the 27B's 3 GiB per prompt chunk. Each prompt is now priced at what it was measured to use: 0.19 GiB for a 59-token prompt on the 4B.
  • Smaller chunks before a refusal. A prompt that does not fit runs in 1,024 or 512-token chunks instead of being refused.
  • Served again. 2,600 to 7,800-token prompts on a 16 GB Mac with Bonsai 2 27B, 6,800 and 7,800-token prompts on a 32 GB Mac with the 27B, and a one-line prompt on an 8 GB Mac with 3 GiB free.

Download DMG · 66 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.12.1: Neuer Speicherwächter und Stabilitätsverbesserungen

MTPLX 2.12.1 ist ein Stabilitätsrelease mit neu geschriebenem Speicherwächter, erhaltenem Cache in langen Sitzungen mit Screenshots, mehr Spielraum auf 128-GB-Macs und mehreren Fehlerbehebungen.

2 October 2026 · build 2012010

A stability release for long coding sessions: the memory guard prices every request before its prompt is read and releases idle conversations before it refuses anything, long sessions keep their cache through screenshots, and Flash-Next keeps its compiled verifier past 140K tokens.

  • Memory guard, rewritten. Every request is priced at its largest moment, idle conversations are released before anything is refused, and a refusal says what holds the memory. A real Pi compaction of a 205K-token Flash-Next session was served with no refusal and swap flat (#567).
  • Long sessions keep their cache. In a Pi session on Flash-Next that grew to 204,981 tokens with 12 screenshots, 86 of 87 turns restored from RAM and all 50,769 verify steps ran compiled.
  • More room on 128 GB Macs. The default memory limit is 90 GiB instead of 96, and --memory-limit max uses everything outside macOS's reserve on a Mac that runs only the server (#548).
  • Fixes. The Metal shared event leak in long non-streaming runs (#544), JSON schema output that looped on whitespace (#547), NaN output with 4-bit and 8-bit KV cache (#526), the app badge stuck on Degraded (#528), screenshots after tool calls in streaming clients (#581), and mtplx start downloading a model you had already picked (#573). …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.12.0: Ternary Bonsai 2 27B, MiMo V2.6 und schnellere Prompt-Verarbeitung

MTPLX 2.12.0 bringt Ternary Bonsai 2 27B mit Bildeingabe für 16-GB-Macs, MiMo V2.6 Qwen 9B, eine um 85 Prozent schnellere Prompt-Verarbeitung bei Flash-Next und den 8-Bit-Build Flash-Next Optimized Quality für Macs ab 256 GB.

23 September 2026 · build 2012009

Ternary Bonsai 2 27B brings a 27B-class model with image input to 16 GB Macs, MiMo V2.6 Qwen 9B adds Xiaomi's coding model built on Qwen 3.5 9B, Flash-Next reads a 4,061-token prompt 85 percent faster than 2.11.3, and Flash-Next Optimized Quality is an 8-bit build for Macs with 256 GB or more.

  • Ternary Bonsai 2 27B. Prism ML's ternary 27B with image input is 8.85 GB with the Qwen3.8-27B draft head and runs on 16 GB Macs. Two new GPU kernels make it faster than the 4-bit 27B: 64.4 against 52.6 tok/s after a 4,061-token prompt and 57.1 against 51.0 after a 16,350-token prompt, with a peak of 11.4 GB against 23.9 (#515).
  • MiMo V2.6 Qwen 9B. Xiaomi's coding and agent fine-tune of Qwen 3.5 9B, as an 8.70 GB 6-bit pack. On 16 to 31 GB Macs it is the second suggestion, after Bonsai 2.
  • Flash-Next reads prompts faster. On the same Mac and prompts as 2.11.3, the first token of a 4,061-token prompt arrives after 2.88 s instead of 5.26 s, and of a 65,502-token prompt after 60.1 s instead of 85.5 s, with decode 13 percent faster at that length. New prompt kernels add another 9 to 16 percent.
  • Flash-Next Optimized Quality. An 8-bit build for Macs with 256 GB or more, a 169.96 GB download. It has not run on a 256 GB Mac yet, so it is listed second there. Forge now converts Flash-Next source models (#508 by @bpmforge). …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.11.3: Decoding-Fehler behoben, Flash-Next bei langem Kontext schneller

MTPLX 2.11.3 behebt acht Exaktheitsfehler beim spekulativen Decoding, beschleunigt Flash-Next bei langem Kontext um bis zu 27 Prozent, unterstützt 261k-Token-Prompts und hält Konversationen mit Tools warm.

17 September 2026 · build 2011050

Eight exactness defects fixed with tests, speculative decoding measured exact by token id on both flagship models, warm agent turns that restore the exact state, Flash-Next decode up to 27 percent faster at long context, 261k-token prompts on Flash-Next, a session cache that cleans itself, and a set of app and agent-harness repairs.

  • Exact speculative decoding, measured. Eight defects in the decode path are fixed, each with a test. At temperature 0 and at the native sampler the fast path matches the plain path by token id on Flash-Next and on the 27B, and warm turns restore the state they were saved with.
  • Faster where it counts. Flash-Next decodes a 9k-token code prompt at 79.3 tok/s against 62.5 on 2.11.2 on the same Mac, a 109k-token OpenCode turn went from 48.8 to 61.8 tok/s, and 261k-token prompts decode instead of running out of GPU memory (PR #482, davidtai).
  • Conversations with tools stay warm. A 19k-token app chat with web search went from 15 s to the first token on every turn to 0.6 s, Hermes turns keep their cache across the model's own token boundaries, and a mid-conversation system message no longer re-prefills the history (#477). …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.11.2: Korrekte 27B-Ausgaben und stabilere lange Agent-Sitzungen

MTPLX 2.11.2 korrigiert falsche Ausgaben des 27B-Modells auf Macs außer M5, behebt Session-Bank- und Speicherprobleme bei langen Agent-Sitzungen und ermöglicht Flash-Next Bildeingabe in Hermes, OpenCode und Pi.

6 September 2026 · build 2011018

A correctness release for every Mac that is not an M5, seven session-bank and memory fixes for long agent sessions, native vision for Flash-Next in Hermes, OpenCode and Pi, a verifier-depth fix that restores Flash-Next's compiled verify route in agent turns, and a set of app and CLI repairs.

  • 27B correct again on M1 to M4. 2.11 turned on a flash-decoding verify route that needs the M5 GPU's tensor units and shipped it without a hardware gate; from 8,192 tokens on, other Macs produced unrelated text, loops or another language. The route now runs only on an M5-class GPU on macOS 26.2 or newer; every other Mac serves the packed kernel from 2.10.2 (#459, #464, #467, #461, #469).
  • Long agent sessions keep their cache. A deep conversation keeps its session at its turn boundaries (#446), later conversations are admitted after a restart (#454), AR-only runtimes restore again (#465), and a request that would cross the memory limit is refused with a structured 507 before prefill instead of swapping the Mac to a panic (#450, #447).
  • Flash-Next sees your agent's screenshots. Hermes, OpenCode and Pi advertise image input when the pack has a vision tower, images keep sparse attention, and a warm image turn is admitted as warm: a fifteen-screenshot Hermes session completed at a 92 GiB peak. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.11: Schnelleres Decoding, lokaler Server, helles Design, zwölf Sprachen

MTPLX 2.11 beschleunigt das Decoding bei beiden Flaggschiffmodellen, vermeidet Leerlauf in Coding-Agent-Sitzungen und bringt einen nur für eigene Seiten erreichbaren lokalen Server, ein helles Erscheinungsbild, zwölf App-Sprachen und viele Fixes.

4 September 2026 · build 2011007

Faster decode at every context on both flagship models, coding-agent sessions without dead time between tool turns, a local server that only its own pages can reach, a light appearance, the app in twelve languages, and the largest fix batch of any release so far.

  • Flash-Next faster at every context. Same-hour pairs against 2.10.2: 53 to 68 tok/s at 16k, 48 to 61 at 100k, 32 to 44 at 206k. A compiled verify lane adapted from PR #391 by @davidtai, with the rows-gather fence, the pack rules and the per-request memory gate built around it.
  • 27B flash-decoding verify. A TensorOps flash-decoding kernel cuts verify attention per layer by 35 percent at 72.7k context, at half the power: 38.6 to 41.4 tok/s at 16k, 20.5 to 23.8 at 88k from the route alone, and 30.9 at 88k with the two opt-in long-context settings.
  • Agent sessions without dead time. Tool-call turns bank in place, warm turns stop waiting on gathers they never use, and a forced tool round in a 41k session drops from two cold re-prefills to 0.58 s. A new agent-session gate runs inside the release script.
  • Light appearance and twelve languages. A curated cream palette under Settings, and onboarding opens with a language step; existing installs are asked once. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.10.2: HTTP 507 bei Speicherablehnung, korrigierte Anthropic-Bridge

MTPLX 2.10.2 liefert ehrliche Speicherablehnungen mit HTTP 507, eine korrigierte Anthropic-Bridge für Claude Code, respektiert explizit deaktivierte Kompilierung und protokolliert Stopp-Gründe.

1 September 2026 · build 2010003

Honest memory refusals, a correct and resilient Anthropic bridge for Claude Code, and sharper stop diagnostics.

  • Honest memory refusals. A prompt that cannot fit is answered upfront with a structured HTTP 507 naming the shortfall, instead of dying mid-prefill and being logged as a client cancellation.
  • Claude Code fixed twice. Context accounting no longer double-counts cached prefixes on session-cache hits, and long first turns (large MCP toolsets reach 165k tokens) survive the client's 300-second stream timeout during prefill. A measured 165k-token first turn completes, then follow-ups serve 99.8 percent of the prompt from cache at 3.9 s to first token.
  • Explicit off means off. An explicit compile kill-switch now wins over profile auto-arming everywhere, including the path that silently re-armed it.
  • Diagnosable stops. Request logs record why a response ended: model EOS, stop sequence, length cap, or repetition stop.

Download DMG · 59 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.10.1: Schnellere lange Prompts auf M4/M5, Bildeingabe für Flash-Next

MTPLX 2.10.1 beschleunigt die Verarbeitung langer Prompts auf M4- und M5-Macs, ermöglicht Bildeingabe für Flash-Next und behebt Probleme auf 96-GB- und M2/M3-Macs sowie leere Antworten.

30 August 2026 · build 2010002

Faster long-prompt processing on M4 and M5 Macs, working image input for Flash-Next, and fixes for 96 GB and M2/M3 machines.

  • Sparse prefill for Flash-Next. A 98k-token prompt processes 35 percent faster at 8 GB lower peak memory, and a full 262k cold prompt completes in 355 s at 87 GB where 2.10.0 climbed to 119 GB and produced nothing.
  • Flash-Next reads images. The packs always shipped their vision weights; the runtime now serves them in app chat, over the API, and in Pi.
  • 96 GB and M2/M3 Macs. The preload memory check scales with the machine so the packs load on 96 GB, and the kernels that crashed M2 and M3 GPUs now fall back automatically.
  • No more empty answers. A first turn with tools declared that ended inside the reasoning channel now continues to a visible answer instead of returning an empty message.

Download DMG · 59 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.10.0: Speicherplanung pro Mac und native Qwen3.8 Flash-Next-Unterstützung

MTPLX 2.10.0 plant den Speicher passend zum Mac, hält die Decode-Geschwindigkeit bei langem Kontext, unterstützt Qwen3.8 Flash-Next nativ und verkürzt die Laufzeit von Coding-Agents deutlich.

29 August 2026 · build 2010000

The engine plans memory for your Mac, decode holds its speed deep into long context, Qwen3.8 Flash-Next runs native from day 0, and coding agents got their wall clock cut by two thirds.

  • Faster everywhere. Against stock 2.9.2 on the same Mac: +15 percent decode at 3k context, +29 at 88k, +54 at 147k, and +41 percent prefill.
  • Memory planned for your Mac. A 48 GB Mac resolves a context window that actually fits and serves 33 tok/s where 2.9.2 sat in swap at 3 to 4. The pressure banner now names whose pressure it is.
  • Qwen3.8 Flash-Next, native. The 125B MoE preview measures 61 tok/s plain and 63 to 76 with its own MTP head on an M5 Max. The 32 GB n-gram table streams from SSD, so both packs fit 96 GB Macs.
  • Agents finish faster and stop dying. The same multi-file task dropped from 150 to 44 seconds, mid-session first token is 0.11 s, and the run killers are fixed: dropped tools, phantom cancels, silent tool errors.
  • Long work holds up. KV q8 costs about 4 percent for double the context headroom, a 34k-token answer no longer decays from 86 to 25 tok/s, rewrites run 19 percent faster than fresh writes, and 100k-token sessions survive a restart.

Download DMG · 59 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.9.2: Agent-Transkripte unverändert, schnelleres greedy Decoding

MTPLX 2.9.2 verändert Agent-Transkripte standardmäßig nicht mehr, beschleunigt greedy Decoding unter 12k Kontext und behebt Fehler in Forge, Vision und Client-Konfigurationen.

25 August 2026 · build 2009010

Your transcript is yours: agent serving is passthrough by default, greedy decoding gets faster below 12k context, and the model forge gets a correctness fix that rescues collapsed draft acceptance.

  • Your transcript is yours. MTPLX no longer compacts tool results, trims file reads, or injects steering text into agent transcripts unless you explicitly turn a rewrite on. The app and the CLI both stopped exporting the legacy compaction settings.
  • Faster greedy decode. Chained greedy drafting is on by default for temperature 0 under 12k context: measured +2.5 to +9.8 percent across 0.5k to 8k prompts, fenced off where it lost.
  • Forge correctness. The MTP norm convention is decided once per tensor set, rescuing packs whose draft acceptance had collapsed to 0 to 2 percent, and the runtime refuses a double-shifted trunk with a clear error.
  • Vision fixes. Images survive message canonicalization on retried turns and near-prefix cache restores. Both were silent vision-drop bugs.
  • Your config is yours too. Managed client configs only update files MTPLX wrote itself and never overwrite a config you have customized.

Download DMG · 60 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.9.1: Abstürze bei langem Kontext behoben, Flight Recorder mit mtplx trace

MTPLX 2.9.1 behebt Abstürze bei langem Kontext, entfernt versteckte Ausgabelimits, erhält Reasoning über Agent-Turns und ergänzt einen Flight Recorder mit mtplx trace.

22 August 2026 · build 2009009

Agent coding sessions run to completion: long-context crash fixes, no hidden output caps, reasoning preserved across turns, and a built-in flight recorder.

  • Long sessions no longer die. Fixed a paged-cache bug that could truncate and crash agent sessions near 19,000 tokens, and a shutdown segfault on quit.
  • No hidden output caps. OpenCode, Pi, and Hermes injected default ceilings are stripped; explicit caps you set are honored on every lane.
  • The model thinks once. Prior reasoning survives across agent turns instead of being re-derived, and warm turns reuse the cache instead of re-reading the whole session.
  • Turbo profile truth. The turbo fast path now applies exactly what it advertises and reports it in /health, if you benchmarked turbo on 2.9.0, re-run it.
  • Flight recorder. Every request records a per-second telemetry log, and mtplx trace turns any coding session into a full diagnosis report. Local-only, a few MB a day.

Download DMG · 60 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.9.0: 15 bis 20 Prozent schnellere Generierung und Modell-Updates per Klick

MTPLX 2.9.0 macht die Generierung um 15 bis 20 Prozent schneller, glättet das Streaming, verkleinert Modell-Updates auf Deltas von 240 bis 450 MB und bietet Modell-Updates per Klick.

20 August 2026 · build 2009001

Faster responses, smoother streaming, smaller model downloads, and one-click model updates.

  • Faster generation. Decode is 15 to 20% faster on typical workloads and up to 60% faster on code-heavy output.
  • Smoother streaming. Visible freezes fell by 95%, and the worst measured stall dropped from 725 ms to 109 ms.
  • Much smaller updates. Existing models now update with a 240 to 450 MB delta instead of downloading 15 to 21 GB again.
  • Smaller models, updated in one click. Qwen 3.8 packs are up to 610 MB smaller, and the app now shows and installs model updates directly.

Download DMG · 60 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.8.3: Streaming-Ruckler und Scrollen während der Generierung verbessert

MTPLX 2.8.3 behebt Ruckler beim Streaming während der App-Nutzung, verbessert das Scrollen während der Generierung und korrigiert leere Transkripte, Denk-Anzeige und Tabellen-Streaming.

18 August 2026 · build 2008003

The streaming-quality release. 2.8.2's field reports kept coming back to one thing: chats that freeze for half a second and land in bursts, but only when a human was actually using the app. Hands-off testing stayed clean, which is exactly how it survived. 2.8.3 closes that whole class, plus every streaming bug found on the way to it.

  • Streaming stays smooth while you touch the app. The window-measurement guard ran in a run-loop phase macOS skips while input events keep arriving, precisely when you scroll or move the mouse. Now it runs every turn: 40 s of continuous wheel-scrolling went from 70 UI stalls (18.7 s frozen) to one, and 30 s of cursor movement from 91 stalls to zero.
  • Scrolling up mid-generation no longer fights you. User scrolling wins instantly, trackpad, momentum, and classic wheel mice alike, and following re-engages when you return to the bottom.
  • The transcript can't go blank, thinking is plain text, tables stream correctly. The mid-generation blank-out, the self-rewriting reasoning ticker, and whitespace-free freezes (tables, URLs, minified code) are all fixed, with a quadratic detokenizer path closed on exactly that content.
  • Every request records a stream-smoothness census. Producer gap percentiles land in every request record, including cancelled runs, which used to log nothing and were exactly the runs people complained about. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.8.1: Cache-Probleme in Agent-Sitzungen behoben, verlässliche Statistiken

MTPLX 2.8.1 behebt stockende Caches in langen Agent-Sitzungen, macht gemeldete Statistiken und /health verlässlich und verhindert, dass ein Bild den Cache eines anderen Bildes liest.

17 August 2026 · build 2008001

The 2.8 line. This release is about trust: people started benchmarking MTPLX seriously and running long agentic sessions against it, and both groups found real problems. 2.8 closes that work, and 2.8.1 is the build that ships it to the desktop, together with a vision cache fix our own release gate caught the same morning.

  • Long agent sessions no longer stall their cache. Sessions used to quietly stop reusing their prompt cache around 38k tokens and re-prefill a growing suffix every turn. The committed frontier now advances on every turn, tool-call turns included, so a 45k-token conversation keeps turn-delta prefills only.
  • Every number the server reports is one you can bench against. mtplx_stats is always populated, temperature 0 is exact, the logprobs contract is parser-safe, and /health now reports any degradation (compiled-verify fallbacks, overridden profile keys, kernel bails) instead of looking like turbo while running slow.
  • A different image can never read another image's cache. The 2.8.1 fix: cache keys for vision turns are derived from the actual image bytes, and the raw session frontier no longer commits image histories. Identical pixels still restore the full prefix; different pixels stop cold before the image. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.7.2: mtplx pull beschädigt Modelldateien nicht mehr

MTPLX 2.7.2 behebt einen Fehler, bei dem mtplx pull geänderte Modelldateien beschädigen konnte, und stellt die Bildfähigkeit der Qwen-3.8-Modelle wieder her.

16 August 2026 · build 2007002

An emergency fix for mtplx pull. On 2.7.1 and older, re-pulling a model that changed upstream can corrupt your local copy. Upgrade before you pull.

  • Pull no longer corrupts files that changed upstream. The downloader treated a complete local file whose size no longer matched the server as an interrupted download, and appended the remote tail onto the old content, corrupting config.json and the safetensors index. Stale files are now re-fetched whole; genuinely interrupted downloads still resume.
  • Already hit by it? Delete the model's config.json and model.safetensors.index.json, then pull again on 2.7.2. The full notes have the details.
  • The Qwen 3.8 models can see again. All six published 3.8 repos were re-published with their vision towers restored, and forge now grafts the tower on every build, failing closed rather than publishing blind. Existing installs pick the repair up as a ~0.9 GB delta.

Download DMG · 56 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.7.1: Korrekturen bei xhigh, KV-Cache und Diagnose

MTPLX 2.7.1 behebt die Auswahl von xhigh, die KV-Cache-Quantisierung für Qwen 3.8, ungenaue Diagnoseangaben und fälschlich angebotene ältere Builds.

15 August 2026 · build 2007001

A bug-fix release. It clears the known-issues list 2.7.0 shipped with.

  • xhigh stays selected. Picking it in Inference settings snapped back to medium, and mtplx config set reasoning_effort xhigh was refused. Both places carried their own copy of the effort list and neither knew about xhigh, so the save was rejected whole and the picker reverted. Every place that accepts an effort level now reads the same list.
  • KV cache quantization reaches Qwen 3.8. The toggle displayed q8 while the launch path recognized only Qwen 3.5 and 3.6.
  • Honest diagnostics. mtplx doctor names the model it actually checked, and turbo's profile note reports the real 32,768-token compiled-verify fence.
  • A new build can't offer you an older one. The build number derived for 2.7.1 came out below the shipped 2.7.0, so a fresh install proposed 2.7.0 to itself.

Download DMG · 56 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.7.0: Qwen3.8-27B als neuer Standard ab 32 GB

MTPLX 2.7.0 unterstützt Qwen3.8-27B mit drei MTPLX-Builds und FP16-Versionen für M1 und M2, setzt es als neuen Standard auf Macs ab 32 GB und erweitert das Compiled Verify auf 32k Kontext.

15 August 2026 · build 27000

Qwen3.8 support 🎉. Qwen3.8-27B came out on 14 August; this release runs it the way the model card says, with three MTPLX builds tuned for it and FP16 versions of all three for M1 and M2 Macs.

  • Qwen 3.8, served properly. Official sampler (1.0 / 0.95 / 20), reasoning effort xhigh, medium and low with medium as the coding default, thinking preserved in history. Bare Speed (16.0 GB), Optimized Speed (20.4 GB, recommended) and Optimized Quality (29.4 GB), each with its calibration stamped in its own metadata.
  • New default. Macs with 32 GB or more now default to Qwen 3.8 Optimized Speed; M1 and M2 get its FP16 build automatically. Under 32 GB still routes to the 9B.
  • Compiled verify to 32k. The compiled verify window moves from 12k to 32k tokens of context: +6.9% at 20k on Qwen 3.8 Bare Speed, peak memory flat at 20k and lower at 30k.
  • Coding agents uncapped. OpenCode and Pi no longer send an output cap for MTPLX models; Pi sessions restore their banked prefix from RAM.
  • Fixes. The SSD session cache no longer walks its whole store on every write or health poll (idle CPU 35% down to 0.2% on an 816k-file bank); the macOS 27 slider crash is fixed (thanks @joshlacal); mtplx pull names the mirror knob when huggingface.co is unreachable.

Download DMG · 56 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.6.0: mtp_batch, Embeddings, Rerank und LFM2

MTPLX 2.6.0 ermöglicht gleichzeitiges spekulatives Decoding über --scheduler-mode mtp_batch, ergänzt Embeddings- und Rerank-Endpunkte, unterstützt LiquidAI LFM2 und korrigiert Temperature-0-Exaktheit.

11 August 2026 · build 26000

Concurrency. Speculative decoding used to be a single-user feature: the moment two requests hit the daemon at once, everyone fell back to plain batching. 2.6.0 removes that trade-off.

  • Concurrent speculative decoding. --scheduler-mode mtp_batch serves independent requests through fixed-width MTP cohorts (three-wide and eight-wide) with per-row stats and honesty controls, and the session bank composes with it.
  • Embeddings and reranking. New /v1/embeddings and /v1/rerank endpoints (Cyb3rb1ade).
  • LiquidAI LFM2. LFM2 and LFM2.5 models run on MTPLX (davidtai).
  • Temperature-0 exactness. A real correctness fix to greedy decoding under prefill partitioning, plus a serial-lane sampling speedup.

Download DMG · 56 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.5.4: Schnellere warme Turns, weniger Stalls

MTPLX 2.5.4 beschleunigt warme Turns in Agent-Sitzungen, vermeidet Hintergrund-Stalls durch SSD-Cache-Arbeit und zeigt beim Start das Cache-Budget an.

7 August 2026 · build 25400

Agent sessions got the attention this cycle, especially Pi. Warm turns stay warm.

  • Faster warm turns. Tool turns restore from the stable boundary instead of re-processing ~200 tokens per round; a postcommit about to finish is briefly waited for instead of thrown away (a 1,449-token re-prefill became 436 tokens, first token 2.7 s down to 1.1 s).
  • No background stalls in your turn. SSD cache work no longer slips into the gap between a request's internal jobs; the SSD tier skips candidates that cannot win.
  • The cache tells you what it is doing. Resolved cache budget printed at startup, per-session and total (#229, #230).

Download DMG · 56 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

MTPLX

MTPLX 2.5.3: Agent-Anfragen warten nicht mehr auf Cache-Arbeit

MTPLX 2.5.3 verhindert, dass Agent-Anfragen hinter Hintergrund-Cache-Arbeit warten, beschleunigt warme Folgeanfragen um 77 bis 92 ms und behebt irreführende API-Antworten.

6 August 2026 · build 25300

A small release focused on the agent lane and the API surface, from a day of head-to-head benchmarking.

  • Agent requests no longer stall behind background cache work. Background commits yield the moment any request is admitted, whichever session it belongs to (worst measured case before: a follow-up turn 44% slower).
  • Warm follow-ups got faster. A byte-identical transcript hits an exact-match encode cache and gets 77 to 92 ms back per request.
  • API honesty. Fixes for places where the API misled external tools.

Download DMG · 56 MB Full release notes

macOS 14+ · Apple Silicon (M-series) · Notarized

Originalquelle(öffnet in neuem Tab)Problem melden