Zum Inhalt springen

Unsloth Release Notes

14 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Unsloth

Unsloth: Command Palette und Desktop-UI/UX-Verbesserungen

Unsloth Desktop erhält eine Command Palette (Cmd/Ctrl+P), teilbare GGUF-Run-Einstellungen, klarere Fehlermeldungen, bis zu 4,1x schnellere Laya-Entscheidungen sowie speicherschonenderes 4-Bit-LoRA-Training für NVFP4-, INT4- und MXFP4-Checkpoints.

This release makes Unsloth Desktop faster and easier to use with improved navigation, shareable run settings, clearer errors, faster Laya decisions, expanded Decision API support, and better 4-bit LoRA training. Cmd/Ctrl+P command palette and shareable GGUF run settings.Laya decisions up to 4.1x faster with hosted provider support.Clearer errors with improved load/generation messages and logs.Lower-memory 4-bit LoRA training for NVFP4, INT4, and MXFP4 checkpoints. Desktop, chat, and models Share GGUF run settings without loading models.Continue responses and fix HTML canvas errors directly in chat.Improved model fit warnings, image loading, and desktop controls.Added FastModel fine-tuning support for T5, T5Gemma, BART, Marian, and more.Added unsloth eval for evaluating checkpoints and LoRA adapters.Improved multi-GPU training and generation support. 4-bit checkpoint training Pre-quantized checkpoints now train using their original packed weights, reducing memory use while preserving accuracy. Qwen3.8-27B-NVFP4 uses 40.2 GB peak memory instead of 72.9 GB.INT4 checkpoints stay packed instead of being re-quantized.Kimi-K2.7-Code and gpt-oss retain packed experts during LoRA training. Decision API, image, and platform updates Laya decision models respond up to 4.1x faster, with hosted providers available through Connections.Image models preserve INT8/FP8 during offloading for faster generation.Qwen-Image-2.1 renders 1536×1536 images up to 3.6x faster on Radeon 8060S.Added AMD Win…

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Unsloth

Unsloth: Laya Decision Models und neue Library

Unsloth unterstützt jetzt Decision Models wie Laya, bietet eine neue Library für Chats, Bilder, Videos und Dokumente, einen Skills Editor, einen Document Viewer, ModelScope-Downloads, Verbesserungen für Apple Silicon sowie rund 4,5x schnellere Bild- und Videogenerierung.

We're adding support for decision models, a unified Library for docs and media, document viewer, many Apple Silicon improvements, creation of Skills, and ~4.5× faster image and video generation. Run and serve Decision Models like Laya (open-source Jev) locally through a TypeSafe-compatible API.Skills Editor to create, edit, and delete Skills directly in Desktop.ModelScope model downloading is now available for users who can't use Hugging FaceDocument Viewer for PDF, Word, Excel, PowerPoint, and chats.New Library: Manage chats, images, videos, and files from the new Library tab.Apple Silicon Improvements including batched serving, structured outputs, and TurboQuant KV cache.Configure how much of a model each GPU receives.Faster Image + Video Generation with ~4.5x faster LTX-2.3 clips and 1.7–6.3x faster VAE decoding.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Unsloth

Unsloth: Qwen-Image-2.1 und benutzerdefinierte Skills

Unsloth kann nun Qwen-Image-2.1 lokal ausführen und bringt benutzerdefinierte Agent Skills, flüssigere Reasoning-Blöcke mit 60 FPS, einfachere Chat- und Projektverwaltung, zuverlässigeres Training sowie bessere Linux- und AMD-Unterstützung.

You can now run Qwen-Image-2.1 locally with Unsloth! This release also includes custom Agent Skills, and easier chat/project management. It also brings 2x faster reasoning blocks (60 FPS vs 30 FPS), more reliable training, and improved Linux installs and updates. Qwen-Image-2.1 support: Run Qwen-Image-2.1 locally with Unsloth for image generation.Custom Agent Skills: Add task-specific skills, reuse skills from Claude Code and .agents folders, and select them in chat with @.Smoother chat experience: Long reasoning blocks now render at 60 FPS, up from 30 FPS. Includes a redesigned thinking UI, smoother code streaming, and more reliable model loading and chat settings.Easier chat and project management: Drag chats to reorder, pin, or move them into projects across all platforms. Edit project names, instructions, and folders directly from the Projects page.Better model and storage handling: Vision models can see images returned by MCP tools. Reuse downloaded models without fetching duplicate copies, and inspect or clear caches from Settings.More reliable training and exports: Resume past training runs with accurate remaining-time estimates. Fixes improve Hugging Face dataset recipes, column mapping, and split/subset selection. Hub uploads exclude files left over from earlier exports.Improved Linux and AMD support: In-app updates for Debian installs, a native ARM64 installer for Ubuntu 24.04+, and a new AMD ROCm Docker image with Unsloth Studio and JupyterLab.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Unsloth

Unsloth: Docker, Multi-User und AMD-Unterstützung

Unsloth bringt ein aktualisiertes Docker-Image mit NVIDIA- und AMD-Unterstützung, Multi-User-Konten mit Isolation, INT8/FP8-Bilddiffusion, ARM64-Windows-CUDA-Support, RDNA1/RDNA2-Support sowie Verbesserungen bei GRPO, Chat und Inferenz.

We're releasing our new updated Docker image along with multi-user accounts, RDNA1+2, FP8/INT8 diffusion support, ARM64 CUDA Windows support and many training, GRPO and inference improvements Highlights: Updated Docker with NVIDIA and AMD support: GuideMulti-user accounts with isolation: Settings > AccountsINT8/FP8 image diffusion inference: 2x fasterARM64 Windows CUDA support for training and inferenceGRPO: Qwen3.5 and latest TRL/vLLM supportAMD RDNA1 and RDNA2 supportBetter Windows NVIDIA GPU detection and recovery Updated Docker image Updated CUDA image and Studio setup, removed bundled caches, and restored Unsloth’s training patches on GPU hosts.Studio data persists on a volume without pinning app code. Improved updates, generated passwords, configurable ports, and package preservation.Added an AMD ROCm image for supported Linux hosts. Multi-user accounts Create accounts in Settings > Accounts with one-time setup codes and individual passwords.Accounts keep work separate while sharing a loaded model when settings match. Single-account behavior is unchanged. Chat + reasoning Edit, reorder, and steer queued prompts, even while a local model loads.Added GGUF reasoning budgets, adjustable chat width, desktop scaling, and a compact composer.Improved long-reasoning responsiveness, chat-history preservation, and tool-generated file and image handling. Hardware + inference MLX gains video input, optional MoE and decode optimizations, and more reliable multimodal chats.Improved GG…

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Unsloth

Unsloth: Qwen3.8-Flash und GLM-5.3 bis zu 2x schneller

Qwen3.8-Flash-Next und GLM-5.3-Flash laufen dank Faster Decoding und MTP, das jetzt standardmäßig aktiv ist, bis zu 2x schneller, außerdem enthält das Release über 170 Verbesserungen bei Training, Chat, Hardware-Support, APIs und Performance.

Run Qwen3.8-Flash-Next and GLM-5.3-Flash up to 2× faster with faster decoding and bonus MTP, now enabled by default. This release also includes 170+ improvements across training, chat, hardware support, APIs, and performance. Faster Qwen and GLM generation with MTP, plus improved tool use and recommended Qwen settings.Much less laggy UI/UX for all chats especially on long conversationsNew support for MiniMax-Music3, Higgs, and MOSS audio models, including progress tracking and clip management.More reliable model loading across local servers and connected providers.Safer chat editing that preserves tool calls, replies, media, and conversation branches.Improved multi-GPU training, memory fitting, model placement, and GGUF exports.Stronger AMD/ROCm installation, detection, and GPU compatibility.More reliable MCP, Deep Research, OAuth, parallel tool calls, and agent workflows.New OpenAI-compatible APIs for video, audio, and MLX-served models.Improved desktop updates, LAN port controls, downloads, and GPU memory cleanup.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Unsloth

Unsloth: Qwen3.8-Flash-Next und GLM-5.3 lokal ausführbar

Qwen3.8-Flash-Next, GLM-5.3-Flash und GLM-5.3 lassen sich nun lokal in Unsloth ausführen, dazu kommen bis zu 5x schnelleres RAM-Offloading, zuverlässigere Auto Compaction, Speicherschätzungen vor dem Laden, JSONL-Chat-Export und über 100 weitere Verbesserungen.

Qwen3.8-Flash-Next, GLM-5.3-Flash + GLM-5.3 can now run locally in Unsloth! Qwen3.8-Flash runs on 75GB RAMGLM-5.3-Flash runs on 102GB RAM + VRAMUp to 5x faster RAM offloadingRepeated Auto Compaction now works reliably100+ chat, reliability and performance improvements Qwen3.8-Flash-Next 125B multimodal reasoning model and early preview of Qwen4's architecture. 1-bit Dynamic GGUF runs on 75GB RAM or unified memoryUp to 262K context with text and imagesNone, Low, Medium and Extra High reasoning modesPreserved Thinking improves consistency in long chats GLM-5.3-Flash 320B multimodal model with 18B active parameters. 1-bit model runs on 102GB combined RAM + VRAMUp to 1M context for text, images and documentsLow, High and Max reasoning modesImproved coding, agent and vision performance Unsloth improvements Large GGUFs automatically split across GPU and system RAMSee memory estimates before loadingChats recover after disconnectsBetter multi-image vision and MCP image supportExport chats as JSONLImproved Auto Compaction controlsFixes for Linux, NVIDIA Wayland, AMD and Windows To run or train the new models, you can download Unsloth Desktop: Download for macOS Download for Windows Download for Linux

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Unsloth

Unsloth: Auto Compaction und LAN-Zugriff

Mit v0.1.801-beta führt Unsloth experimentelle Auto Compaction für lange Chats, Remote- und LAN-Zugriff als Preview, schnelleren Chat, Unterstützung für eigene llama.cpp-Builds, zusätzliche Inferenz-Schalter und Unsloth Dynamic v3.0 ein.

Thanks for the support for Qwen3.8-27B and Unsloth Desktop last week! For our new v0.1.801-beta release, we merged 200+ PRs to introduce many new features, fixes including: Auto Compaction (Experimental) for longer chats beyond context limitsRemote & LAN Access (Preview) for easy network access without Cloudflare linksFaster Chat - Improved streaming performance, reduced UI lag, and smoother long conversations.Support for custom llama.cpp builds. Toggles for Cache RAM, Mmap, Mlock, Checkpoints, Spe Decoding KV Cache, Vision On / OffUnsloth Dynamic v3.0 is released. New Qwen3.8-27B Dynamic v3.0 GGUFs deliver >10% higher top-1 accuracy compared to everyone else. Works with Unsloth. Auto compaction (Experimental) Long chats can now exceed context limits by moving older turns into a searchable archive. Older turns are removed only when needed and remain searchable.Fresh context epochs improve recall without permanently trimming chats.Uses retrieval instead of summarization for better accuracy. Remote & LAN access (Preview) Access Unsloth from other devices on your network. New remote access settings.LAN control, QR codes, and auto-start options.Disabled by default for security. Chat improvements Faster long chats and improved threading.Projects organize chats, files, and workspaces.Added prompt queueing, shortcuts, edit_file, and better tool support. Hardware, inference, API Support for custom llama.cpp builds and Intel XPU.More inference controls (cache, mmap, mlock, checkpoint…

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Unsloth

Unsloth: Qwen3.8 lokal und weitere Neuerungen

Qwen3.8-27B und Qwen3.8-2.4T laufen jetzt lokal in Unsloth (Qwen3.8-27B auch mit Fine-Tuning), dazu kommen zusätzliche llama-server-Argumente, Tool Calling für externe Provider, Codex-Login, schnellere FP8- und GGUF-Inferenz sowie behobene Bypass-Berechtigungen.

Qwen3.8-27B and Qwen3.8-2.4T can now be run locally in Unsloth! Run on 17GB RAM via Unsloth Dynamic GGUFs. You can also fine-tune Qwen3.8-27B in Unsloth. Qwen3.8-27B is by far the strongest model for its size. We also uploaded NVFP4 quants. Other Unsloth updates include: Qwen3.8-27B + extra llama-server arguments allowed + custom VRAM toggleExternal provider has tool calling + tool support + login with CodexFast FP8 10x faster MiniMax-H3 inference (3 minutes vs 30)10% faster inference for GGUFs + Bypass permissions fixedConnected API providers can use their own Search or Unsloth Desktop's built-in Search and tools. Tool results are passed back to the model so it can continue multi-step tasks.Sign in with a Codex subscription and use Codex tools inside Chat. To run or train Qwen3.8, you can download Unsloth Desktop: Download for macOS Download for Windows Download for Linux

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Unsloth

Unsloth Desktop: Open-Source-App für Mac, Windows und Linux

Unsloth Desktop ist erschienen, eine Open-Source-App für Mac, Windows und Linux zum lokalen Ausführen und Trainieren von Modellen, mit Unterstützung für MLX, GGUF, Bild-, Video- und Audiomodelle, Websuche, Deep Research, RAG, MCP und OpenAI-kompatibler API.

Introducing Unsloth Desktop 🦥 - the first desktop app to run and train models locally. Open-source. Runs on Mac, Windows and Linux • Supports MLX, diffusion image/video, audio, GGUF • Connect Claude Code and Codex to local LLMs • 50% more accurate, self-healing tool calls + sandboxed code exec • Works for CPU + multiGPU setups - NVIDIA, AMD, Intel, Mac • Train models 2× faster with 70% less VRAM • Private web search, deep research, RAG, MCP + exports (NVFP4, GGUF) • Use Unsloth’s OpenAI-compatible API and cloud models • Securely deploy LLMs remotely and access anywhere You can download Unsloth Desktop now: Download for macOS Download for Windows Download for Linux

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Unsloth

Meta Muse Glimmer lokal in Unsloth ausführen und feintunen

Meta Muse Glimmer, ein dichtes 30B-Modell unter Apache-2.0-Lizenz, lässt sich in Unsloth mit etwa 20GB RAM/VRAM lokal ausführen und mit 20GB VRAM feintunen.

Meta has released Muse Glimmer, the first open model from Meta Superintelligence Labs. Muse Glimmer is a dense, 30B parameter model by Meta, built for local agentic and coding workflows. Released under the Apache 2.0 license you can run Muse Glimmer 30B via Unsloth and Unsloth Dynamic quants for the best performance. Muse Glimmer 30B can run locally on 20GB RAM/VRAM setups, including Mac and GPUs. You can also train and fine-tune Muse Glimmer 30B using Unsloth on 20GB VRAM. You can run and fine-tune Muse Glimmer via Unsloth Desktop app: Download for macOS Download for Windows Download for Linux

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Unsloth

Kimi K3, Deep Research und paralleler Chat in Unsloth

Unsloth kann Kimi K3 mit Dynamic GGUFs lokal ausführen, mehrere Chats parallel generieren lassen und bietet einen neuen Deep-Research-Modus sowie besseren AMD- und Intel-GPU-Support, DoRA-Training und diverse Fixes.

Hey everyone! Kimi K3 can now run locally with Unsloth Dynamic GGUFs, Unsloth can keep multiple chats generating in parallel, and the new Deep Research mode plans, reads and cites sources using your local model. This release also brings better AMD and Intel GPU support, DoRA training, and many installer, MLX, export and inference fixes. Kimi K3 Moonshot AI’s Kimi K3 is a 2.8T-parameter MoE model with 104B active parameters, native vision support and a 1M context window. Kimi K3 is thinking-only and Unsloth supports low, high and max reasoning efforts. You can run our Kimi K3 Dynamic GGUFs through Unsloth. Unsloth automatically detects multi-GPU setups and can offload model layers to system memory. Kimi K3 is a very large model, so plan your hardware accordingly: UD-IQ1_S is 595GB in disk space.UD-Q4_K_XL is 1.51TB in disk space.For lossless inference, use UD-Q8_K_XL, which is 1.56TB in disk space. Read our full Kimi K3 guide and learn more about Unsloth Dynamic 2.0 GGUFs. Parallel Chat Unsloth can now run multiple conversations at once. Starting a New Chat leaves the previous answer generating, and each active conversation gets its own progress indicator and Stop control. 4 llama-server slots by default, adjustable in the web UI.Tools, uploads, self-healing and agents stay isolated between chats.Stop one chat without interrupting others or restarting the server.Unsloth reduces the slot count when memory is limited.Reloading the model still stops active chats after confir…

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Unsloth

AMD-Support: Training und Inferenz auf AMD-GPUs

Unsloth ermöglicht jetzt lokales Training und lokale Inferenz auf AMD-GPUs unter Windows, WSL und Linux mit bis zu 2x Geschwindigkeit und 70% weniger VRAM, außerdem kamen im Juli-Update RDNA2-, Gorgon-Halo- und Vulkan-Support sowie Verbesserungen bei Erkennung und Installation hinzu.

Hey everyone! This release brings local LLM training and inference to AMD GPUs across Windows, WSL and Linux. Starting today, our AMD collaboration, custom Triton kernels, and math algorithms enables you to train and run 500+ models across AMD's Radeon, Instinct, Vulkan, Ryzen and data center GPUs, up to 2× faster with 70% less VRAM and no accuracy loss. Optimized ROCm builds also support GGUF & Safetensors inference. July 23 Update Added RDNA2, Gorgon Halo, Vulkan support + fixed AMD installing not detecting GPUs on Strix Halo / other AMD GPUsBetter RDNA4, HIP / ROCm failure auto fixing and catching2x faster unified memory AMD safetensors loading + much faster gradient checkpointing for unified memory devicesAdded voice dictation / whisper.cpp preliminary support for fast text to speechFixed rollback environments during installs eating 5GB of disk space - now auto cleans Train LLMs Locally on AMD Train, run RL, chat with and deploy models locally on AMD GPUs.More reliable AMD GPU detection and installation across Windows, WSL and Linux.Improved ROCm compatibility for AMD MI300X and MI325X GPUs.Remote access Unsloth via unsloth studio --secure through free HTTPS via Cloudflare Run Larger Models on Your Hardware Use automatic GPU placement or choose exactly which GPUs and model layers to use.Move MoE expert layers into system memory to help larger models fit.Split models across multiple GPUs or use Tensor Parallelism.Save hardware settings separately for each model and quant.…

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Unsloth

Personalisierung, NVFP4 und neue Anzeigesprachen

Unsloth bietet nun anpassbare Farbpaletten und Schriften, sieben neue Anzeigesprachen, einen Voice-Tab, eine vierstufige Tool-Call-Berechtigungsauswahl, GPU-beschleunigte Intel-Inferenz über Vulkan, Support für Inkling und erweiterte Dynamic-NVFP4-Quants.

Hey guys we got lots of new update for Unsloth, especially customization. Unsloth now yours to personalize: three color palettes plus custom colors and fonts, seven new display languages, and a new Voice settings tab for dictation and read-aloud. Agents get safer with a four-level tool-call permission selector (Ask, Approve for me, Off, Full access) and workspace isolation, and Intel GPUs finally get GPU-accelerated inference through new Vulkan llama.cpp support. We’ve also added native support for Inkling, a 975B-parameter open model with 41B active parameters and up to a 1M-token context window. Licensed under Apache 2.0, it accepts text, images, and audio and generates text. Dynamic NVFP4 We expanded our NVFP4 collection with quantized versions of Qwen3.6, Qwen3.5, Inkling, GLM-4.7 Flash, and Gemma 4. Read the Dynamic NVFP4 guide for details. Personalize Your Unsloth Unsloth is no longer limited to light and dark mode: Choose from Standard, Classic, and Minimal palettes, each with light and dark variants.Customize accent, background, and foreground colors.Import your own UI, heading, chat, and code fonts.Adjust font size, contrast, motion, cursor behavior, and font smoothing.Search settings and sync preferences across devices. Light and dark modes keep their own customization values. Seven New Languages Unsloth now supports French, German, Spanish, Hindi, Arabic, Russian, and Korean, alongside Chinese, Japanese, and Portuguese. Browser-language auto-detection is now the de…

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Unsloth

DeepSeek-V4 und NVFP4 in Unsloth

Unsloth unterstützt nun DeepSeek-V4-Flash, exportiert NVFP4-, FP8- und imatrix-GGUFs, kann als llama-swap-artiges API-System dienen, bietet Japanisch und brasilianisches Portugiesisch und macht GRPO 1,3x sowie MoE-Training 3–5x schneller.

Unsloth can now export NVFP4, FP8, and imatrix GGUFs after training; act as a llama-swap API system; add Japanese and Brazilian Portuguese support; and includes MLX, safetensors, tool calling, healing support, and more. Unsloth core makes GRPO 1.3x faster, adds HTTP fallback for stalled downloads, improves offline mode, speeds up MoE training by 3-5x, and fixes many bugs. This release series uses unsloth>=2026.7.1. DeepSeek-V4-Flash is now supported with Thinking toggles and our improved chat template fixes. Smarter OpenAI-Compatible API Serving Run one local API endpoint with safer model swapping and better agent-tool recovery. API requests can opt into automatic switching between downloaded local GGUFs, while unknown model names safely keep using the current model./v1/models now returns clean model IDs and the local GGUF catalog instead of local .gguf paths.Idle auto-unload can free VRAM after inactivity, and tool-call healing can now be controlled per request. Export Improvements Exports are more flexible and avoid unnecessary downloads. Select multiple export formats at once, including portable FP8/INT8, GGUF LoRA, source-matched exports, imatrix GGUF, and compressed FP8/FP4.Multi-checkpoint exports avoid more repeated base-model downloads.FP8, INT8, and GGUF-LoRA exports now respect trust_remote_code, and GGUF export handles missing quantization settings more reliably. RAG and File Chat File chat is more useful on real documents. RAG attachments can now use whole-docum…

Originalquelle(öffnet in neuem Tab)Problem melden