Zum Inhalt springen

Hugging Face Release Notes

65 Einträge aus 3 Quellen. Zuletzt aktualisiert:

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.10.4 (Version 5.10.3): Fixes für vLLM-Synchronisation

Das Patch-Release enthält mehrere Korrekturen für die Synchronisation mit vLLM, darunter Fixes für InternVL-Modelle, Token-IDs im ProcessorMixin, Offsets in der Verarbeitung, die PEFT-Untergrenze und den Mistral-Backend; auf PyPI erscheint es als 5.10.4, da 5.10.3 nicht existiert.

Patch release v5.10.4

Update: Note that on pypi 5.10.3 doesn't exist and this this saved under 5.10.4 (so essentially a minor version skipped). Sorry about that, that's on me. Just wanted to clarify to make this less confusing!

A few fixes needed for vLLM to sync with transformers :hugs:

  • [fix] regression introduced by #45534 #46456 by @eustlb (#46456)
  • Fix {image/video/audio}_token_ids in ProcessorMixin #46500 by @hmellor (#46500)
  • Fix InternVL models #46524 by @hmellor (#46524)
  • Fix the offsets in processing #46525 by @zucchini-nlp (#46525)
  • Fix peft lower bound #46605 by @hmellor (#46605)
  • mistral common backend fix #46667 by @itazap (#46667)

Full Changelog: https://github.com/huggingface/transformers/compare/v5.10.2...v5.10.3

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.12.0: Neues Modell MiniMax-M3-VL und weitere Ergänzungen

Transformers 5.12.0 fügt das Vision-Language-Modell MiniMax-M3-VL hinzu und aktualisiert die Dokumentation und Tests für PP-OCRv6.

Release v5.12.0

New Model additions

MiniMax-M3-VL

<img width="886" height="583" alt="image" src="https://github.com/user-attachments/assets/ae9dd96f-6877-4531-a06b-a756686f24e5" />

MiniMax-M3-VL is the vision-language member of the MiniMax-M3 family that pairs a CLIP-style vision tower with 3D rotary position embeddings with the MiniMax-M3 text backbone. It uses a mixed dense/sparse Mixture-of-Experts decoder with SwiGLU-OAI gated experts and a lightning indexer for block-sparse attention. The model processes images through a Conv3d patch embedding system and includes specialized components for efficient multimodal understanding and generation.

Links: Documentation

  • Add minimax m3vl (#46600) by @ArthurZucker in #46600

PP-OCRv6: update documentation and slow tests (#46576)

<img width="3840" height="1494" alt="image" src="https://github.com/user-attachments/assets/e62284ec-78bf-49cb-8aa2-deccc665372f" /> …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.11.0: Neue Modelle DiffusionGemma und DeepSeek-V3.2

Transformers 5.11.0 ergänzt DiffusionGemma, das Text blockweise per Diffusionssampler für schnellere Generierung erzeugt, sowie DeepSeek-V3.2.

Release v5.11.0

New Model additions

DiffusionGemma

<img width="1240" height="700" alt="image" src="https://github.com/user-attachments/assets/5081e449-6374-4076-bd96-d295c8334ca4" />

DiffusionGemma is engineered to reduce the sequential bottlenecks of standard causal language models by employing an encoder-decoder architecture specifically optimized for inference speed. During inference, DiffusionGemma leverages multi-canvas sampling, where rather than generating one token at a time, the model iteratively denoises a full block of tokens using a diffusion sampler. This block-autoregressive approach facilitates text generation at higher speeds compared to traditional sequential generation methods.

Links: Documentation

  • GPU go brr (#46540) by @gante in #46540

DeepSeek-V3.2

<img width="1135" height="671" alt="image" src="https://github.com/user-attachments/assets/24c9694d-eeae-402c-9a98-f7a3971dd9d0" /> …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.10.2: Korrektur der Modellkonvertierung für CLIP-Modelle

Das Patch-Release 5.10.2 behebt einen gravierenden Fehler bei der Konvertierung von CLIP-bezogenen Modellen wie sam3, weshalb ein Update empfohlen wird.

Patch release v5.10.2

There was a big bug in the model conversion of models related to clip, this affected models like sam3 and others. Please make sure to update :pray:

  • Fix conversion for clip models by @zucchini-nlp (#46406)

Full Changelog: https://github.com/huggingface/transformers/compare/v5.10.1...v5.10.2

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.10.1: Gemma 4 Unified und Gemma 4 MTP, 5.10.0 zurückgezogen

Version 5.10.1 ersetzt das zurückgezogene 5.10.0, das auf einem beschädigten Branch veröffentlicht wurde, und fügt Gemma 4 12B Unified, ein encoderfreies multimodales Modell, sowie Gemma 4 MTP hinzu.

Release v5.10.1

v5.10.0 was yanked as we publish on a corrupted branch. Sorry everyone, this happens when we rush a release!!!

New Model additions

Gemma4 unified+ Gemma4 MTP

<img width="2000" height="400" alt="image" src="https://github.com/user-attachments/assets/5e3ee940-f78d-4343-ac7a-889930800aa6" />

Gemma 4 12B Unified is an encoder-free multimodal model with pretrained and instruction-tuned variants. Unlike standard Gemma 4, which uses dedicated encoder towers, Gemma 4 12B Unified projects raw inputs directly into the language model's embedding space through lightweight linear pipelines. This results in a simpler architecture while maintaining strong multimodal performance.

Key differences from standard Gemma 4:

  • No Vision Tower: Raw pixel patches are projected directly into LM space via a Dense + LayerNorm pipeline with factorized 2D positional embeddings, replacing the vision encoder.
  • No Audio Tower: Raw 16 kHz waveform samples are chunked into fixed-length frames and projected through a simple RMSNorm → Linear pipeline, replacing the mel spectrogram + Conformer encoder.
  • Shared Multimodal Pipeline: Both vision and audio use the same Gemma4UnifiedMultimodalEmbedder (RMSNorm → Linear) for the final projection to text hidden space.

You can find the original Gemma 4 12B Unified checkpoints under the Gemma 4 release. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.9.0: Neue Modelle Cohere2Moe, Parakeet tdt und HRM-Text

Transformers 5.9.0 fügt neue Modelle hinzu, darunter das MoE-Modell Cohere2Moe (Command A+), Parakeet tdt und das autoregressive HRM-Text.

Release v5.9.0

New Model additions

Cohere2Moe

Command A+ is a Mixture-of-Experts (MoE) language model from Cohere that features a hybrid attention pattern combining sliding window and full attention layers. The model incorporates both shared and routed experts and supports a very large context window for processing extensive text sequences.

Links: Documentation

  • Add new cohere2_moe model (#46115) by @Cyrilvallez in #46115

Parakeet tdt (#44171)

  • Parakeet tdt (#44171) by @lmaksym

HRM-Text

HRM-Text is an improved autoregressive language-modeling variant of the Hierarchical Reasoning Model (HRM) that uses a hierarchical recurrent forward pass with two transformer stacks - one for slow, abstract planning (H) and one for fast, detailed computation (L) - reused inside a nested recurrence. It features PrefixLM attention where instruction tokens attend bidirectionally while response tokens attend causally, per-head sigmoid output gates, and parameterless RMSNorm. The model is designed as a base language model without instruction tuning or chat templates.

Links: Documentation | Paper …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.8.1: Fix der DeepSeek-V4-Integration

Das Patch-Release 5.8.1 korrigiert hauptsächlich die DeepSeek-V4-Integration und behebt zudem einen falschen WeightConverter-Regex-Treffer bei shared_experts sowie einen fatal_error-Fall im ContinuousBatchingManager.

Patch release v5.8.1

This release is mainly to fix the Deepseek V4 integration!!!

<img width="714" height="774" alt="image" src="https://github.com/user-attachments/assets/0d85e891-a0ff-436e-a9d4-b6633096f2b5" />
  • [fix] Add fatal_error to ContinuousBatchingManager so the serving... by @qgallouedec, @remi-or
  • Fix WeightConverter regex incorrectly matching shared_experts as experts by @silencelamb, @claude
  • Fix deepseek v4 by @ArthurZucker (#45892)
  • Deepseek v4 csa mask collapse by @ArthurZucker, @Sawyer117 (#45928)

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.8.0: DeepSeek-V4 und Gemma 4 Assistant

Transformers 5.8.0 fügt DeepSeek-V4 (Flash, Pro und Base-Varianten) sowie Gemma 4 Assistant als neue Modelle hinzu.

Release v5.8.0

New Model additions

DeepSeek-V4

<img width="6604" height="3574" alt="image" src="https://github.com/user-attachments/assets/4c0fdb29-f770-463c-a97b-d24438896a4c" />

DeepSeek-V4 is the next-generation MoE (Mixture of Experts) language model from DeepSeek that introduces several architectural innovations over DeepSeek-V3. The architecture replaces Multi-head Latent Attention (MLA) with a hybrid local + long-range attention design, swaps residual connections for Manifold-Constrained Hyper-Connections (mHC), and bootstraps the first few MoE layers with a static token-id → expert-id hash table. This implementation covers DeepSeek-V4-Flash, DeepSeek-V4-Pro, and their -Base pretrained variants, which share the same architecture but differ in width, depth, expert count and weights.

Links: Documentation | Paper

  • Add DeepSeek V4 (#45643) by @ArthurZucker in #45643

Gemma 4 Assistant

<img width="2000" height="400" alt="image" src="https://github.com/user-attachments/assets/02c79b0b-a172-4495-b09d-a6a4b625ee66" /> …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.38.0 mit neuen Bild- und Audio-Pipelines

Diffusers 0.38.0 ergänzt neue Pipelines wie LLaDA2, Nucleus-MoE, Ernie-Image und LongCat-AudioDiT sowie Verbesserungen an der Kernbibliothek.

New Pipelines

LLaDA2

LLaDA2 is a family of discrete diffusion language models that generate text through block-wise iterative refinement. Instead of autoregressive token-by-token generation, LLaDA2 starts with a fully masked sequence and progressively unmasks tokens by confidence over multiple refinement steps.

Nucleus-MoE

NucleusMoE-Image is a 2B active 17B parameter model trained with efficiency at its core. Our novel architecture highlights the scalability of a sparse MoE architecture for Image generation.

Thanks to @sippycoder for the contribution.

Ernie-Image

ERNIE-Image is a powerful and highly efficient image generation model with 8B parameters.

Thanks to @HsiaWinter for the contribution.

LongCat-AudioDiT …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.7.0: Neue Modelle Laguna und DEIMv2

Transformers 5.7.0 ergänzt das Mixture-of-Experts-Sprachmodell Laguna von Poolside sowie das Modell DEIMv2.

Release v5.7.0

New Model additions

Laguna

<img width="699" height="176" alt="image" src="https://github.com/user-attachments/assets/d3bae269-bea7-4ddf-a53f-d4718befdb17" />

Laguna is Poolside's mixture-of-experts language model family that extends standard SwiGLU MoE transformers with two key innovations. It features per-layer head counts allowing different decoder layers to have different query-head counts while sharing the same KV cache shape, and implements a sigmoid MoE router with auxiliary-loss-free load balancing that uses element-wise sigmoid of gate logits plus learned per-expert bias for router scoring.

Links: Documentation

  • Laguna XS.2 implementation (#45673) by @joerowell in #45673

DEIMv2

<img width="2874" height="908" alt="image" src="https://github.com/user-attachments/assets/fc8c59fe-f964-42ce-ae8e-c7fcace9beb7" /> …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.6.2: Qwen 3.5 und 3.6 MoE mit FP8 funktionieren wieder

Das Patch-Release 5.6.2 behebt, dass Qwen 3.5 und 3.6 MoE (nur Text) mit FP8 nicht funktionierten, durch eine korrigierte Konfigurationsauslese und Fehlerbehandlung für kernels.

Patch release v5.6.2

Qwen 3.5 and 3.6 MoE (text-only) were broken when using with FP8. It should now work again with this :saluting_face:

Full Changelog: https://github.com/huggingface/transformers/compare/v5.6.1...v5.6.2

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.6.1: Fix für defekten Flash-Attention-Pfad

Das Patch-Release 5.6.1 behebt einen AttributeError bei s_aux=None in flash_attention_forward, durch den der Flash-Attention-Pfad nicht funktionierte.

Patch release v5.6.1

Flash attention path was broken! Sorry everyone for this one 🤗

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.6.0: Neue Modelle OpenAI Privacy Filter und QianfanOCR

Transformers 5.6.0 fügt den OpenAI Privacy Filter zur Erkennung und Maskierung personenbezogener Daten in Text sowie das OCR-Modell QianfanOCR von Baidu hinzu.

Release v5.6.0

New Model additions

OpenAI Privacy Filter

OpenAI Privacy Filter is a bidirectional token-classification model for personally identifiable information (PII) detection and masking in text. It is intended for high-throughput data sanitization workflows where teams need a model that they can run on-premises that is fast, context-aware, and tunable. The model labels an input sequence in a single forward pass, then decodes coherent spans with a constrained Viterbi procedure, predicting probability distributions over 8 privacy-related output categories for each input token.

Links: Documentation

  • [Privacy Filter] Add model (#45580) by @vasqu in #45580

QianfanOCR

Qianfan-OCR is a 4B-parameter end-to-end document intelligence model developed by Baidu that performs direct image-to-text conversion without traditional multi-stage OCR pipelines. It supports a broad range of prompt-driven tasks including structured document parsing, table extraction, chart understanding, document question answering, and key information extraction all within one unified model. The model features a unique "Layout-as-Thought" capability that generates structured layout representations before producing final outputs, making it particularly effective for complex documents with mixed element types. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.5.4: Fixes für Tokenizer, Training und Qwen2.5-VL

Das Patch-Release 5.5.4 behebt eine Tokenizer-Regression bei Kimi-K2.5, einen IndexError mit DeepSpeed ZeRO-3 bei aktiven Rotary-Kernels, ein Problem bei GAS im Training und eine fälschlich auf Standbilder angewandte temporale RoPE-Skalierung bei Qwen2.5-VL.

Patch release v5.5.4

This is mostly some fixes that are good to have asap, mostly for tokenizers; ** Fix Kimi-K2.5 tokenizer regression and _patch_mistral_regex Attribute… (#45305) by ArthurZucker

For training: ** Fix #45305 + add regression test GAS (#45349) by florian6973, SunMarc ** Fix IndexError with DeepSpeed ZeRO-3 when kernels rotary is active (#…) by ArthurZucker

And for Qwen2.5-VL : ** Fix Qwen2.5-VL temporal RoPE scaling applied to still images (#45330) by Kash6, zucchini-nlp

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.5.3: device_map-Unterstützung für Gemma4 repariert

Das kleine Patch-Release 5.5.3 behebt die Unterstützung von device_map (auto) für Gemma4.

Small patch release to fix device_map support for Gemma4! It contains the following commit:

  • [gemma4] Fix device map auto (#45347) by @Cyrilvallez

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.5.2: Gemma4-Optimierungen und Fixes

Das Patch-Release 5.5.2 optimiert gemma4, behebt die Inferenz mit use_cache=False trotz geteilter k/v-Zustände zwischen Layern, ergänzt MoE im Gemma4-TP-Plan und korrigiert Conversion Mappings für VLMs mit uneinheitlich serialisierten Gewichtsnamen.

Small patch dedicated to optimizing gemma4, fixing inference with use_cache=False due to k/v states sharing between layers, as well as conversion mappings for some models that would inconsistently serialize their weight names. It contains the following PRs:

  • Add MoE to Gemma4 TP plan (#45219) by @sywangyi and @Cyrilvallez
  • [gemma4] Dissociate kv states sharing from the Cache (#45312) by @Cyrilvallez
  • [gemma4] Remove all shared weights, and silently skip them during loading (#45336) by @Cyrilvallez
  • Fix conversion mappings for vlms (#45340) by @Cyrilvallez

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.5.1: Fixes für vLLM und Gemma4

Das kleine Patch-Release 5.5.1 behebt den Export für gemma4, ergänzt Integrationstests und korrigiert vLLM-CI-Probleme.

Patch release v5.5.1

This patch is very small and focuses on vLLM and Gemma4!

** Fix export for gemma4 and add Integration tests (#45285) by @Cyrilvallez ** Fix vllm cis (#45139) by @ArthurZucker

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.5.0: Neues Modell Gemma4

Version 5.5.0 fügt das multimodale Modell Gemma4 hinzu, das Bilder unterschiedlicher Größe mit einem festen Token-Budget verarbeitet und dafür eine Bildverarbeitung sowie 2D-RoPE nutzt.

Release v5.5.0

<img width="2786" height="1504" alt="image" src="https://github.com/user-attachments/assets/6c8c878f-042b-4858-9f64-73fd9ccd7e4b" />

New Model additions

Gemma4

Gemma 4 is a multimodal model with pretrained and instruction-tuned variants, available in 1B, 13B, and 27B parameters. The architecture is mostly the same as the previous Gemma versions. The key differences are a vision processor that can output images of fixed token budget and a spatial 2D RoPE to encode vision-specific information across height and width axis.

<img width="1478" height="1374" alt="image" src="https://github.com/user-attachments/assets/9d88bd1b-02ea-4829-b7d0-fac0e347d436" />

You can find all the original Gemma 4 checkpoints under the Gemma 4 release.

The key difference from previous Gemma releases is the new design to process images of different sizes using a fixed-budget number of tokens. Unlike many models that squash every image into a fixed square (like 224×224), Gemma 4 keeps the image's natural aspect ratio while making it the right size. There a a couple constraints to follow:

  • The total number of pixels must fit within a patch budget
  • Both height and width must be divisible by 48 (= patch size 16 × pooling kernel 3)

[!IMPORTANT] …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.4.0: PaddlePaddle-Modelle, Mistral 4, VidEoMT und mehr

Version 5.4.0 ergänzt neue Modelle, darunter VidEoMT, einen leichtgewichtigen Encoder-only-Transformer für Online-Videosegmentierung, sowie UVDoc, PI0, SLANeXt, Mistral 4 und Jina Embeddings v3.

New Model additions

VidEoMT

<img width="1480" height="460" alt="image" src="https://github.com/user-attachments/assets/bec6fc25-b0ab-4227-8c2b-a838554f37f3" />

Video Encoder-only Mask Transformer (VidEoMT) is a lightweight encoder-only model for online video segmentation built on a plain Vision Transformer (ViT). It eliminates the need for dedicated tracking modules by introducing a lightweight query propagation mechanism that carries information across frames and employs a query fusion strategy that combines propagated queries with temporally-agnostic learned queries. VidEoMT achieves competitive accuracy while being 5x-10x faster than existing approaches, running at up to 160 FPS with a ViT-L backbone.

Links: Documentation | Paper

  • Add VidEoMT (#44285) by @NielsRogge in #44285

UVDoc

<img width="1765" height="875" alt="image" src="https://github.com/user-attachments/assets/365e510e-8fb8-46cb-8f4b-e8b7082f0ae2" /> …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers Version 0.37.1

Das Update behebt das Laden von ModularPipelines mit AutoModel-Typhinweisen, das Laden von Flux-Klein-LoRA und einen ungeschützten torchvision-Import in Cosmos Predict 2.5.

  • Fix for loading ModularPipelines with AutoModel type hints in their modular_model_index.json #13271
  • Fix Flux Klein LoRA loading #13313
  • Fix unguarded torchvision import in Cosmos Predict 2.5 #13321

Originalquelle(öffnet in neuem Tab)Problem melden