Zum Inhalt springen

Transformers Updates & Release Notes

30 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Folge Transformers, um die Release Notes in deinen Feed zu holen.

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.7.0: Neue Modelle Laguna und DEIMv2

Transformers 5.7.0 ergänzt das Mixture-of-Experts-Sprachmodell Laguna von Poolside sowie das Modell DEIMv2.

Release v5.7.0

New Model additions

Laguna

<img width="699" height="176" alt="image" src="https://github.com/user-attachments/assets/d3bae269-bea7-4ddf-a53f-d4718befdb17" />

Laguna is Poolside's mixture-of-experts language model family that extends standard SwiGLU MoE transformers with two key innovations. It features per-layer head counts allowing different decoder layers to have different query-head counts while sharing the same KV cache shape, and implements a sigmoid MoE router with auxiliary-loss-free load balancing that uses element-wise sigmoid of gate logits plus learned per-expert bias for router scoring.

Links: Documentation

  • Laguna XS.2 implementation (#45673) by @joerowell in #45673

DEIMv2

<img width="2874" height="908" alt="image" src="https://github.com/user-attachments/assets/fc8c59fe-f964-42ce-ae8e-c7fcace9beb7" /> …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.6.2: Qwen 3.5 und 3.6 MoE mit FP8 funktionieren wieder

Das Patch-Release 5.6.2 behebt, dass Qwen 3.5 und 3.6 MoE (nur Text) mit FP8 nicht funktionierten, durch eine korrigierte Konfigurationsauslese und Fehlerbehandlung für kernels.

Patch release v5.6.2

Qwen 3.5 and 3.6 MoE (text-only) were broken when using with FP8. It should now work again with this :saluting_face:

Full Changelog: https://github.com/huggingface/transformers/compare/v5.6.1...v5.6.2

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.6.1: Fix für defekten Flash-Attention-Pfad

Das Patch-Release 5.6.1 behebt einen AttributeError bei s_aux=None in flash_attention_forward, durch den der Flash-Attention-Pfad nicht funktionierte.

Patch release v5.6.1

Flash attention path was broken! Sorry everyone for this one 🤗

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.6.0: Neue Modelle OpenAI Privacy Filter und QianfanOCR

Transformers 5.6.0 fügt den OpenAI Privacy Filter zur Erkennung und Maskierung personenbezogener Daten in Text sowie das OCR-Modell QianfanOCR von Baidu hinzu.

Release v5.6.0

New Model additions

OpenAI Privacy Filter

OpenAI Privacy Filter is a bidirectional token-classification model for personally identifiable information (PII) detection and masking in text. It is intended for high-throughput data sanitization workflows where teams need a model that they can run on-premises that is fast, context-aware, and tunable. The model labels an input sequence in a single forward pass, then decodes coherent spans with a constrained Viterbi procedure, predicting probability distributions over 8 privacy-related output categories for each input token.

Links: Documentation

  • [Privacy Filter] Add model (#45580) by @vasqu in #45580

QianfanOCR

Qianfan-OCR is a 4B-parameter end-to-end document intelligence model developed by Baidu that performs direct image-to-text conversion without traditional multi-stage OCR pipelines. It supports a broad range of prompt-driven tasks including structured document parsing, table extraction, chart understanding, document question answering, and key information extraction all within one unified model. The model features a unique "Layout-as-Thought" capability that generates structured layout representations before producing final outputs, making it particularly effective for complex documents with mixed element types. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.5.4: Fixes für Tokenizer, Training und Qwen2.5-VL

Das Patch-Release 5.5.4 behebt eine Tokenizer-Regression bei Kimi-K2.5, einen IndexError mit DeepSpeed ZeRO-3 bei aktiven Rotary-Kernels, ein Problem bei GAS im Training und eine fälschlich auf Standbilder angewandte temporale RoPE-Skalierung bei Qwen2.5-VL.

Patch release v5.5.4

This is mostly some fixes that are good to have asap, mostly for tokenizers; ** Fix Kimi-K2.5 tokenizer regression and _patch_mistral_regex Attribute… (#45305) by ArthurZucker

For training: ** Fix #45305 + add regression test GAS (#45349) by florian6973, SunMarc ** Fix IndexError with DeepSpeed ZeRO-3 when kernels rotary is active (#…) by ArthurZucker

And for Qwen2.5-VL : ** Fix Qwen2.5-VL temporal RoPE scaling applied to still images (#45330) by Kash6, zucchini-nlp

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.5.3: device_map-Unterstützung für Gemma4 repariert

Das kleine Patch-Release 5.5.3 behebt die Unterstützung von device_map (auto) für Gemma4.

Small patch release to fix device_map support for Gemma4! It contains the following commit:

  • [gemma4] Fix device map auto (#45347) by @Cyrilvallez

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.5.2: Gemma4-Optimierungen und Fixes

Das Patch-Release 5.5.2 optimiert gemma4, behebt die Inferenz mit use_cache=False trotz geteilter k/v-Zustände zwischen Layern, ergänzt MoE im Gemma4-TP-Plan und korrigiert Conversion Mappings für VLMs mit uneinheitlich serialisierten Gewichtsnamen.

Small patch dedicated to optimizing gemma4, fixing inference with use_cache=False due to k/v states sharing between layers, as well as conversion mappings for some models that would inconsistently serialize their weight names. It contains the following PRs:

  • Add MoE to Gemma4 TP plan (#45219) by @sywangyi and @Cyrilvallez
  • [gemma4] Dissociate kv states sharing from the Cache (#45312) by @Cyrilvallez
  • [gemma4] Remove all shared weights, and silently skip them during loading (#45336) by @Cyrilvallez
  • Fix conversion mappings for vlms (#45340) by @Cyrilvallez

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.5.1: Fixes für vLLM und Gemma4

Das kleine Patch-Release 5.5.1 behebt den Export für gemma4, ergänzt Integrationstests und korrigiert vLLM-CI-Probleme.

Patch release v5.5.1

This patch is very small and focuses on vLLM and Gemma4!

** Fix export for gemma4 and add Integration tests (#45285) by @Cyrilvallez ** Fix vllm cis (#45139) by @ArthurZucker

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.5.0: Neues Modell Gemma4

Version 5.5.0 fügt das multimodale Modell Gemma4 hinzu, das Bilder unterschiedlicher Größe mit einem festen Token-Budget verarbeitet und dafür eine Bildverarbeitung sowie 2D-RoPE nutzt.

Release v5.5.0

<img width="2786" height="1504" alt="image" src="https://github.com/user-attachments/assets/6c8c878f-042b-4858-9f64-73fd9ccd7e4b" />

New Model additions

Gemma4

Gemma 4 is a multimodal model with pretrained and instruction-tuned variants, available in 1B, 13B, and 27B parameters. The architecture is mostly the same as the previous Gemma versions. The key differences are a vision processor that can output images of fixed token budget and a spatial 2D RoPE to encode vision-specific information across height and width axis.

<img width="1478" height="1374" alt="image" src="https://github.com/user-attachments/assets/9d88bd1b-02ea-4829-b7d0-fac0e347d436" />

You can find all the original Gemma 4 checkpoints under the Gemma 4 release.

The key difference from previous Gemma releases is the new design to process images of different sizes using a fixed-budget number of tokens. Unlike many models that squash every image into a fixed square (like 224×224), Gemma 4 keeps the image's natural aspect ratio while making it the right size. There a a couple constraints to follow:

  • The total number of pixels must fit within a patch budget
  • Both height and width must be divisible by 48 (= patch size 16 × pooling kernel 3)

[!IMPORTANT] …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.4.0: PaddlePaddle-Modelle, Mistral 4, VidEoMT und mehr

Version 5.4.0 ergänzt neue Modelle, darunter VidEoMT, einen leichtgewichtigen Encoder-only-Transformer für Online-Videosegmentierung, sowie UVDoc, PI0, SLANeXt, Mistral 4 und Jina Embeddings v3.

New Model additions

VidEoMT

<img width="1480" height="460" alt="image" src="https://github.com/user-attachments/assets/bec6fc25-b0ab-4227-8c2b-a838554f37f3" />

Video Encoder-only Mask Transformer (VidEoMT) is a lightweight encoder-only model for online video segmentation built on a plain Vision Transformer (ViT). It eliminates the need for dedicated tracking modules by introducing a lightweight query propagation mechanism that carries information across frames and employs a query fusion strategy that combines propagated queries with temporally-agnostic learned queries. VidEoMT achieves competitive accuracy while being 5x-10x faster than existing approaches, running at up to 160 FPS with a ViT-L backbone.

Links: Documentation | Paper

  • Add VidEoMT (#44285) by @NielsRogge in #44285

UVDoc

<img width="1765" height="875" alt="image" src="https://github.com/user-attachments/assets/365e510e-8fb8-46cb-8f4b-e8b7082f0ae2" /> …

Originalquelle(öffnet in neuem Tab)Problem melden

Transformers Updates & Release Notes (Hugging Face) – Oktober 2026 (Seite 2) | updatefeed