Zum Inhalt springen

Transformers Updates & Release Notes

30 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Folge Transformers, um die Release Notes in deinen Feed zu holen.

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers v5.19.0: EmbeddingGemma2, Router-Logits als Breaking Change

Transformers v5.19.0 ergänzt das multimodale Embedding-Modell EmbeddingGemma2, das Text, Bilder, Audio und Video in einen gemeinsamen 768-dimensionalen Vektorraum abbildet, und gibt bei allen MoE-Modellen mit Router-Logits diese nun bei output_router_logits=True zurück, was ein Breaking Change ist.

Release v5.19.0

New Model additions

EmbeddingGemma2

<img width="2716" height="2308" alt="image" src="https://github.com/user-attachments/assets/84734e74-163d-4d12-b166-ffcf6749d563" />

EmbeddingGemma 2 is a multimodal embedding model from Google built on the Gemma 4 architecture. It encodes text, images, audio, and video, individually or combined in one input, into a shared 768-dimensional vector space for cross-modal retrieval, semantic similarity, clustering, and classification. It uses Matryoshka Representation Learning, so embeddings can be truncated to 512, 256, or 128 dimensions. It also offers configurable visual and video token budgets, and unused vision or audio towers can be disabled at load time to save memory.

Links: Documentation

  • Smthn smthn (#49364) by @vasqu in #49364

Breaking changes

All MoE models whose routers compute logits now return them when output_router_logits=True, following the Qwen3-MoE pattern (a router_logits recorder on the base model, MoeModelOutputWithPast from the backbone, and a MoE causal LM output from the head), so code that relied on the previous outputs or their absence should read the router logits from these output classes.

  • 🚨 Return router logits from every MoE model that computes them (#48920) by @qgallouedec …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.18.0: Nemotron 3 Diarization und NemotronH Omni

Transformers 5.18.0 ergänzt unter anderem die neuen Modelle Nemotron 3 Diarization für Streaming-Sprecherdiarisierung mit bis zu acht Sprechern sowie NemotronH Omni von NVIDIA.

New Model additions

Nemotron 3 Diarization

<img width="1680" height="900" alt="image" src="https://github.com/user-attachments/assets/fe735cb3-9e5b-43ad-8f60-9dec8425aec7" />

Nemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine "who spoke when" in real-world audio. It supports both streaming and offline inference, handles up to eight speakers, and orders speaker outputs by each speaker's first arrival in the input audio.

The model uses the Arrival-Order Speaker Cache (AOSC) 1 and FIFO queue introduced for Streaming Sortformer 1, 2. A single checkpoint supports configurable latency profiles, from an 80 ms input buffer to a 30.4 s offline-style buffer, and configurable output frame resolution in multiples of 10 ms. With chunked inference, the maximum audio duration is not limited.

Links: Documentation

  • Add Nemotron3Diarization (#49056) by @eustlb in #49056

NemotronH Omni

NemotronH Omni is a multimodal reasoning model from NVIDIA that pairs the NemotronH hybrid …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.17.0: neues Modell HYV4 (Hy4-Preview)

Transformers 5.17.0 ergänzt das neue Modell HYV4 (Hy4-Preview), ein Mixture-of-Experts-Sprachmodell mit 780B Parametern, 49B aktiven Parametern pro Token und 1M Tokens Kontextfenster.

Release v5.17.0

New Model additions

HYV4

<img width="1503" height="827" alt="image" src="https://github.com/user-attachments/assets/e6ed85ee-eb1d-40eb-a0d4-c649f6337ca9" />

Hy4-Preview is a 780B-parameter mixture-of-experts language model that activates 49B parameters per token. Each MoE layer holds 256 routed experts plus one always-active shared expert and routes every token to 8 of them. The context window is 1M tokens.

The architecture combines four features:

  • Multi-head Latent Attention (MLA) compresses keys and values into a low-rank latent (kv_lora_rank) that kv_b_proj expands back to one key/value per query head.
  • DeepSeek Sparse Attention (DSA) selects index_topk keys per query with a lightweight indexer. Following IndexShare, only the layers marked "full" in indexer_types run an indexer; "shared" layers reuse the previous full layer's selection.
  • Gated MLA with learnable attention sinks, where each head owns a sink logit that participates in the softmax and contributes no value, as in GPT-OSS.
  • Independent Hyper-Connections (iHC) replace the plain residual path with hc_mult parallel residual streams that are collapsed before, and redistributed after, every sublayer.

The implementation does not execute the multi-token prediction (MTP) layers. Released checkpoints …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.16.1: GLM-5.3-Flash und kleine Fixes

Transformers v5.16.1 unterstützt nun GLM-5.3-Flash, ein nativ multimodales Modell der GLM-5-Reihe, und enthält kleine Fixes zur Rückwärtskompatibilität bei TP sowie ein angepinntes hf-Kernel aus Sicherheitsgründen.

Release v5.16.1

This is a special release as we include GLM! (and a few small fixes)

GLM-5.3-Flash

<img width="4239" height="2643" alt="image" src="https://github.com/user-attachments/assets/17bc9c29-758b-44c8-8230-42f945ded209" />

GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.

GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Together with our latest 30T-token multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver more intelligence with less compute.

Links: Documentation

  • [Glm 5.3 Flash] GLM 5.3 Flash Support (#48342) by @Dovis01 in #48342

Small patch fixes

Mainly BC behavior for TP and pinning a hf kernel for security reasons :hugs: …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.16.0: neues Modell Qwen4-Exp

Transformers v5.16.0 ergänzt das neue Modell Qwen4-Exp, das auf der Qwen3.5-Architektur aufbaut und lineare mit sparsamer Attention kombiniert.

Release v5.16.0

New Model additions

Qwen4-Exp

<img width="2241" height="693" alt="image" src="https://github.com/user-attachments/assets/c838b5ba-ffea-42da-baa9-3f66178e3671" />

Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE).

GR is a Qwen-developed residual architecture that combines Hyper-Connection with GatedNorm. It mixes multiple residual streams with fine-grained elementwise gating before each attention and Mixture-of-Experts (MoE) block, then controls how much of the block output is injected back into each stream.

QSA uses multiple query heads to score compressed key blocks, selects the most relevant contiguous token blocks, and keeps the incomplete trailing block uncompressed. This block-level selection reduces indexing overhead and improves memory locality for long sequences. Combined with Gated DeltaNet, QSA makes Qwen4-Exp the first hybrid architecture to integrate linear and sparse attention, substantially improving inference efficiency for long-context workloads.

PLE enriches selected decoder layers with layer-specific lexical features derived from hashed token n-grams and a dilated depthwise convolution.

Links: Documentation …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.15.1: Fixes für DFlash, MTP, Gemma4 und Lanczos

Transformers 5.15.1 behebt Probleme mit DFlash- und MTP-Candidate-Generatoren, einen Gerätekonflikt bei Gemma4-Videos sowie Fehler bei der Bildverarbeitung auf Beschleunigern mit Lanczos-Filter, indem auf bicubic ausgewichen wird.

Patch release v5.15.1

This patch most notably solves a few issues with DFlash and MTP candidate generators, as well as an issue where images could sometimes not be processed on accelerator if using Lanczos filter.

It contains the following commits:

  • Fix DFlash candidate token device mismatch with device_map="auto" (#47877) by @sywangyi and @Cyrilvallez
  • Align logit distributions for CandidateGenerators using sampling (#48007) by @Cyrilvallez
  • Fix MTP config when mlp_layer_types is absent (#48015) by @Cyrilvallez
  • Fallback from 'lanczos' to 'bicubic' when on cuda (#48026) by @zucchini-nlp
  • Fix gemma4 video to device (#47896) by @guarin

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.15.0: Muse Glimmer, GraniteMoeSWA und weitere Modelle

Transformers v5.15.0 ergänzt neue Modelle, darunter Metas multimodales Muse Glimmer mit 30B Parametern unter Apache-2.0-Lizenz sowie GraniteMoeSWA, GraniteSWA und A.X-K1/A.X-K2.

Release v5.15.0

New Model additions

Meta Muse Glimmer

Muse Glimmer, released today, is Meta’s new multimodal model, especially designed for agentic use cases. Distilled from Muse to 30B parameters, and released under the Apache 2.0 license, it can be deployed to local setups for privacy-aware applications such as coding, document analysis, personal assistants, Claw- or Hermes-like setups.

Muse Glimmer is a dense 30B parameter model consisting of:

  • 2B ViT-style encoder for vision (Perception Encoder)
  • 28B parameter text decoder

We're covering it in the following blogpost: http://hf.co/blog/muse-glimmer

<img width="960" height="1787" alt="image" src="https://github.com/user-attachments/assets/3d8e548e-f84f-4269-8bd0-a12722d7ab01" />

GraniteMoeSWA & GraniteSWA

<img width="1013" height="389" alt="image" src="https://github.com/user-attachments/assets/2c2b87f0-466a-413a-a4be-25ceae49c9a5" />

Links: Documentation

  • Add Granite-swa and Granitemoe-swa model support (#47179) by @daviswer in #47179

Links: Documentation

  • Add Granite-swa and Granitemoe-swa model support (#47179) by @daviswer in #47179

A.X-K1 & A.X-K2 …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.14.1: Fixes für Inkling, FP8-Kernels und deepgemm

Transformers 5.14.1 behebt Probleme bei der Inkling-Integration, darunter Fehler bei assistierter Generierung mit EncoderDecoderCache und beim sdpa-Prefill mit position_bias, und enthält zudem Fixes für FP8-Kernels und deepgemm auf mehreren Geräten.

Patch release v5.14.1

This patch solves a few issues which appeared when integrating Inkling model, most notably an issue affecting models using EncoderDecoderCache during assisted generation. It also fixes an issue that could appear during prefill with StaticCache and sdpa without padding for Inkling which uses a position_bias. It contains the following commits:

  • Fix sdpa prefill with position_bias (#47359) by @Cyrilvallez
  • Fix assisted decoding for models with EncoderDecoder cache & OlmoHybrid (#47361) by @Cyrilvallez
  • [FP8] Bump kernels version (#47344) by @vasqu
  • Fix deepgemm on multiple devices (#47323) by @IlyasMoutawwakil

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.14.0: Inkling, TIPSv2 und TIPSv2 DPT

Transformers v5.14.0 ergänzt neue Modelle, darunter das multimodale Inkling von Thinking Machines mit 975B Parametern (41B aktiv) sowie TIPSv2 und TIPSv2 DPT.

Release v5.14.0

New Model additions

Inkling (fresh from Thinking Machines): 975B total, 41B active

  • Add Inkling model #47347 by @molbap @Cyrilvallez @eustlb and @zucchini-nlp
<img width="3840" height="2160" alt="image" src="https://github.com/user-attachments/assets/051f819a-512f-4987-9bee-6e2fa2af3db7" />

Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI- powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-following, and other natural language and multimodal tasks. It is released with open weights to support research, fine-tuning and integration into third-party products by downstream developers.

TIPSv2

<img width="1555" height="1306" alt="image" src="https://github.com/user-attachments/assets/2d9f21e5-05f8-4c36-93ef-22f03c089f52" />

Links: Documentation

  • Add TIPSv2 (#46347) by @Ternura143 in #46347

TIPSv2 DPT …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.13.1: Kompatibilität mit der neuesten vllm-Version

Transformers 5.13.1 sorgt für Kompatibilität mit der neuesten vllm-Version, indem Handhabung von Legacy-Layer-Typen, Custom Code und str-Schlüsseln in _LazyAutoMapping.register korrigiert wird.

Patch release v5.13.1

This patch is focused on enabling transformers for the latest release of vllm!

  • Be more defensive with remap_legacy_layer_types for custom models (#47245) from @hmellor
  • Fix custom code which doesn't know about the new linear layer type names (#47174) from @hmellor
  • Fix case where _LazyAutoMapping.register is passed a str key (#47148) from @hmellor

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.13.0: Neue Modelle, darunter Kimi K2.5 bis K2.7

Transformers 5.13.0 fügt unter anderem die Architektur für Kimi K2.5 hinzu, die auch von K2.6 und K2.7 genutzt wird, sowie weitere neue Modelle wie MiMo-V2-Flash.

Release v5.13.0

New Model additions

KimiK 2.5, 2.6, and 2.7

<img width="1097" height="400" alt="image" src="https://github.com/user-attachments/assets/c24d2232-a9b4-413b-a2c8-58d013b6dfbd" />

This release includes the architecture for Kimi 2.5 which is used by 2.5-2.7:

Kimi K2.5 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration. The model was proposed in Kimi K2.5: Visual Agentic Intelligence and further improved in [Kimi K2.6: Advancing Open-Source Coding](Kimi K2.5: Visual Agentic Intelligence).

Kimi K2.5 achieves significant improvements on complex, end-to-end coding tasks, generalizing robustly across programming languages (Rust, Go, Python) and domains spanning front-end, DevOps, and performance optimization. The model is capable of transforming simple prompts and visual inputs into production-ready interfaces and lightweight full-stack workflows, generating structured layouts, interactive elements, and rich animations with deliberate aesthetic precision.

Links: Documentation

  • Add new model: Kimi2-6 (#45630) by @zucchini-nlp in #45630

MiMo-V2-Flash …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.12.1: PEFT-Untergrenze und Mistral-Tokenizer-Fix

Das Patch-Release 5.12.1 hebt die Untergrenze für PEFT an und behebt einen Fehler, durch den der Auto-Tokenizer den Mistral-Tokenizer bei installiertem mistral-common nicht korrekt auflöste.

Patch release v5.12.1

Updated the lower bound for PEFT and a fix for auto tokenizer to properly resolve the mistral tokenizer (when mistral-common is installed). This is similar to v.5.10.3 minus the fixes that were already included in the main release - vLLM will first target 5.10.3 :hugs:

  • Fix peft lower bound #46605 by @hmellor (#46605)
  • mistral common backend fix #46667 by @itazap (#46667)

Full Changelog: https://github.com/huggingface/transformers/compare/v5.12.0...v5.12.1

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.10.4 (Version 5.10.3): Fixes für vLLM-Synchronisation

Das Patch-Release enthält mehrere Korrekturen für die Synchronisation mit vLLM, darunter Fixes für InternVL-Modelle, Token-IDs im ProcessorMixin, Offsets in der Verarbeitung, die PEFT-Untergrenze und den Mistral-Backend; auf PyPI erscheint es als 5.10.4, da 5.10.3 nicht existiert.

Patch release v5.10.4

Update: Note that on pypi 5.10.3 doesn't exist and this this saved under 5.10.4 (so essentially a minor version skipped). Sorry about that, that's on me. Just wanted to clarify to make this less confusing!

A few fixes needed for vLLM to sync with transformers :hugs:

  • [fix] regression introduced by #45534 #46456 by @eustlb (#46456)
  • Fix {image/video/audio}_token_ids in ProcessorMixin #46500 by @hmellor (#46500)
  • Fix InternVL models #46524 by @hmellor (#46524)
  • Fix the offsets in processing #46525 by @zucchini-nlp (#46525)
  • Fix peft lower bound #46605 by @hmellor (#46605)
  • mistral common backend fix #46667 by @itazap (#46667)

Full Changelog: https://github.com/huggingface/transformers/compare/v5.10.2...v5.10.3

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.12.0: Neues Modell MiniMax-M3-VL und weitere Ergänzungen

Transformers 5.12.0 fügt das Vision-Language-Modell MiniMax-M3-VL hinzu und aktualisiert die Dokumentation und Tests für PP-OCRv6.

Release v5.12.0

New Model additions

MiniMax-M3-VL

<img width="886" height="583" alt="image" src="https://github.com/user-attachments/assets/ae9dd96f-6877-4531-a06b-a756686f24e5" />

MiniMax-M3-VL is the vision-language member of the MiniMax-M3 family that pairs a CLIP-style vision tower with 3D rotary position embeddings with the MiniMax-M3 text backbone. It uses a mixed dense/sparse Mixture-of-Experts decoder with SwiGLU-OAI gated experts and a lightning indexer for block-sparse attention. The model processes images through a Conv3d patch embedding system and includes specialized components for efficient multimodal understanding and generation.

Links: Documentation

  • Add minimax m3vl (#46600) by @ArthurZucker in #46600

PP-OCRv6: update documentation and slow tests (#46576)

<img width="3840" height="1494" alt="image" src="https://github.com/user-attachments/assets/e62284ec-78bf-49cb-8aa2-deccc665372f" /> …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.11.0: Neue Modelle DiffusionGemma und DeepSeek-V3.2

Transformers 5.11.0 ergänzt DiffusionGemma, das Text blockweise per Diffusionssampler für schnellere Generierung erzeugt, sowie DeepSeek-V3.2.

Release v5.11.0

New Model additions

DiffusionGemma

<img width="1240" height="700" alt="image" src="https://github.com/user-attachments/assets/5081e449-6374-4076-bd96-d295c8334ca4" />

DiffusionGemma is engineered to reduce the sequential bottlenecks of standard causal language models by employing an encoder-decoder architecture specifically optimized for inference speed. During inference, DiffusionGemma leverages multi-canvas sampling, where rather than generating one token at a time, the model iteratively denoises a full block of tokens using a diffusion sampler. This block-autoregressive approach facilitates text generation at higher speeds compared to traditional sequential generation methods.

Links: Documentation

  • GPU go brr (#46540) by @gante in #46540

DeepSeek-V3.2

<img width="1135" height="671" alt="image" src="https://github.com/user-attachments/assets/24c9694d-eeae-402c-9a98-f7a3971dd9d0" /> …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.10.2: Korrektur der Modellkonvertierung für CLIP-Modelle

Das Patch-Release 5.10.2 behebt einen gravierenden Fehler bei der Konvertierung von CLIP-bezogenen Modellen wie sam3, weshalb ein Update empfohlen wird.

Patch release v5.10.2

There was a big bug in the model conversion of models related to clip, this affected models like sam3 and others. Please make sure to update :pray:

  • Fix conversion for clip models by @zucchini-nlp (#46406)

Full Changelog: https://github.com/huggingface/transformers/compare/v5.10.1...v5.10.2

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.10.1: Gemma 4 Unified und Gemma 4 MTP, 5.10.0 zurückgezogen

Version 5.10.1 ersetzt das zurückgezogene 5.10.0, das auf einem beschädigten Branch veröffentlicht wurde, und fügt Gemma 4 12B Unified, ein encoderfreies multimodales Modell, sowie Gemma 4 MTP hinzu.

Release v5.10.1

v5.10.0 was yanked as we publish on a corrupted branch. Sorry everyone, this happens when we rush a release!!!

New Model additions

Gemma4 unified+ Gemma4 MTP

<img width="2000" height="400" alt="image" src="https://github.com/user-attachments/assets/5e3ee940-f78d-4343-ac7a-889930800aa6" />

Gemma 4 12B Unified is an encoder-free multimodal model with pretrained and instruction-tuned variants. Unlike standard Gemma 4, which uses dedicated encoder towers, Gemma 4 12B Unified projects raw inputs directly into the language model's embedding space through lightweight linear pipelines. This results in a simpler architecture while maintaining strong multimodal performance.

Key differences from standard Gemma 4:

  • No Vision Tower: Raw pixel patches are projected directly into LM space via a Dense + LayerNorm pipeline with factorized 2D positional embeddings, replacing the vision encoder.
  • No Audio Tower: Raw 16 kHz waveform samples are chunked into fixed-length frames and projected through a simple RMSNorm → Linear pipeline, replacing the mel spectrogram + Conformer encoder.
  • Shared Multimodal Pipeline: Both vision and audio use the same Gemma4UnifiedMultimodalEmbedder (RMSNorm → Linear) for the final projection to text hidden space.

You can find the original Gemma 4 12B Unified checkpoints under the Gemma 4 release. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.9.0: Neue Modelle Cohere2Moe, Parakeet tdt und HRM-Text

Transformers 5.9.0 fügt neue Modelle hinzu, darunter das MoE-Modell Cohere2Moe (Command A+), Parakeet tdt und das autoregressive HRM-Text.

Release v5.9.0

New Model additions

Cohere2Moe

Command A+ is a Mixture-of-Experts (MoE) language model from Cohere that features a hybrid attention pattern combining sliding window and full attention layers. The model incorporates both shared and routed experts and supports a very large context window for processing extensive text sequences.

Links: Documentation

  • Add new cohere2_moe model (#46115) by @Cyrilvallez in #46115

Parakeet tdt (#44171)

  • Parakeet tdt (#44171) by @lmaksym

HRM-Text

HRM-Text is an improved autoregressive language-modeling variant of the Hierarchical Reasoning Model (HRM) that uses a hierarchical recurrent forward pass with two transformer stacks - one for slow, abstract planning (H) and one for fast, detailed computation (L) - reused inside a nested recurrence. It features PrefixLM attention where instruction tokens attend bidirectionally while response tokens attend causally, per-head sigmoid output gates, and parameterless RMSNorm. The model is designed as a base language model without instruction tuning or chat templates.

Links: Documentation | Paper …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.8.1: Fix der DeepSeek-V4-Integration

Das Patch-Release 5.8.1 korrigiert hauptsächlich die DeepSeek-V4-Integration und behebt zudem einen falschen WeightConverter-Regex-Treffer bei shared_experts sowie einen fatal_error-Fall im ContinuousBatchingManager.

Patch release v5.8.1

This release is mainly to fix the Deepseek V4 integration!!!

<img width="714" height="774" alt="image" src="https://github.com/user-attachments/assets/0d85e891-a0ff-436e-a9d4-b6633096f2b5" />
  • [fix] Add fatal_error to ContinuousBatchingManager so the serving... by @qgallouedec, @remi-or
  • Fix WeightConverter regex incorrectly matching shared_experts as experts by @silencelamb, @claude
  • Fix deepseek v4 by @ArthurZucker (#45892)
  • Deepseek v4 csa mask collapse by @ArthurZucker, @Sawyer117 (#45928)

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Transformers von Hugging Face

Transformers 5.8.0: DeepSeek-V4 und Gemma 4 Assistant

Transformers 5.8.0 fügt DeepSeek-V4 (Flash, Pro und Base-Varianten) sowie Gemma 4 Assistant als neue Modelle hinzu.

Release v5.8.0

New Model additions

DeepSeek-V4

<img width="6604" height="3574" alt="image" src="https://github.com/user-attachments/assets/4c0fdb29-f770-463c-a97b-d24438896a4c" />

DeepSeek-V4 is the next-generation MoE (Mixture of Experts) language model from DeepSeek that introduces several architectural innovations over DeepSeek-V3. The architecture replaces Multi-head Latent Attention (MLA) with a hybrid local + long-range attention design, swaps residual connections for Manifold-Constrained Hyper-Connections (mHC), and bootstraps the first few MoE layers with a static token-id → expert-id hash table. This implementation covers DeepSeek-V4-Flash, DeepSeek-V4-Pro, and their -Base pretrained variants, which share the same architecture but differ in width, depth, expert count and weights.

Links: Documentation | Paper

  • Add DeepSeek V4 (#45643) by @ArthurZucker in #45643

Gemma 4 Assistant

<img width="2000" height="400" alt="image" src="https://github.com/user-attachments/assets/02c79b0b-a172-4495-b09d-a6a4b625ee66" /> …

Originalquelle(öffnet in neuem Tab)Problem melden