Zum Inhalt springen

Hugging Face Release Notes

65 Einträge aus 3 Quellen. Zuletzt aktualisiert:

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Transformers von Hugging Face

Transformers v5.19.0: EmbeddingGemma2, Router-Logits als Breaking Change

Transformers v5.19.0 ergänzt das multimodale Embedding-Modell EmbeddingGemma2, das Text, Bilder, Audio und Video in einen gemeinsamen 768-dimensionalen Vektorraum abbildet, und gibt bei allen MoE-Modellen mit Router-Logits diese nun bei output_router_logits=True zurück, was ein Breaking Change ist.

Release v5.19.0

New Model additions

EmbeddingGemma2

<img width="2716" height="2308" alt="image" src="https://github.com/user-attachments/assets/84734e74-163d-4d12-b166-ffcf6749d563" />

EmbeddingGemma 2 is a multimodal embedding model from Google built on the Gemma 4 architecture. It encodes text, images, audio, and video, individually or combined in one input, into a shared 768-dimensional vector space for cross-modal retrieval, semantic similarity, clustering, and classification. It uses Matryoshka Representation Learning, so embeddings can be truncated to 512, 256, or 128 dimensions. It also offers configurable visual and video token budgets, and unused vision or audio towers can be disabled at load time to save memory.

Links: Documentation

  • Smthn smthn (#49364) by @vasqu in #49364

Breaking changes

All MoE models whose routers compute logits now return them when output_router_logits=True, following the Qwen3-MoE pattern (a router_logits recorder on the base model, MoeModelOutputWithPast from the backbone, and a MoE causal LM output from the head), so code that relied on the previous outputs or their absence should read the router logits from these output classes.

  • 🚨 Return router logits from every MoE model that computes them (#48920) by @qgallouedec …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Diffusers von Hugging Face

Diffusers 0.41.0: Qwen-Image 2.1 und neue LTX-2.5-DFR-Pipelines

Diffusers 0.41.0 integriert Qwen-Image 2.1 mit Text-zu-Bild-Generierung, Bildbearbeitung, nativer Transparenz (RGBA) und LoRA-Training und ergänzt unter anderem Tensor-Parallel-Checkpoint-Loading mit geringerem Speicherbedarf sowie neue LTX-2.5-DFR-Pipelines.

[!TIP] This release brings Qwen-Image 2.1 to Diffusers, with text-to-image generation, image editing, native transparency, and LoRA training.

Starting with this release, we’re adopting the Transformers release philosophy: coordinating minor releases around new model integrations, while including the other changes merged since the previous release. Patch releases remain focused on fixes. See our release policy.

Qwen-Image 2.1

Qwen-Image 2.1 unifies image generation and editing in one model. Its visual generation component has 7B parameters, and it supports multiple reference images and native RGBA output for transparent images. Check out the docs for more details.

Thanks to @naykun for the model integration (#14804).

Other highlights

  • Tensor-parallel checkpoint loading: each rank reads its own slice of sharded weights directly, reducing the memory needed during loading. (#14544)
  • LTX-2.5 DFR: new pipelines support generation with keyframe slots and composable spatial and temporal refinement. (#14567) …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Hugging Face

P(doom) im Hugging-Face-Profil angeben

Nutzer können in ihrem Hugging-Face-Profil ihren P(doom) angeben, also ihre geschätzte Wahrscheinlichkeit für eine existenzielle KI-Katastrophe, wobei die Antwort zu einer Community-Umfrage beiträgt, deren anonymisierte, aggregierte Ergebnisse in zwei Wochen veröffentlicht werden.

[Upvote

155](/login?next=%2Fchangelog%2Fpdoom-survey)

You can now add your P(doom) to your Hugging Face profile: your estimated probability of AI causing an existential catastrophe. Set it from your profile.

Your answer also contributes to a community survey on AI risk perception. Only aggregated, anonymized results will be published, in two weeks.

image image

Sep 25, 26

[Upvote

194](/login?next=%2Fchangelog) …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Transformers von Hugging Face

Transformers 5.18.0: Nemotron 3 Diarization und NemotronH Omni

Transformers 5.18.0 ergänzt unter anderem die neuen Modelle Nemotron 3 Diarization für Streaming-Sprecherdiarisierung mit bis zu acht Sprechern sowie NemotronH Omni von NVIDIA.

New Model additions

Nemotron 3 Diarization

<img width="1680" height="900" alt="image" src="https://github.com/user-attachments/assets/fe735cb3-9e5b-43ad-8f60-9dec8425aec7" />

Nemotron 3 Diarization is an open-weight streaming speaker diarization model designed to determine "who spoke when" in real-world audio. It supports both streaming and offline inference, handles up to eight speakers, and orders speaker outputs by each speaker's first arrival in the input audio.

The model uses the Arrival-Order Speaker Cache (AOSC) 1 and FIFO queue introduced for Streaming Sortformer 1, 2. A single checkpoint supports configurable latency profiles, from an 80 ms input buffer to a 30.4 s offline-style buffer, and configurable output frame resolution in multiples of 10 ms. With chunked inference, the maximum audio duration is not limited.

Links: Documentation

  • Add Nemotron3Diarization (#49056) by @eustlb in #49056

NemotronH Omni

NemotronH Omni is a multimodal reasoning model from NVIDIA that pairs the NemotronH hybrid …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Hugging Face

Live-Ressourcenauslastung bei Jobs

Die Job-Seite zeigt neben den Logs ein neues Resources-Panel, das während der Laufzeit CPU-, RAM-, GPU- und VRAM-Auslastung live und im Abstand weniger Sekunden aktualisiert anzeigt.

[Upvote

194](/login?next=%2Fchangelog%2Flive-resource-usage-on-jobs)

The job page has a new Resources panel next to the logs. While your job runs, it shows live CPU, RAM, GPU and VRAM usage, updated every couple of seconds.

Screenshot 2026-09-25 at 14.50.06 Screenshot 2026-09-25 at 14.50.04

Sep 24, 26

[Upvote

65](/login?next=%2Fchangelog)

  • …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Hugging Face

Hugging Face jetzt im Google Cloud Marketplace

Organisationen können Hugging Face jetzt über den Google Cloud Marketplace abonnieren, sodass Inference Providers, Inference Endpoints, Spaces, Jobs, Speicher und Team- oder Enterprise-Plätze über die bestehende Google-Cloud-Rechnung abgerechnet werden.

[Upvote

37](/login?next=%2Fchangelog%2Fgcp-marketplace)

Organizations can now subscribe with their Google Cloud account. Inference Providers, Inference Endpoints, Spaces, Jobs, storage and Team or Enterprise seats all land on the Google Cloud invoice they already pay, with no new vendor to onboard and no credit card. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Hugging Face

Vorschau für LeRobot-Episoden

LeRobot-Datasets öffnen sich jetzt mit einer Episodenvorschau, in der alle Kameras über eine gemeinsame Zeitleiste synchron abgespielt werden, und Dataset-Listen zeigen Robotertyp, Episodenanzahl und eine Vorschau der ersten Episode.

[Upvote

65](/login?next=%2Fchangelog%2Fpreview-lerobot-episodes)

LeRobot datasets now open with an episode preview on the dataset page. Every camera plays on one shared transport, so you can scrub through an episode and see each angle at the same moment, and the episode list beside it previews each take on hover.

Dataset listings show it too: a robotics dataset now carries its robot type, its episode count and a preview of its first episode, so you can tell what's inside before opening it.

Browse them all at https://huggingface.co/datasets?other=LeRobot …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Transformers von Hugging Face

Transformers 5.17.0: neues Modell HYV4 (Hy4-Preview)

Transformers 5.17.0 ergänzt das neue Modell HYV4 (Hy4-Preview), ein Mixture-of-Experts-Sprachmodell mit 780B Parametern, 49B aktiven Parametern pro Token und 1M Tokens Kontextfenster.

Release v5.17.0

New Model additions

HYV4

<img width="1503" height="827" alt="image" src="https://github.com/user-attachments/assets/e6ed85ee-eb1d-40eb-a0d4-c649f6337ca9" />

Hy4-Preview is a 780B-parameter mixture-of-experts language model that activates 49B parameters per token. Each MoE layer holds 256 routed experts plus one always-active shared expert and routes every token to 8 of them. The context window is 1M tokens.

The architecture combines four features:

  • Multi-head Latent Attention (MLA) compresses keys and values into a low-rank latent (kv_lora_rank) that kv_b_proj expands back to one key/value per query head.
  • DeepSeek Sparse Attention (DSA) selects index_topk keys per query with a lightweight indexer. Following IndexShare, only the layers marked "full" in indexer_types run an indexer; "shared" layers reuse the previous full layer's selection.
  • Gated MLA with learnable attention sinks, where each head owns a sink logit that participates in the softmax and contributes no value, as in GPT-OSS.
  • Independent Hyper-Connections (iHC) replace the plain residual path with hc_mult parallel residual streams that are collapsed before, and redistributed after, every sublayer.

The implementation does not execute the multi-token prediction (MTP) layers. Released checkpoints …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Transformers von Hugging Face

Transformers 5.16.0: neues Modell Qwen4-Exp

Transformers v5.16.0 ergänzt das neue Modell Qwen4-Exp, das auf der Qwen3.5-Architektur aufbaut und lineare mit sparsamer Attention kombiniert.

Release v5.16.0

New Model additions

Qwen4-Exp

<img width="2241" height="693" alt="image" src="https://github.com/user-attachments/assets/c838b5ba-ffea-42da-baa9-3f66178e3671" />

Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding (PLE).

GR is a Qwen-developed residual architecture that combines Hyper-Connection with GatedNorm. It mixes multiple residual streams with fine-grained elementwise gating before each attention and Mixture-of-Experts (MoE) block, then controls how much of the block output is injected back into each stream.

QSA uses multiple query heads to score compressed key blocks, selects the most relevant contiguous token blocks, and keeps the incomplete trailing block uncompressed. This block-level selection reduces indexing overhead and improves memory locality for long sequences. Combined with Gated DeltaNet, QSA makes Qwen4-Exp the first hybrid architecture to integrate linear and sparse attention, substantially improving inference efficiency for long-context workloads.

PLE enriches selected decoder layers with layer-specific lexical features derived from hashed token n-grams and a dilated depthwise convolution.

Links: Documentation …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Transformers von Hugging Face

Transformers 5.16.1: GLM-5.3-Flash und kleine Fixes

Transformers v5.16.1 unterstützt nun GLM-5.3-Flash, ein nativ multimodales Modell der GLM-5-Reihe, und enthält kleine Fixes zur Rückwärtskompatibilität bei TP sowie ein angepinntes hf-Kernel aus Sicherheitsgründen.

Release v5.16.1

This is a special release as we include GLM! (and a few small fixes)

GLM-5.3-Flash

<img width="4239" height="2643" alt="image" src="https://github.com/user-attachments/assets/17bc9c29-758b-44c8-8230-42f945ded209" />

GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.

GLM-5.3-Flash starts from a newly trained base model, with its architecture and training recipe redesigned around capability and efficiency. For the first time in the GLM series, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. The model also adopts Manifold-Constrained Hyper-Connections (mHC) to further improve scaling efficiency. Together with our latest 30T-token multimodal pre-training corpus, these changes enable GLM-5.3-Flash to deliver more intelligence with less compute.

Links: Documentation

  • [Glm 5.3 Flash] GLM 5.3 Flash Support (#48342) by @Dovis01 in #48342

Small patch fixes

Mainly BC behavior for TP and pinning a hf kernel for security reasons :hugs: …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Diffusers von Hugging Face

Diffusers 0.40.0: Neue Pipelines und Tensor-Parallelismus

Diffusers 0.40.0 bringt neue Pipelines wie LTX2.5, MiniMax H3 und Wan Animate 2, erklärt Modular Diffusers für stabil und bietet minimale Unterstützung für Tensor-Parallelismus.

[!TIP] This release features several new pipelines, including LTX2.5, MiniMax H3, and Wan Animate 2. We're also graduating Modular Diffusers out of the experimental phase and announcing its stable support. Additionally, this release includes minimal support for tensor-parallel. There's a lot more that went down in this release. So, please consult the notes for details.

New Pipelines

MiniMax-H3

MiniMax-H3 generates video and its soundtrack together. A single transformer denoises one packed sequence containing the text conditioning, the conditioning media, and the target video and audio latents — there is no separate vocoder and no post-hoc audio pass. Its conditioner is a Qwen3VLForConditionalGeneration whose unnormalized 50th-decoder-layer hidden state is read instead of the last one.

MiniMax-H3 is integrated as Modular Diffusers blocks only — MiniMaxH3Blocks and their MiniMaxH3ModularPipeline are the whole integration. The conversion ships both checkpoint partitions in one repository and exposes three workflows (t2va, fl2va, ref2va) that can be pruned at from_pretrained time so only that task's components are declared and downloaded.

MiniMax Music 3 …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Transformers von Hugging Face

Transformers 5.15.1: Fixes für DFlash, MTP, Gemma4 und Lanczos

Transformers 5.15.1 behebt Probleme mit DFlash- und MTP-Candidate-Generatoren, einen Gerätekonflikt bei Gemma4-Videos sowie Fehler bei der Bildverarbeitung auf Beschleunigern mit Lanczos-Filter, indem auf bicubic ausgewichen wird.

Patch release v5.15.1

This patch most notably solves a few issues with DFlash and MTP candidate generators, as well as an issue where images could sometimes not be processed on accelerator if using Lanczos filter.

It contains the following commits:

  • Fix DFlash candidate token device mismatch with device_map="auto" (#47877) by @sywangyi and @Cyrilvallez
  • Align logit distributions for CandidateGenerators using sampling (#48007) by @Cyrilvallez
  • Fix MTP config when mlp_layer_types is absent (#48015) by @Cyrilvallez
  • Fallback from 'lanczos' to 'bicubic' when on cuda (#48026) by @zucchini-nlp
  • Fix gemma4 video to device (#47896) by @guarin

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Hugging Face

Feingranulare Funktionszugriffe pro Resource Group

Der Zugriff auf Funktionen lässt sich nun pro Resource Group statt nur über die Organisationsrolle steuern, etwa Jobs für alle offen, Inference Endpoints nur für Admins und Blog-Veröffentlichung nur für die Resource Group blog-writer.

[Upvote

417](/login?next=%2Fchangelog%2Fgranular-feature-access)

You can now control feature access per resource group rather than across the whole organization. Before, the only way to limit a feature was by organization role.

This means you can leave Jobs open to everyone, restrict Inference Endpoints to admins, and give blog publishing rights only to the blog-writer resource group users.

Feature access per resource group …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Transformers von Hugging Face

Transformers 5.15.0: Muse Glimmer, GraniteMoeSWA und weitere Modelle

Transformers v5.15.0 ergänzt neue Modelle, darunter Metas multimodales Muse Glimmer mit 30B Parametern unter Apache-2.0-Lizenz sowie GraniteMoeSWA, GraniteSWA und A.X-K1/A.X-K2.

Release v5.15.0

New Model additions

Meta Muse Glimmer

Muse Glimmer, released today, is Meta’s new multimodal model, especially designed for agentic use cases. Distilled from Muse to 30B parameters, and released under the Apache 2.0 license, it can be deployed to local setups for privacy-aware applications such as coding, document analysis, personal assistants, Claw- or Hermes-like setups.

Muse Glimmer is a dense 30B parameter model consisting of:

  • 2B ViT-style encoder for vision (Perception Encoder)
  • 28B parameter text decoder

We're covering it in the following blogpost: http://hf.co/blog/muse-glimmer

<img width="960" height="1787" alt="image" src="https://github.com/user-attachments/assets/3d8e548e-f84f-4269-8bd0-a12722d7ab01" />

GraniteMoeSWA & GraniteSWA

<img width="1013" height="389" alt="image" src="https://github.com/user-attachments/assets/2c2b87f0-466a-413a-a4be-25ceae49c9a5" />

Links: Documentation

  • Add Granite-swa and Granitemoe-swa model support (#47179) by @daviswer in #47179

Links: Documentation

  • Add Granite-swa and Granitemoe-swa model support (#47179) by @daviswer in #47179

A.X-K1 & A.X-K2 …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Transformers von Hugging Face

Transformers 5.14.1: Fixes für Inkling, FP8-Kernels und deepgemm

Transformers 5.14.1 behebt Probleme bei der Inkling-Integration, darunter Fehler bei assistierter Generierung mit EncoderDecoderCache und beim sdpa-Prefill mit position_bias, und enthält zudem Fixes für FP8-Kernels und deepgemm auf mehreren Geräten.

Patch release v5.14.1

This patch solves a few issues which appeared when integrating Inkling model, most notably an issue affecting models using EncoderDecoderCache during assisted generation. It also fixes an issue that could appear during prefill with StaticCache and sdpa without padding for Inkling which uses a position_bias. It contains the following commits:

  • Fix sdpa prefill with position_bias (#47359) by @Cyrilvallez
  • Fix assisted decoding for models with EncoderDecoder cache & OlmoHybrid (#47361) by @Cyrilvallez
  • [FP8] Bump kernels version (#47344) by @vasqu
  • Fix deepgemm on multiple devices (#47323) by @IlyasMoutawwakil

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Transformers von Hugging Face

Transformers 5.14.0: Inkling, TIPSv2 und TIPSv2 DPT

Transformers v5.14.0 ergänzt neue Modelle, darunter das multimodale Inkling von Thinking Machines mit 975B Parametern (41B aktiv) sowie TIPSv2 und TIPSv2 DPT.

Release v5.14.0

New Model additions

Inkling (fresh from Thinking Machines): 975B total, 41B active

  • Add Inkling model #47347 by @molbap @Cyrilvallez @eustlb and @zucchini-nlp
<img width="3840" height="2160" alt="image" src="https://github.com/user-attachments/assets/051f819a-512f-4987-9bee-6e2fa2af3db7" />

Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI- powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-following, and other natural language and multimodal tasks. It is released with open weights to support research, fine-tuning and integration into third-party products by downstream developers.

TIPSv2

<img width="1555" height="1306" alt="image" src="https://github.com/user-attachments/assets/2d9f21e5-05f8-4c36-93ef-22f03c089f52" />

Links: Documentation

  • Add TIPSv2 (#46347) by @Ternura143 in #46347

TIPSv2 DPT …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Transformers von Hugging Face

Transformers 5.13.1: Kompatibilität mit der neuesten vllm-Version

Transformers 5.13.1 sorgt für Kompatibilität mit der neuesten vllm-Version, indem Handhabung von Legacy-Layer-Typen, Custom Code und str-Schlüsseln in _LazyAutoMapping.register korrigiert wird.

Patch release v5.13.1

This patch is focused on enabling transformers for the latest release of vllm!

  • Be more defensive with remap_legacy_layer_types for custom models (#47245) from @hmellor
  • Fix custom code which doesn't know about the new linear layer type names (#47174) from @hmellor
  • Fix case where _LazyAutoMapping.register is passed a str key (#47148) from @hmellor

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Transformers von Hugging Face

Transformers 5.13.0: Neue Modelle, darunter Kimi K2.5 bis K2.7

Transformers 5.13.0 fügt unter anderem die Architektur für Kimi K2.5 hinzu, die auch von K2.6 und K2.7 genutzt wird, sowie weitere neue Modelle wie MiMo-V2-Flash.

Release v5.13.0

New Model additions

KimiK 2.5, 2.6, and 2.7

<img width="1097" height="400" alt="image" src="https://github.com/user-attachments/assets/c24d2232-a9b4-413b-a2c8-58d013b6dfbd" />

This release includes the architecture for Kimi 2.5 which is used by 2.5-2.7:

Kimi K2.5 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration. The model was proposed in Kimi K2.5: Visual Agentic Intelligence and further improved in [Kimi K2.6: Advancing Open-Source Coding](Kimi K2.5: Visual Agentic Intelligence).

Kimi K2.5 achieves significant improvements on complex, end-to-end coding tasks, generalizing robustly across programming languages (Rust, Go, Python) and domains spanning front-end, DevOps, and performance optimization. The model is capable of transforming simple prompts and visual inputs into production-ready interfaces and lightweight full-stack workflows, generating structured layouts, interactive elements, and rich animations with deliberate aesthetic precision.

Links: Documentation

  • Add new model: Kimi2-6 (#45630) by @zucchini-nlp in #45630

MiMo-V2-Flash …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Diffusers von Hugging Face

Diffusers 0.39.0 mit neuen Bild- und Video-Pipelines

Diffusers 0.39.0 führt neue Pipelines ein, darunter NVIDIAs Cosmos 3 als einheitliches World Foundation Model und das Text-zu-Bild-Modell Ideogram 4 mit LoRA-Unterstützung, außerdem Verbesserungen an der Kernbibliothek.

New Pipelines

Cosmos 3

Cosmos 3 is NVIDIA's unified world foundation model (WFM) for Physical AI — a single omni-model built on a Mixture-of-Transformers (MoT) architecture that combines world generation, physical reasoning, and action generation, replacing the separate Predict, Reason, and Transfer models from earlier Cosmos releases. A single Cosmos3OmniTransformer runs a Qwen-style language model in parallel with a diffusion generation pathway, joined by a 3D multimodal RoPE. This release also lands video-to-video and action-conditioned generation, and a sound encoder.

Thanks to @atharvajoshi10, @yzhautouskay, and @MaciejBalaNV for the contributions.

Ideogram 4

Ideogram 4 is a flow-matching text-to-image model that uses a multimodal text encoder and an asymmetric classifier-free guidance scheme: a dedicated unconditional_transformer produces the negative branch with zeroed text features, while the main transformer consumes the full packed text + image sequence. The pipeline ships with structured prompt upsampling and LoRA loading support. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Transformers von Hugging Face

Transformers 5.10.4 (Version 5.10.3): Fixes für vLLM-Synchronisation

Das Patch-Release enthält mehrere Korrekturen für die Synchronisation mit vLLM, darunter Fixes für InternVL-Modelle, Token-IDs im ProcessorMixin, Offsets in der Verarbeitung, die PEFT-Untergrenze und den Mistral-Backend; auf PyPI erscheint es als 5.10.4, da 5.10.3 nicht existiert.

Patch release v5.10.4

Update: Note that on pypi 5.10.3 doesn't exist and this this saved under 5.10.4 (so essentially a minor version skipped). Sorry about that, that's on me. Just wanted to clarify to make this less confusing!

A few fixes needed for vLLM to sync with transformers :hugs:

  • [fix] regression introduced by #45534 #46456 by @eustlb (#46456)
  • Fix {image/video/audio}_token_ids in ProcessorMixin #46500 by @hmellor (#46500)
  • Fix InternVL models #46524 by @hmellor (#46524)
  • Fix the offsets in processing #46525 by @zucchini-nlp (#46525)
  • Fix peft lower bound #46605 by @hmellor (#46605)
  • mistral common backend fix #46667 by @itazap (#46667)

Full Changelog: https://github.com/huggingface/transformers/compare/v5.10.2...v5.10.3

Originalquelle(öffnet in neuem Tab)Problem melden