Angaben zum Datum
Datum aus der Quelle.
Aufgenommen am .
Transformers v5.19.0: EmbeddingGemma2, Router-Logits als Breaking Change
Transformers v5.19.0 ergänzt das multimodale Embedding-Modell EmbeddingGemma2, das Text, Bilder, Audio und Video in einen gemeinsamen 768-dimensionalen Vektorraum abbildet, und gibt bei allen MoE-Modellen mit Router-Logits diese nun bei output_router_logits=True zurück, was ein Breaking Change ist.
Release v5.19.0
New Model additions
EmbeddingGemma2
<img width="2716" height="2308" alt="image" src="https://github.com/user-attachments/assets/84734e74-163d-4d12-b166-ffcf6749d563" />EmbeddingGemma 2 is a multimodal embedding model from Google built on the Gemma 4 architecture. It encodes text, images, audio, and video, individually or combined in one input, into a shared 768-dimensional vector space for cross-modal retrieval, semantic similarity, clustering, and classification. It uses Matryoshka Representation Learning, so embeddings can be truncated to 512, 256, or 128 dimensions. It also offers configurable visual and video token budgets, and unused vision or audio towers can be disabled at load time to save memory.
Links: Documentation
- Smthn smthn (#49364) by @vasqu in #49364
Breaking changes
All MoE models whose routers compute logits now return them when output_router_logits=True, following the Qwen3-MoE pattern (a router_logits recorder on the base model, MoeModelOutputWithPast from the backbone, and a MoE causal LM output from the head), so code that relied on the previous outputs or their absence should read the router logits from these output classes.
- 🚨 Return router logits from every MoE model that computes them (#48920) by @qgallouedec …