Zum Inhalt springen

Diffusers Updates & Release Notes

30 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.41.0: Qwen-Image 2.1 und neue LTX-2.5-DFR-Pipelines

Diffusers 0.41.0 integriert Qwen-Image 2.1 mit Text-zu-Bild-Generierung, Bildbearbeitung, nativer Transparenz (RGBA) und LoRA-Training und ergänzt unter anderem Tensor-Parallel-Checkpoint-Loading mit geringerem Speicherbedarf sowie neue LTX-2.5-DFR-Pipelines.

[!TIP] This release brings Qwen-Image 2.1 to Diffusers, with text-to-image generation, image editing, native transparency, and LoRA training.

Starting with this release, we’re adopting the Transformers release philosophy: coordinating minor releases around new model integrations, while including the other changes merged since the previous release. Patch releases remain focused on fixes. See our release policy.

Qwen-Image 2.1

Qwen-Image 2.1 unifies image generation and editing in one model. Its visual generation component has 7B parameters, and it supports multiple reference images and native RGBA output for transparent images. Check out the docs for more details.

Thanks to @naykun for the model integration (#14804).

Other highlights

  • Tensor-parallel checkpoint loading: each rank reads its own slice of sharded weights directly, reducing the memory needed during loading. (#14544)
  • LTX-2.5 DFR: new pipelines support generation with keyframe slots and composable spatial and temporal refinement. (#14567) …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.40.0: Neue Pipelines und Tensor-Parallelismus

Diffusers 0.40.0 bringt neue Pipelines wie LTX2.5, MiniMax H3 und Wan Animate 2, erklärt Modular Diffusers für stabil und bietet minimale Unterstützung für Tensor-Parallelismus.

[!TIP] This release features several new pipelines, including LTX2.5, MiniMax H3, and Wan Animate 2. We're also graduating Modular Diffusers out of the experimental phase and announcing its stable support. Additionally, this release includes minimal support for tensor-parallel. There's a lot more that went down in this release. So, please consult the notes for details.

New Pipelines

MiniMax-H3

MiniMax-H3 generates video and its soundtrack together. A single transformer denoises one packed sequence containing the text conditioning, the conditioning media, and the target video and audio latents — there is no separate vocoder and no post-hoc audio pass. Its conditioner is a Qwen3VLForConditionalGeneration whose unnormalized 50th-decoder-layer hidden state is read instead of the last one.

MiniMax-H3 is integrated as Modular Diffusers blocks only — MiniMaxH3Blocks and their MiniMaxH3ModularPipeline are the whole integration. The conversion ships both checkpoint partitions in one repository and exposes three workflows (t2va, fl2va, ref2va) that can be pruned at from_pretrained time so only that task's components are declared and downloaded.

MiniMax Music 3 …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.39.0 mit neuen Bild- und Video-Pipelines

Diffusers 0.39.0 führt neue Pipelines ein, darunter NVIDIAs Cosmos 3 als einheitliches World Foundation Model und das Text-zu-Bild-Modell Ideogram 4 mit LoRA-Unterstützung, außerdem Verbesserungen an der Kernbibliothek.

New Pipelines

Cosmos 3

Cosmos 3 is NVIDIA's unified world foundation model (WFM) for Physical AI — a single omni-model built on a Mixture-of-Transformers (MoT) architecture that combines world generation, physical reasoning, and action generation, replacing the separate Predict, Reason, and Transfer models from earlier Cosmos releases. A single Cosmos3OmniTransformer runs a Qwen-style language model in parallel with a diffusion generation pathway, joined by a 3D multimodal RoPE. This release also lands video-to-video and action-conditioned generation, and a sound encoder.

Thanks to @atharvajoshi10, @yzhautouskay, and @MaciejBalaNV for the contributions.

Ideogram 4

Ideogram 4 is a flow-matching text-to-image model that uses a multimodal text encoder and an asymmetric classifier-free guidance scheme: a dedicated unconditional_transformer produces the negative branch with zeroed text features, while the main transformer consumes the full packed text + image sequence. The pipeline ships with structured prompt upsampling and LoRA loading support. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.38.0 mit neuen Bild- und Audio-Pipelines

Diffusers 0.38.0 ergänzt neue Pipelines wie LLaDA2, Nucleus-MoE, Ernie-Image und LongCat-AudioDiT sowie Verbesserungen an der Kernbibliothek.

New Pipelines

LLaDA2

LLaDA2 is a family of discrete diffusion language models that generate text through block-wise iterative refinement. Instead of autoregressive token-by-token generation, LLaDA2 starts with a fully masked sequence and progressively unmasks tokens by confidence over multiple refinement steps.

Nucleus-MoE

NucleusMoE-Image is a 2B active 17B parameter model trained with efficiency at its core. Our novel architecture highlights the scalability of a sparse MoE architecture for Image generation.

Thanks to @sippycoder for the contribution.

Ernie-Image

ERNIE-Image is a powerful and highly efficient image generation model with 8B parameters.

Thanks to @HsiaWinter for the contribution.

LongCat-AudioDiT …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers Version 0.37.1

Das Update behebt das Laden von ModularPipelines mit AutoModel-Typhinweisen, das Laden von Flux-Klein-LoRA und einen ungeschützten torchvision-Import in Cosmos Predict 2.5.

  • Fix for loading ModularPipelines with AutoModel type hints in their modular_model_index.json #13271
  • Fix Flux Klein LoRA loading #13313
  • Fix unguarded torchvision import in Cosmos Predict 2.5 #13321

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.37.0 mit Modular Diffusers und neuen Pipelines

Diffusers 0.37.0 führt Modular Diffusers zum Zusammenstellen von Pipelines aus wiederverwendbaren Blöcken ein und bringt neue Bild- und Video-Pipelines, darunter Z Image Omni Base, sowie Verbesserungen an der Kernbibliothek.

Modular Diffusers

Modular Diffusers introduces a new way to build diffusion pipelines by composing reusable blocks. Instead of writing entire pipelines from scratch, you can now mix and match building blocks to create custom workflows tailored to your specific needs! This complements the existing DiffusionPipeline class, providing a more flexible way to create custom diffusion pipelines.

Find more details on how to get started with Modular Diffusers here, and also check out the announcement post.

New Pipelines and Models

Image 🌆

  • Z Image Omni Base: Z-Image is the foundation model of the Z-Image family, engineered for good quality, robust generative diversity, broad stylistic coverage, and precise prompt adherence. While Z-Image-Turbo is built for speed, Z-Image is a full-capacity, undistilled transformer designed to be the backbone for creators, researchers, and developers who require the highest level of creative freedom. Thanks to @RuoyiDufor for contributing this in #12857. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.36.0: Neue Pipelines, Caching-Methode und Trainingsskript

Diffusers 0.36.0 bringt neue Bild- und Videopipelines wie Flux2, Z-Image und QwenImage Edit Plus, eine neue Caching-Methode, ein neues Trainingsskript und neue auf kernels basierende Attention-Backends.

The release features a number of new image and video pipelines, a new caching method, a new training script, new kernels - powered attention backends, and more. It is quite packed with a lot of new stuff, so make sure you read the release notes fully 🚀

New image pipelines

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.35.2: Fehlerbehebungen für transformers-Modelle und Imports

Version 0.35.2 behebt Fehler beim Laden von transformers-Modellen mit offload_state_dict, ein Importproblem mit TRANSFORMERS_FLAX_WEIGHTS_NAME, die Kompatibilität mit PyTorch 2.3.1 sowie ein Problem mit scale_shift_factor auf der CPU bei Wan und LTX.

All commits

  • Release: v0.35.1-patch by @sayakpaul (direct commit on v0.35.2-patch)
  • handle offload_state_dict when initing transformers models by @sayakpaul in #12438
  • [CI] Fix TRANSFORMERS_FLAX_WEIGHTS_NAME import issue by @DN6 in #12354
  • Fix PyTorch 2.3.1 compatibility: add version guard for torch.library.… by @Aishwarya0811 in #12206
  • fix scale_shift_factor being on cpu for wan and ltx by @vladmandic in #12347
  • Release: v0.35.2-patch by @sayakpaul (direct commit on v0.35.2-patch)

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.35.0: Qwen-Image, Flux-Kontext und Wan 2.2

Diffusers 0.35.0 führt die neuen Pipelines Wan 2.2, Flux-Kontext, Qwen-Image und Qwen-Image-Edit sowie neue Trainingsskripte und Verbesserungen im Bedienkomfort ein.

This release comes packed with new image generation and editing pipelines, a new video pipeline, new training scripts, quality-of-life improvements, and much more. Read the rest of the release notes fully to not miss out on the fun stuff.

New pipelines 🧨

We welcomed new pipelines in this release:

  • Wan 2.2
  • Flux-Kontext
  • Qwen-Image
  • Qwen-Image-Edit

Wan 2.2 📹

This update to Wan provides significant improvements in video fidelity, prompt adherence, and style. Please check out the official doc to learn more.

Flux-Kontext 🎇

Flux-Kontext is a 12-billion-parameter rectified flow transformer capable of editing images based on text instructions. Please check out the official doc to learn more about it.

Qwen-Image 🌅

After a successful run of delivering language models and vision-language models, the Qwen team is back with an image generation model, which is Apache-2.0 licensed! It achieves significant advances in complex text rendering and precise image editing. To learn more about this powerful model, refer to our docs. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.34.0: Neue Bild- und Videomodelle, besseres torch.compile

Diffusers 0.34.0 ergänzt neue Videopipelines wie Wan VACE für kontrollierbare Videogenerierung und Cosmos Predict2 Video2World und verbessert die Unterstützung von torch.compile.

📹 New video generation pipelines

Wan VACE

Wan VACE supports various generation techniques which achieve controllable video generation. It comes in two variants: a 1.3B model for fast iteration & prototyping, and a 14B for high quality generation. Some of the capabilities include:

  • Control to Video (Depth, Pose, Sketch, Flow, Grayscale, Scribble, Layout, Boundary Box, etc.). Recommended library for preprocessing videos to obtain control videos: huggingface/controlnet_aux
  • Image/Video to Video (first frame, last frame, starting clip, ending clip, random clips)
  • Inpainting and Outpainting
  • Subject to Video (faces, object, characters, etc.)
  • Composition to Video (reference anything, animate anything, swap anything, expand anything, move anything, etc.)

The code snippets available in this pull request demonstrate some examples of how videos can be generated with controllability signals.

Check out the docs to learn more.

Cosmos Predict2 Video2World …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.33.0: Neue Modelle, Speicheroptimierungen, Caching und Remote VAEs

Diffusers 0.33.0 bringt neue Videopipelines wie Wan 2.1, LTX Video 0.9.5 mit LTXConditionPipeline und Hunyuan Image to Video sowie Speicheroptimierungen, Caching-Methoden, Remote VAEs und neue Trainingsskripte.

New Pipelines for Video Generation

Wan 2.1

Wan2.1 is a comprehensive and open suite of video foundation models that pushes the boundaries of video generation. The model release includes 4 different model variants and three different pipelines for Text to Video, Image to Video and Video to Video.

  • Wan-AI/Wan2.1-T2V-1.3B-Diffusers
  • Wan-AI/Wan2.1-T2V-14B-Diffusers
  • Wan-AI/Wan2.1-I2V-14B-480P-Diffusers
  • Wan-AI/Wan2.1-I2V-14B-720P-Diffusers

Check out the docs here to learn more.

LTX Video 0.9.5

LTX Video 0.9.5 is the updated version of the super-fast LTX Video model series. The latest model introduces additional conditioning options, such as keyframe-based animation and video extension (both forward and backward).

To support these additional conditioning inputs, we’ve introduced the LTXConditionPipeline and LTXVideoCondition object.

To learn more about the usage, check out the docs here.

Hunyuan Image to Video …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.32.2: Fixes für Flux, LoRA-Laden und HunyuanVideo

Version 0.32.2 behebt Regressionen beim Laden von Comfy-UI-Single-File-Checkpoints für Flux und von LoRAs mit bitsandbytes-4bit-Flux-Modellen, ergänzt unload_lora_weights für Flux Control und behebt Fehler bei HunyuanVideo mit Batchgröße > 1, das nun auch LoRAs aus dem Original-Code laden kann.

Fixes for Flux Single File loading, LoRA loading for 4bit BnB Flux, Hunyuan Video

This patch release

  • Fixes a regression in loading Comfy UI format single file checkpoints for Flux
  • Fixes a regression in loading LoRAs with bitsandbytes 4bit quantized Flux models
  • Adds unload_lora_weights for Flux Control
  • Fixes a bug that prevents Hunyuan Video from running with batch size > 1
  • Allow Hunyuan Video to load LoRAs created from the original repository code

All commits

  • [Single File] Fix loading Flux Dev finetunes with Comfy Prefix by @DN6 in #10545
  • [CI] Update HF Token on Fast GPU Model Tests by @DN6 #10570
  • [CI] Update HF Token in Fast GPU Tests by @DN6 #10568
  • Fix batch > 1 in HunyuanVideo by @hlky in #10548
  • Fix HunyuanVideo produces NaN on PyTorch<2.5 by @hlky in #10482
  • Fix hunyuan video attention mask dim by @a-r-r-o-w in #10454
  • [LoRA] Support original format loras for HunyuanVideo by @a-r-r-o-w in #10376
  • [LoRA] feat: support loading loras into 4bit quantized Flux models. by @sayakpaul in #10578
  • [LoRA] clean up load_lora_into_text_encoder() and fuse_lora() copied from by @sayakpaul in #10495
  • [LoRA] feat: support unload_lora_weights() for Flux Control. by @sayakpaul in #10206
  • Fix Flux multiple Lora loading bug by @maxs-kan in #10388
  • [LoRA] fix: lora unloading when using expanded Flux LoRAs. by @sayakpaul in #10397

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.32.1: Fixes für den TorchAO Quantizer

Version 0.32.1 behebt Fehler des TorchAO Quantizers: Der Import unter PyTorch vor 2.3.0 funktioniert wieder, die Quantisierung wird korrekt ausgeführt und die Verwendung von Device Map löst nun einen Fehler aus.

TorchAO Quantizer fixes

This patch release fixes a few bugs related to the TorchAO Quantizer introduced in v0.32.0.

  • Importing Diffusers would raise an error in PyTorch versions lower than 2.3.0. This should no longer be a problem.
  • Device Map does not work as expected when using the quantizer. We now raise an error if it is used. Support for using device maps with different quantization backends will be added in the near future.
  • Quantization was not performed due to faulty logic. This is now fixed and better tested.

Refer to our documentation to learn more about how to use different quantization backends.

All commits

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.32.0: Neue Video- und Bildpipelines, Quantisierungs-Backends

Diffusers 0.32.0 bringt neue Videopipelines (Mochi-1, Allegro, LTXVideo, HunyuanVideo), neue Bildpipelines wie SANA, neue Quantisierungs-Backends und neue Trainingsskripte.

https://github.com/user-attachments/assets/34d5f7ca-8e33-4401-8109-5c245ce7595f

This release took a while, but it has many exciting updates. It contains several new pipelines for image and video generation, new quantization backends, and more.

Going forward, to provide more transparency to the community about ongoing developments and releases in Diffusers, we will be making use of a roadmap tracker.

New Video Generation Pipelines 📹

Open video generation models are on the rise, and we’re pleased to provide comprehensive integration support for all of them. The following video pipelines are bundled in this release:

Check out this section to learn more about the fine-tuning options available for these new video models.

New Image Generation Pipelines

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.31.0: Stable Diffusion 3.5 Large, CogView3 und Quantisierung

Diffusers 0.31.0 unterstützt Stable Diffusion 3.5 Large und das neue Text-zu-Bild-Modell CogView3-plus und bringt zudem Quantisierung und neue Trainingsskripte.

v0.31.0: Stable Diffusion 3.5 Large, CogView3, Quantization, Training Scripts, and more

Stable Diffusion 3.5 Large

Stability AI’s latest text-to-image generation model is Stable Diffusion 3.5 Large. SD3.5 Large is the next iteration of Stable Diffusion 3. It comes with two checkpoints (both of which have 8B params):

  • A regular one
  • A timestep-distilled one enabling few-step inference

Make sure to fill up the form by going to the model page, and then run huggingface-cli login before running the code below.

# make sure to update diffusers
# pip install -U diffusers
import torch
from diffusers import StableDiffusion3Pipeline

pipe = StableDiffusion3Pipeline.from_pretrained(
	"stabilityai/stable-diffusion-3.5-large", torch_dtype=torch.bfloat16
).to("cuda")

image = pipe(
    prompt="a photo of a cat holding a sign that says hello world",
    negative_prompt="",
    num_inference_steps=40,
    height=1024,
    width=1024,
    guidance_scale=4.5,
).images[0]

image.save("sd3_hello_world.png")

Follow the documentation to know more.

Cogview3-plus

We added a new text-to-image model, Cogview3-plus, from the THUDM team! The model is DiT-based and supports image generation from 512 to 2048px. Thanks to @zRzRzRzRzRzRzR for contributing it!

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.30.3: CogVideoX Image-to-Video und Video-to-Video

Version 0.30.3 ergänzt die Pipelines CogVideoXImageToVideoPipeline und CogVideoXVideoToVideoPipeline sowie Tiled Encoding im CogVideoX-VAE über vae.enable_tiling().

This patch release adds Diffusers support for the upcoming CogVideoX-5B-I2V release (an Image-to-Video generation model)! The model weights will be available by end of the week on the HF Hub at THUDM/CogVideoX-5b-I2V (Link). Stay tuned for the release!

This release features two new pipelines:

  • CogVideoXImageToVideoPipeline
  • CogVideoXVideoToVideoPipeline

Additionally, we now have support for tiled encoding in the CogVideoX VAE. This can be enabled by calling the vae.enable_tiling() method, and it is used in the new Video-to-Video pipeline to encode sample videos to latents in a memory-efficient manner.

CogVideoXImageToVideoPipeline

The code below demonstrates how to use the new image-to-video pipeline:

import torch
from diffusers import CogVideoXImageToVideoPipeline
from diffusers.utils import export_to_video, load_image

pipe = CogVideoXImageToVideoPipeline.from_pretrained("THUDM/CogVideoX-5b-I2V", torch_dtype=torch.bfloat16)
pipe.to("cuda")

# Optionally, enable memory optimizations.
# If enabling CPU offloading, remember to remove `pipe.to("cuda")` above
pipe.enable_model_cpu_offload()
pipe.vae.enable_tiling()

prompt = "An astronaut hatching from an egg, on the surface of the moon, the darkness and depth of space realised in the background. High quality, ultrarealistic detail and breath-taking movie-like camera shot."
image = load_image( …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.30.2: Standard-Repository für Single File aktualisiert

Version 0.30.2 aktualisiert das Runway-Repository für Single-File-Laden, behebt das Wiederholen von Flux-CLIP-Prompt-Embeds bei num_images_per_prompt > 1 und korrigiert cache_dir sowie local_files_only beim IP-Adapter-Image-Encoder.

All commits

  • update runway repo for single_file by @yiyixuxu in #9323
  • Fix Flux CLIP prompt embeds repeat for num_images_per_prompt > 1 by @DN6 in #9280
  • [IP Adapter] Fix cache_dir and local_files_only for image encoder by @asomoza in #9272

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.30.1: CogVideoX-5B und Fehlerbehebungen

Version 0.30.1 unterstützt CogVideoX-5B und führt VAE-Tiling über enable_tiling() ein, wodurch der Speicherbedarf mit CPU-Offloading auf 12 GB für CogVideoX-2B und 21 GB für CogVideoX-5B sinkt.

CogVideoX-5B

This patch release adds diffusers support for the upcoming CogVideoX-5B release! The model weights will be available next week on the Huggingface Hub at THUDM/CogVideoX-5b. Stay tuned for the release!

Additionally, we have implemented VAE tiling feature, which reduces the memory requirement for CogVideoX models. With this update, the total memory requirement is now 12GB for CogVideoX-2B and 21GB for CogVideoX-5B (with CPU offloading). To Enable this feature, simply call enable_tiling() on the VAE.

The code below shows how to generate a video with CogVideoX-5B

import torch
from diffusers import CogVideoXPipeline
from diffusers.utils import export_to_video

prompt = "Tracking shot,late afternoon light casting long shadows,a cyclist in athletic gear pedaling down a scenic mountain road,winding path with trees and a lake in the background,invigorating and adventurous atmosphere."

pipe = CogVideoXPipeline.from_pretrained(
    "THUDM/CogVideoX-5b",
    torch_dtype=torch.bfloat16
)

pipe.enable_model_cpu_offload()
pipe.vae.enable_tiling()

video = pipe(
    prompt=prompt,
    num_videos_per_prompt=1,
    num_inference_steps=50,
    num_frames=49,
    guidance_scale=6,
).frames[0]

export_to_video(video, "output.mp4", fps=8)

https://github.com/user-attachments/assets/c2d4f7e8-ef86-4da6-8085-cb9f83f47f34 …

Originalquelle(öffnet in neuem Tab)Problem melden