Zum Inhalt springen

Diffusers Updates & Release Notes

30 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.30.0: Neue Pipelines wie Flux, Stable Audio und CogVideoX

Diffusers 0.30.0 bringt neue Pipelines für Audio (Stable Audio), Video (Latte, CogVideoX) und Bilder (Lumina, Kolors, AuraFlow, Flux) sowie neue Methoden wie FreeNoise und SparseCtrl und Refactorings.

New pipelines

Untitled

Image taken from the Lumina’s GitHub.

This release features many new pipelines. Below, we provide a list:

Audio pipelines 🎼

Video pipelines 📹

  • Latte (thanks to @maxin-cn for the contribution through #8404)
  • CogVideoX (thanks to @zRzRzRzRzRzRzR for the contribution through #9082)

Image pipelines 🎇

Be sure to check out the respective docs to know more about these pipelines. Some additional pointers are below for curious minds:

  • Lumina introduces a new DiT architecture that is multilingual in nature.
  • Kolors is inspired by SDXL and is also multilingual in nature. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.29.2: Fixes für Deprecation und LoRA

Version 0.29.2 behebt einen Shape-Fehler bei SD3 ohne T5 und num_images_per_prompt > 1, korrigiert das Laden von LoRA und DoRA, überarbeitet die LoRA-Konvertierung und entfernt eine Deprecation zur Ausgabeklasse von transformer2d.

All commits

  • [SD3] Fix mis-matched shape when num_images_per_prompt > 1 using without T5 (text_encoder_3=None) by @Dalanke in #8558
  • [LoRA] refactor lora conversion utility. by @sayakpaul in #8295
  • [LoRA] fix conversion utility so that lora dora loads correctly by @sayakpaul in #8688
  • [Chore] remove deprecation from transformer2d regarding the output class. by @sayakpaul in #8698
  • [LoRA] fix vanilla fine-tuned lora loading. by @sayakpaul in #8691
  • Release: v0.29.2 by @sayakpaul (direct commit on v0.29.2-patch)

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.29.1: SD3 ControlNet und erweiterter from_single_file-Support

Version 0.29.1 bringt die SD3-ControlNet-Pipeline, erweiterte from_single_file-Unterstützung für alle SD3-Single-File-Checkpoints, lange Prompts mit dem T5-Textencoder und Fehlerbehebungen.

SD3 CntrolNet

<img width="624" alt="image" src="https://github.com/huggingface/diffusers/assets/46553287/db384753-cfbb-488c-bc74-8280f9bee24e">
import torch
from diffusers import StableDiffusion3ControlNetPipeline
from diffusers.models import SD3ControlNetModel, SD3MultiControlNetModel
from diffusers.utils import load_image

controlnet = SD3ControlNetModel.from_pretrained("InstantX/SD3-Controlnet-Canny", torch_dtype=torch.float16)

pipe = StableDiffusion3ControlNetPipeline.from_pretrained(
    "stabilityai/stable-diffusion-3-medium-diffusers", controlnet=controlnet, torch_dtype=torch.float16
)
pipe.to("cuda")
control_image = load_image("https://huggingface.co/InstantX/SD3-Controlnet-Canny/resolve/main/canny.jpg")
prompt = "A girl holding a sign that says InstantX"
image = pipe(prompt, control_image=control_image, controlnet_conditioning_scale=0.7).images[0]
image.save("sd3.png")

📜 Refer to the official docs here to learn more about it.

Thanks to @haofanwang @wangqixun from the @ResearcherXman team for contributing this pipeline!

Expanded single file support

We now support all available single-file checkpoints for sd3 in diffusers! To load the single file checkpoint with t5

import torch
from diffusers import StableDiffusion3Pipeline

pipe = StableDiffusion3Pipeline.from_single_file( …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.29.0: Stable Diffusion 3

Diffusers 0.29.0 unterstützt Stable Diffusion 3 von Stability AI für Text-zu-Bild-Generierung, wobei das gesperrte Modell zuvor auf Hugging Face freigeschaltet und per Login verwendet werden muss.

This release emphasizes Stable Diffusion 3, Stability AI’s latest iteration of the Stable Diffusion family of models. It was introduced in Scaling Rectified Flow Transformers for High-Resolution Image Synthesis by Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yannik Marek, and Robin Rombach.

As the model is gated, before using it with diffusers, you first need to go to the Stable Diffusion 3 Medium Hugging Face page, fill in the form and accept the gate. Once you are in, you need to log in so that your system knows you’ve accepted the gate.

huggingface-cli login

The code below shows how to perform text-to-image generation with SD3:

import torch
from diffusers import StableDiffusion3Pipeline

pipe = StableDiffusion3Pipeline.from_pretrained("stabilityai/stable-diffusion-3-medium-diffusers", torch_dtype=torch.float16)
pipe = pipe.to("cuda")

image = pipe(
    "A cat holding a sign that says hello world",
    negative_prompt="",
    num_inference_steps=28,
    guidance_scale=7.0,
).images[0]
image

image …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.28.2: Fehler bei from_single_file mit CLIP-Checkpoints behoben

Version 0.28.2 ändert den Checkpoint-Schlüssel, mit dem CLIP-Modelle in Single-File-Checkpoints erkannt werden, und behebt so einen Fehler bei from_single_file.

  • Change checkpoint key used to identify CLIP models in single file checkpoints by @DN6 in #8319

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.28.1: HunyuanDiT und Transformer2D-Modellvarianten

Version 0.28.1 führt die mehrsprachige Hunyuan-DiT-Pipeline von Tencent ein und ergänzt Klassenvarianten für Transformer2DModel.

This patch release primarily introduces the Hunyuan DiT pipeline from the Tencent team.

Hunyuan DiT

image

Hunyuan DiT is a transformer-based diffusion pipeline, introduced in the Hunyuan-DiT : A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding paper by the Tencent Hunyuan.

import torch
from diffusers import HunyuanDiTPipeline

pipe = HunyuanDiTPipeline.from_pretrained(
    "Tencent-Hunyuan/HunyuanDiT-Diffusers", torch_dtype=torch.float16
)
pipe.to("cuda")

# You may also use English prompt as HunyuanDiT supports both English and Chinese
# prompt = "An astronaut riding a horse"
prompt = "一个宇航员在骑马"
image = pipe(prompt).images[0]

🧠 This pipeline has support for multi-linguality.

📜 Refer to the official docs here to learn more about it.

Thanks to @gnobitab, for contributing Hunyuan DiT in #8240.

All commits

  • Release: v0.28.0 by @sayakpaul (direct commit on v0.28.1-patch)
  • [Core] Introduce class variants for Transformer2DModel by @sayakpaul in #7647
  • resolve comflicts by @toshas (direct commit on v0.28.1-patch)
  • Tencent Hunyuan Team: add HunyuanDiT related updates by @gnobitab in #8240 …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.28.0: Marigold, PixArt Sigma, AnimateDiff SDXL und mehr

Diffusers 0.28.0 führt mit Marigold die erste offizielle Pipeline für diskriminative Aufgaben wie Tiefen- und Oberflächennormalen-Schätzung ein und bringt laut Titel außerdem PixArt Sigma, AnimateDiff SDXL, InstantStyle und ein VQGAN-Trainingsskript.

Diffusion models are known for their abilities in the space of generative modeling. This release of diffusers introduces the first official pipeline (Marigold) for discriminative tasks such as depth estimation and surface normals’ estimation!

Starting this release, we will also highlight the changes and features from the library that make it easy to integrate community checkpoints, features, and so on. Read on!

Marigold

Proposed in Marigold: Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation, Marigold introduces a diffusion model and associated fine-tuning protocol for monocular depth estimation. It can also be extended to perform surface normals’ estimation.

marigold

(Image taken from the official repository)

The code snippet below shows how to use this pipeline for depth estimation:

import diffusers
import torch

pipe = diffusers.MarigoldDepthPipeline.from_pretrained(
    "prs-eth/marigold-depth-lcm-v1-0", variant="fp16", torch_dtype=torch.float16
).to("cuda")

image = diffusers.utils.load_image("https://marigoldmonodepth.github.io/images/einstein.jpg")
depth = pipe(image)

vis = pipe.image_processor.visualize_depth(depth.prediction)
vis[0].save("einstein_depth.png") …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.27.2: Fehlerbehebungen für add_noise, StableCascade und LoRA

Version 0.27.2 behebt einen Fehler in der Scheduler-Funktion add_noise, Probleme mit Embeddings im Stable-Cascade-Decoder bei mehreren Bild-Embeddings pro Prompt sowie Probleme mit cross_attention_kwargs und dem scale-Argument bei LoRA.

All commits

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.27.1: Klarere Handhabung des scale-Arguments bei LoRA

Version 0.27.1 entfernt das LoRA-scale-Argument aus den weitergereichten Argumenten, damit es nicht mehr an andere Komponenten propagiert wird, und räumt so die Verwirrung um dieses Argument auf.

All commits

  • Release: v0.27.0 by @DN6 (direct commit on v0.27.1-patch)
  • [LoRA] pop the LoRA scale so that it doesn't get propagated to the weeds by @sayakpaul in #7338
  • Release: 0.27.1-patch by @sayakpaul (direct commit on v0.27.1-patch)

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.27.0: Stable Cascade, Playground v2.5, EDM-Training und mehr

Diffusers 0.27.0 ergänzt Unterstützung für das Text-zu-Bild-Modell Stable Cascade (nicht-kommerzielle Lizenz, torch>=2.2.0 für bfloat16) sowie Playground v2.5, EDM-Style-Training und IP-Adapter-Image-Embeds.

Stable Cascade

We are adding support for a new text-to-image model building on Würstchen called Stable Cascade, which comes with a non-commercial license. The Stable Cascade line of pipelines differs from Stable Diffusion in that they are built upon three distinct models and allow for hierarchical compression of image patients, achieving remarkable outputs.

from diffusers import StableCascadePriorPipeline, StableCascadeDecoderPipeline
import torch

prior = StableCascadePriorPipeline.from_pretrained(
    "stabilityai/stable-cascade-prior",
    torch_dtype=torch.bfloat16,
).to("cuda")

prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image_emb = prior(prompt=prompt).image_embeddings[0]

decoder = StableCascadeDecoderPipeline.from_pretrained(
    "stabilityai/stable-cascade",
    torch_dtype=torch.bfloat16,
).to("cuda")

image = pipe(image_embeddings=image_emb, prompt=prompt).images[0]
image

📜 Check out the docs here to know more about the model.

Note: You will need a torch>=2.2.0 to use the torch.bfloat16 data type with the Stable Cascade pipeline.

Playground v2.5 …

Originalquelle(öffnet in neuem Tab)Problem melden