Zum Inhalt springen

Hugging Face Release Notes

65 Einträge aus 3 Quellen. Zuletzt aktualisiert:

Folge Hugging Face, um die Release Notes in deinen Feed zu holen.

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.28.1: HunyuanDiT und Transformer2D-Modellvarianten

Version 0.28.1 führt die mehrsprachige Hunyuan-DiT-Pipeline von Tencent ein und ergänzt Klassenvarianten für Transformer2DModel.

This patch release primarily introduces the Hunyuan DiT pipeline from the Tencent team.

Hunyuan DiT

image

Hunyuan DiT is a transformer-based diffusion pipeline, introduced in the Hunyuan-DiT : A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding paper by the Tencent Hunyuan.

import torch
from diffusers import HunyuanDiTPipeline

pipe = HunyuanDiTPipeline.from_pretrained(
    "Tencent-Hunyuan/HunyuanDiT-Diffusers", torch_dtype=torch.float16
)
pipe.to("cuda")

# You may also use English prompt as HunyuanDiT supports both English and Chinese
# prompt = "An astronaut riding a horse"
prompt = "一个宇航员在骑马"
image = pipe(prompt).images[0]

🧠 This pipeline has support for multi-linguality.

📜 Refer to the official docs here to learn more about it.

Thanks to @gnobitab, for contributing Hunyuan DiT in #8240.

All commits

  • Release: v0.28.0 by @sayakpaul (direct commit on v0.28.1-patch)
  • [Core] Introduce class variants for Transformer2DModel by @sayakpaul in #7647
  • resolve comflicts by @toshas (direct commit on v0.28.1-patch)
  • Tencent Hunyuan Team: add HunyuanDiT related updates by @gnobitab in #8240 …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.28.0: Marigold, PixArt Sigma, AnimateDiff SDXL und mehr

Diffusers 0.28.0 führt mit Marigold die erste offizielle Pipeline für diskriminative Aufgaben wie Tiefen- und Oberflächennormalen-Schätzung ein und bringt laut Titel außerdem PixArt Sigma, AnimateDiff SDXL, InstantStyle und ein VQGAN-Trainingsskript.

Diffusion models are known for their abilities in the space of generative modeling. This release of diffusers introduces the first official pipeline (Marigold) for discriminative tasks such as depth estimation and surface normals’ estimation!

Starting this release, we will also highlight the changes and features from the library that make it easy to integrate community checkpoints, features, and so on. Read on!

Marigold

Proposed in Marigold: Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation, Marigold introduces a diffusion model and associated fine-tuning protocol for monocular depth estimation. It can also be extended to perform surface normals’ estimation.

marigold

(Image taken from the official repository)

The code snippet below shows how to use this pipeline for depth estimation:

import diffusers
import torch

pipe = diffusers.MarigoldDepthPipeline.from_pretrained(
    "prs-eth/marigold-depth-lcm-v1-0", variant="fp16", torch_dtype=torch.float16
).to("cuda")

image = diffusers.utils.load_image("https://marigoldmonodepth.github.io/images/einstein.jpg")
depth = pipe(image)

vis = pipe.image_processor.visualize_depth(depth.prediction)
vis[0].save("einstein_depth.png") …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.27.2: Fehlerbehebungen für add_noise, StableCascade und LoRA

Version 0.27.2 behebt einen Fehler in der Scheduler-Funktion add_noise, Probleme mit Embeddings im Stable-Cascade-Decoder bei mehreren Bild-Embeddings pro Prompt sowie Probleme mit cross_attention_kwargs und dem scale-Argument bei LoRA.

All commits

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.27.1: Klarere Handhabung des scale-Arguments bei LoRA

Version 0.27.1 entfernt das LoRA-scale-Argument aus den weitergereichten Argumenten, damit es nicht mehr an andere Komponenten propagiert wird, und räumt so die Verwirrung um dieses Argument auf.

All commits

  • Release: v0.27.0 by @DN6 (direct commit on v0.27.1-patch)
  • [LoRA] pop the LoRA scale so that it doesn't get propagated to the weeds by @sayakpaul in #7338
  • Release: 0.27.1-patch by @sayakpaul (direct commit on v0.27.1-patch)

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Diffusers von Hugging Face

Diffusers 0.27.0: Stable Cascade, Playground v2.5, EDM-Training und mehr

Diffusers 0.27.0 ergänzt Unterstützung für das Text-zu-Bild-Modell Stable Cascade (nicht-kommerzielle Lizenz, torch>=2.2.0 für bfloat16) sowie Playground v2.5, EDM-Style-Training und IP-Adapter-Image-Embeds.

Stable Cascade

We are adding support for a new text-to-image model building on Würstchen called Stable Cascade, which comes with a non-commercial license. The Stable Cascade line of pipelines differs from Stable Diffusion in that they are built upon three distinct models and allow for hierarchical compression of image patients, achieving remarkable outputs.

from diffusers import StableCascadePriorPipeline, StableCascadeDecoderPipeline
import torch

prior = StableCascadePriorPipeline.from_pretrained(
    "stabilityai/stable-cascade-prior",
    torch_dtype=torch.bfloat16,
).to("cuda")

prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k"
image_emb = prior(prompt=prompt).image_embeddings[0]

decoder = StableCascadeDecoderPipeline.from_pretrained(
    "stabilityai/stable-cascade",
    torch_dtype=torch.bfloat16,
).to("cuda")

image = pipe(image_embeddings=image_emb, prompt=prompt).images[0]
image

📜 Check out the docs here to know more about the model.

Note: You will need a torch>=2.2.0 to use the torch.bfloat16 data type with the Stable Cascade pipeline.

Playground v2.5 …

Originalquelle(öffnet in neuem Tab)Problem melden