Zum Inhalt springen

Fireworks AI Release Notes

79 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Dokumentation zum Learning-Rate-Scheduler für SFT

Die Learning-Rate-Scheduler-Einstellungen constant, linear und cosine für Supervised-Fine-Tuning-Jobs sind jetzt über firectl und das REST-API-Objekt lrScheduler dokumentiert.

<Badge color="purple">Training</Badge>

SFT learning rate scheduler documentation

Documented learning rate scheduler settings for supervised fine-tuning jobs, including constant, linear, and cosine schedules via firectl and the REST API lrScheduler object.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Serverless-Einstellung von Kimi K2.5 und Qwen 3.6 Plus

Kimi K2.5 und Qwen 3.6 Plus sind auf Serverless veraltet; empfohlen wird die Migration zu Kimi K2.6 bzw. Qwen 3.7 Plus.

<Badge color="blue">Inference</Badge>

Serverless deprecation: Kimi K2.5 and Qwen 3.6 Plus

Kimi K2.5 and Qwen 3.6 Plus are deprecated from serverless.

Recommended migrations

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Serverless-Einstellung von MiniMax M2.5

MiniMax M2.5 ist auf Serverless veraltet, Nutzer sollen zu MiniMax M2.7 migrieren.

<Badge color="blue">Inference</Badge>

Serverless deprecation: MiniMax M2.5

MiniMax M2.5 is deprecated from serverless. Migrate to MiniMax M2.7.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Audio-Inferenz und Bildgenerierung veraltet

Audio-Inferenz und Bildgenerierung sind als veraltet gekennzeichnet.

<Badge color="blue">Inference</Badge>

Audio inference and image generation deprecation

Audio inference and image generation are deprecated.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Serverless-Abkündigung älterer Modelle zum 14. Mai 2026

Mehrere ältere Serverless-Modelle, darunter DeepSeek V3.1/V3.2, GLM 4.7, GLM 5, Qwen3 8B und Llama 3.3 70B Instruct, werden am 14. Mai 2026 abgeschaltet, sodass Nutzer vorher auf empfohlene Ersatzmodelle migrieren müssen, während Dedicated Deployments unberührt bleiben.

Several legacy serverless models will be decommissioned on May 14, 2026 to make room for newer, higher-performance releases. This applies only to serverless usage; dedicated deployments are unaffected.

Action required

If you use any of the models below on serverless, migrate to a recommended replacement before May 14, 2026. After that date, they will no longer be available via serverless endpoints.

Serverless migration guide

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Serverless-Abschaltung älterer Modelle zum 14. Mai 2026

Mehrere ältere Serverless-Modelle (u. a. DeepSeek V3.1/V3.2, GLM 4.7, GLM 5, Qwen3 8B) werden am 14. Mai 2026 abgeschaltet, dedizierte Deployments bleiben unberührt, und Nutzer sollen vorher auf empfohlene Ersatzmodelle wie Kimi K2.6, GLM 5.1 oder GPT-OSS 20B migrieren.

<Badge color="blue">Inference</Badge>

Serverless deprecation: legacy models removed May 14, 2026

Several legacy serverless models will be decommissioned on May 14, 2026 to make room for newer, higher-performance releases. This applies only to serverless usage; dedicated deployments are unaffected.

Action required

If you use any of the models below on serverless, migrate to a recommended replacement before May 14, 2026. After that date, they will no longer be available via serverless endpoints.

Serverless migration guide

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Video- und Audio-Modelle, AWS-S3-Training und SSO-JIT-Provisioning

Multimodale Modelle wie Qwen3 Omni und Molmo2 verarbeiten nun Video- und Audioeingaben über die Chat Completions API, Trainingsdatasets können in eigenen AWS-S3-Buckets liegen, und Enterprise-SSO unterstützt Just-In-Time-Benutzerprovisionierung.

<Badge color="blue">Inference</Badge> <Badge color="purple">Training</Badge> <Badge color="gray">Platform</Badge>

Video & Audio Models, AWS S3 Training Integration, and SSO Improvements

Video & Audio Input Models

You can now query multimodal models with video and audio inputs for video captioning, scene analysis, and multimodal question answering. Deploy models like Qwen3 Omni and Molmo2 to process video and audio content directly using the Chat Completions API.

See the Video & Audio Inputs guide for deployment instructions and code examples.

AWS S3 Integration for Training Datasets

Training datasets can now be stored in your own AWS S3 buckets using GCP-to-AWS OIDC federation. This Bring Your Own Bucket (BYOB) approach keeps your data private while enabling secure access during Supervised Fine-Tuning and Reinforcement Fine-Tuning jobs—no long-lived credentials required.

See the Secure Training (BYOB) documentation for IAM role setup and usage examples.

Just-In-Time (JIT) User Provisioning for SSO (Enterprise)

JIT user provisioning automatically creates user accounts when users sign in through SSO for the first time. Enable this when configuring your identity provider to eliminate manual user creation.

See the SSO documentation for setup instructions.

📚 Documentation Updates …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Warm-Start-Training und Azure Federated Identity

Reinforcement-Fine-Tuning-Jobs lassen sich per --warm-start-from von SFT-Checkpoints starten, und Modell-Uploads aus Azure Blob Storage unterstützen Azure-AD-Federated-Identity statt SAS-Tokens.

<Badge color="purple">Training</Badge> <Badge color="gray">Platform</Badge>

Warm-Start Training and Azure Model Uploads

Warm-Start Training for Reinforcement Fine-Tuning

You can now warm-start Reinforcement Fine-Tuning jobs from previously supervised fine-tuned checkpoints using the --warm-start-from flag. This enables a streamlined SFT-to-RFT workflow where you first train a model with supervised fine-tuning, then continue training with reinforcement learning.

See the Warm-Start Training guide for details.

Azure Federated Identity for Model Uploads

Model uploads from Azure Blob Storage now support Azure AD federated identity authentication as an alternative to SAS tokens. This eliminates the need for credential rotation and enables secure, credential-less authentication.

See the Uploading Custom Models documentation for setup instructions.

📚 Documentation Updates

  • Warm-Start Training: New guide for SFT-to-RFT workflows (Warm-Start Training)
  • Azure Federated Identity: Setup instructions for Azure AD authentication (Uploading Custom Models)
  • Preserved Thinking: Multi-turn reasoning with preserved thinking context (Reasoning)
  • GLM 4.7: Added to models supporting reasoning_effort parameter …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Playground-Kategorien, neue Rollen und Fine-Tuning-Verbesserungen

Der Playground erhält Kategorie-Tabs (LLM, Image, TTS, STT), es gibt die neuen Benutzerrollen Contributor und Inference, und Fine-Tuning-Jobs lassen sich stoppen, fortsetzen und klonen sowie RFT-Ausgabedatasets herunterladen.

<Badge color="blue">Inference</Badge> <Badge color="purple">Training</Badge> <Badge color="gray">Platform</Badge>

Playground Categories, New User Roles, Fine-Tuning Improvements, and New Models

Playground Categories

The Playground now features category tabs (LLM, Image, TTS, STT) in the header for easier switching between model types. The playground automatically detects the appropriate category based on the selected model and provides smart defaults for each category.

User Roles: Contributor and Inference

New user roles provide more granular access control for team collaboration:

  • Contributor: Read and write access to resources without administrative privileges
  • Inference: Read-only access with the ability to run inference on deployments

Assign these roles when inviting team members to provide appropriate access levels.

Fine-Tuning Improvements

Fine-tuning workflows have been enhanced with several new capabilities:

  • Stop and Resume Jobs: Stop running fine-tuning jobs and resume them later from where they left off. Available for Supervised Fine-Tuning and Reinforcement Fine-Tuning jobs.
  • Clone Jobs: Quickly create new fine-tuning jobs based on existing job configurations using the Clone action.
  • Download Output Datasets: Download output datasets from Reinforcement Fine-Tuning jobs, including individual files or bulk download as a ZIP archive. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Reasoning-Guide, Prompt-Caching-Updates und neue Modelle

Es gibt einen neuen Reasoning-Guide, aktualisierte Prompt-Caching-Dokumentation (gecachte Prompt-Tokens auf Serverless kosten 50 % weniger) sowie die neuen Modelle Devstral Small 2 24B Instruct 2512 und NVIDIA Nemotron Nano 3 30B A3B.

<Badge color="blue">Inference</Badge> <Badge color="purple">Training</Badge> <Badge color="gray">Platform</Badge>

Reasoning Guide, Prompt Caching Updates, New Models and CLI Updates

Reasoning Guide

A new Reasoning guide is now available in the documentation. This comprehensive guide covers:

  • Accessing reasoning_content from thinking/reasoning models
  • Controlling reasoning effort with the reasoning_effort parameter
  • Streaming with reasoning content
  • Interleaved thinking for multi-step tool-calling workflows

The guide provides code examples using the Fireworks Python SDK and explains how to work with models that support extended reasoning capabilities.

Prompt Caching Updates

Prompt caching documentation has been updated with expanded guidance:

  • Cached prompt tokens on serverless now cost 50% less than uncached tokens
  • Session affinity routing via the user field or x-session-affinity header for improved cache hit rates
  • Prompt optimization techniques for maximizing cache efficiency

See the Prompt Caching guide for details.

✨ New Models

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

DeepSeek V3.2 auf Serverless, Preisanzeige für Cache und neue Modelle

DeepSeek V3.2 ist auf Serverless verfügbar, die Model Library zeigt gecachte und ungecachte Input-Token-Preise an, das Evaluations-Dashboard erhält Statusspalte und Filter, und Ministral 3 14B und 8B Instruct 2512 sind neu.

<Badge color="blue">Inference</Badge> <Badge color="purple">Training</Badge> <Badge color="gray">Platform</Badge>

DeepSeek V3.2 on Serverless, Cached Token Pricing, and New Models

☁️ Serverless

Cached Token Pricing Display

The Model Library and model detail pages now display cached and uncached input token pricing for serverless models that support prompt caching. This gives you better visibility into potential cost savings when using prompt caching with supported models.

Evaluations Dashboard Improvements

The Evaluations dashboard has been enhanced with new filtering and status tracking capabilities:

  • Status column showing evaluator build state (Active, Building, Failed)
  • Quick filters to filter evaluators and evaluation jobs by status
  • Improved table layout with actions integrated into the status column

✨ New Models

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Audit Logs, Dataset-Download und gewichtetes RFT-Training

Die Web-App bietet jetzt durchsuchbare Audit Logs und Dataset-Downloads, und Reinforcement Fine-Tuning unterstützt Gewichtung pro Beispiel.

<Badge color="blue">Inference</Badge> <Badge color="purple">Training</Badge> <Badge color="gray">Platform</Badge>

Audit Logs, Dataset Download, Weighted Training for Reinforcement Fine-Tuning, and New Model

Audit Logs in Web App

You can now view and search audit logs directly from the Fireworks web app. The new Audit Logs page provides:

  • Search and filter logs by status and timeframe
  • Detailed view panel for individual log entries
  • Easy navigation from the console sidebar under Account settings

See the Audit Logs documentation for more information.

Dataset Download

You can now download datasets directly from the Fireworks web app. The new download functionality allows you to:

  • Download individual files from a dataset
  • Download all files at once with "Download All"
  • Access downloads from the Datasets table in the dashboard

Weighted Training for Reinforcement Fine-Tuning

Reinforcement Fine-Tuning now supports per-example weighting, giving you more control over which samples have greater influence during training. This feature mirrors the weighted training functionality already available in Supervised Fine-Tuning.

See the Weighted Training documentation for details on the weight field format.

✨ New Models …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Evaluator-Verbesserungen, Kimi K2 Thinking auf Serverless und neue API-Endpunkte

Evaluatoren lassen sich über GitHub-Templates erstellen, Integrationen mit W&B und MLflow sind dokumentiert, und Kimi K2 Thinking sowie KAT Dev 32B und 72B Exp sind verfügbar, Kimi K2 Thinking auch auf Serverless.

<Badge color="blue">Inference</Badge> <Badge color="purple">Training</Badge> <Badge color="gray">Platform</Badge>

Evaluator Improvements, Kimi K2 Thinking on Serverless, and New API Endpoints

Improved Evaluator Creation Experience

The evaluator creation workflow has been significantly enhanced with GitHub template integration. You can now:

  • Fork evaluator templates directly from GitHub repositories
  • Browse and preview templates before using them
  • Create evaluators with a streamlined save dialog
  • View evaluators in a new sortable and paginated table

MLOps & Observability Integrations

New documentation for integrating Fireworks with MLOps and observability tools:

  • Weights & Biases (W&B) integration for experiment tracking during fine-tuning
  • MLflow integration for model management and experiment logging

✨ New Models

☁️ Serverless

📚 New REST API Endpoints …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Build SDK wird eingestellt, verbessertes RFT

Das Build SDK (letzte Version 0.19.20) wird zugunsten eines neuen, aus der REST API generierten Python SDK ab Version 1.0.0 eingestellt, und die Reinforcement-Fine-Tuning-Erfahrung wurde unter anderem mit Multi-Turn-Training verbessert.

<Badge color="purple">Training</Badge>

☀️ Sunsetting Build SDK

The Build SDK is being deprecated in favor of a new Python SDK generated directly from our REST API. The new SDK is more up-to-date, flexible, and continuously synchronized with our REST API. Please note that the last version of the Build SDK will be 0.19.20, and the new SDK will start at 1.0.0. Python package managers will not automatically update to the new SDK, so you will need to manually update your dependencies and refactor your code.

Existing codebases using the Build SDK will continue to function as before and will not be affected unless you choose to upgrade to the new SDK version.

The new SDK replaces the Build SDK's LLM and Dataset classes with REST API-aligned methods. If you upgrade to version 1.0.0 or later, you will need to migrate your code.

🚀 Improved RFT Experience

We've drastically improved the RFT experience with better reliability, developer-friendly SDK for hooking up your existing agents, support for multi-turn training, better observability in our Web App, and better overall developer experience.

See Reinforcement Fine-Tuning for more details.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

SFT mit separaten Thinking Traces für Reasoning-Modelle

Supervised Fine-Tuning unterstützt nun separate Thinking Traces für Reasoning-Modelle wie DeepSeek R1, GPT OSS und Qwen3 Thinking sowie Multi-Turn-Fine-Tuning für die GPT-OSS-Familie.

<Badge color="purple">Training</Badge>

Supervised Fine-Tuning

We now support supervised fine tuning with separate thinking traces for reasoning models (e.g. DeepSeek R1, GPT OSS, Qwen3 Thinking etc) that ensures training-inference consistency. An example including thinking traces would look like:

  {
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "What is the capital of France?"}, 
      {"role": "assistant", "content": "Paris.", "reasoning_content": "The user is asking about the capital city of France, it should be Paris."}
    ]
  }
  {
    "messages": [
      {"role": "user", "content": "What is 1+1?"},
      {"role": "assistant", "content": "2", "weight": 0, "reasoning_content": "The user is asking about the result of 1+1, the answer is 2."},
      {"role": "user", "content": "Now what is 2+2?"},
      {"role": "assistant", "content": "4", "reasoning_content": "The user is asking about the result of 2+2, the answer should be 4."}
    ]
  }

We are also properly supporting multi-turn fine tuning (with or without thinking traces) for GPT OSS model family that ensures training-inference consistency.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

SFT-Unterstützung für Qwen3 MoE und GPT OSS

Supervised Fine-Tuning unterstützt nun Qwen3-MoE-Modelle und GPT-OSS-Modelle, wobei GPT OSS vorerst nur Single-Turn ohne Thinking Traces abdeckt.

<Badge color="purple">Training</Badge>

Supervised Fine-Tuning

We now support Qwen3 MoE model (Qwen3 dense models are already supported) and GPT OSS models for supervised fine-tuning. GPT OSS model fine tunning support is single-turn without thinking traces at the moment.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Fine-Tuning von Vision-Language-Modellen und .apply()-Pflicht im Build SDK

Vision-Language-Modelle der Qwen-2.5-VL-Familie (3B bis 72B) lassen sich mit Bild- und Textdaten feinabstimmen, und das Build SDK verlangt für On-Demand-Deployments nun einen expliziten .apply()-Aufruf.

<Badge color="purple">Training</Badge> <Badge color="gray">Platform</Badge>

🎨 Vision-Language Model Fine-Tuning

You can now fine-tune Vision-Language Models (VLMs) on Fireworks AI using the Qwen 2.5 VL model family. This extends our Supervised Fine-tuning V2 platform to support multimodal training with both images and text data.

Supported models:

  • Qwen 2.5 VL 3B Instruct
  • Qwen 2.5 VL 7B Instruct
  • Qwen 2.5 VL 32B Instruct
  • Qwen 2.5 VL 72B Instruct

Features:

  • Fine-tune on datasets containing both images and text in JSONL format with base64-encoded images
  • Support for up to 64K context length during training
  • Built on the same Supervised Fine-tuning V2 infrastructure as text models

See the VLM fine-tuning documentation for setup instructions and dataset formatting requirements.

🔧 Build SDK: Deployment Configuration Application Requirement

The Build SDK now requires you to call .apply() to apply any deployment configurations to Fireworks when using deployment_type="on-demand" or deployment_type="on-demand-lora". This change ensures explicit control over when deployments are created and helps prevent accidental deployment creation.

Key changes:

  • .apply() is now required for on-demand and on-demand-lora deployments
  • Serverless deployments do not require .apply() calls …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Eigene Rollout- und Reward-Entwicklung für Reinforcement Learning

Nutzer können eigene Rollout- und Reward-Logik entwickeln, während Fireworks Training und Deployment übernimmt, und dafür die neue Methode LLM.reinforcement_step() verwenden.

<Badge color="purple">Training</Badge>

🚀 Bring Your Own Rollout and Reward Development for Reinforcement Learning

You can now develop your own custom rollout and reward functionality while using Fireworks to manage the training and deployment of your models. This gives you full control over your reinforcement learning workflows while leveraging Fireworks' infrastructure for model training and deployment.

See the new LLM.reinforcement_step() method and ReinforcementStep class for usage examples and details.

Originalquelle(öffnet in neuem Tab)Problem melden