Zum Inhalt springen

Fireworks AI Release Notes

79 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Folge Fireworks AI, um die Release Notes in deinen Feed zu holen.

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

SFT V2 unterstützt Llama 4 MoE

Supervised Fine-Tuning V2 unterstützt nun die Llama-4-MoE-Modelle Scout und Maverick, jeweils nur Text.

<Badge color="purple">Training</Badge>

Supervised Fine-Tuning V2

We now support Llama 4 MoE model supervised fine-tuning (Llama 4 Scout, Llama 4 Maverick, Text only).

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Build SDK: Überarbeitete Deployment-Logik der LLM-Klasse

Die Deployment-Logik der LLM-Klasse im Build SDK wurde überarbeitet: id ist bei "on-demand" und base_id bei "on-demand-lora" nun Pflicht, deployment_display_name ist optional und standardmäßig der Dateiname, und ein Deployment mit gleicher id wird wiederverwendet statt neu angelegt.

<Badge color="gray">Platform</Badge>

🏗️ Build SDK LLM Deployment Logic Refactor

Based on early feedback from users and internal testing, we've refactored the LLM class deployment logic in the Build SDK to make it easier to understand.

Key changes:

  • The id parameter is now required when deployment_type is "on-demand"
  • The base_id parameter is now required when deployment_type is "on-demand-lora"
  • The deployment_display_name parameter is now optional and defaults to the filename where the LLM was instantiated

A new deployment will be created if a deployment with the same id does not exist. Otherwise, the existing deployment will be reused.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Responses API jetzt im Python SDK verfügbar

Die Responses API lässt sich jetzt im Python SDK verwenden.

<Badge color="blue">Inference</Badge>

🚀 Support for Responses API in Python SDK

You can now use the Responses API in the Python SDK. This is useful if you want to use the Responses API in your own applications.

See the Responses API guide for usage examples and details.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Anmeldung bei Fireworks mit LinkedIn

Die Anmeldung bei Fireworks ist jetzt mit einem LinkedIn-Konto möglich, auf der Login-Seite sowie per firectl login in der CLI, wobei die primäre LinkedIn-E-Mail-Adresse zur Kontoidentifikation dient.

<Badge color="gray">Platform</Badge>

Support for LinkedIn authentication

You can now log in to Fireworks using your LinkedIn account. This is useful if you already have a LinkedIn account and want to use it to log in to Fireworks.

To log in with LinkedIn, go to the Fireworks login page and click the "Continue with LinkedIn" button.

You can also log in with LinkedIn from the CLI using the firectl login command.

How it works:

  • Fireworks uses your LinkedIn primary email address for account identification
  • You can switch between different Fireworks accounts by changing your LinkedIn primary email
  • See our LinkedIn authentication FAQ for detailed instructions on managing email addresses

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

GitHub-Anmeldung und Einstellung von Document Inlining

Fireworks unterstützt jetzt die Anmeldung per GitHub-Konto (Web und firectl login), während Document Inlining (#transform=inline) eingestellt wurde und nicht mehr verfügbar ist.

<Badge color="blue">Inference</Badge> <Badge color="gray">Platform</Badge>

Support for GitHub authentication

You can now log in to Fireworks using your GitHub account. This is useful if you already have a GitHub account and want to use it to log in to Fireworks.

To log in with GitHub, go to the Fireworks login page and click the "Continue with GitHub" button.

You can also log in with GitHub from the CLI using the firectl login command.

🚨 Document Inlining Deprecation

Document Inlining has been deprecated and is no longer available on the Fireworks platform. This feature allowed LLMs to process images and PDFs through the chat completions API by appending #transform=inline to document URLs.

Migration recommendations:

  • For image processing: Use Vision Language Models (VLMs) like Qwen2.5-VL 32B Instruct
  • For PDF processing: Use dedicated PDF processing libraries combined with text-based LLMs
  • For structured extraction: Leverage our structured responses capabilities

For assistance with migration, please contact our support team or visit our Discord community.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Build SDK mit reward-kit und neue Responses API

Das Build SDK integriert nun reward-kit für die Entwicklung eigener Evaluatoren für Reinforcement Fine-Tuning, und die neue Responses API unterstützt mehrstufige Konversationen mit previous_response_id, Streaming und Steuerung der Speicherung über store.

<Badge color="blue">Inference</Badge> <Badge color="purple">Training</Badge>

🎯 Build SDK: Reward-kit integration for evaluator development

The Build SDK now natively integrates with reward-kit to simplify evaluator development for Reinforcement Fine-Tuning (RFT). You can now create custom evaluators in Python with automatic dependency management and seamless deployment to Fireworks infrastructure.

Key features:

  • Native reward-kit integration for evaluator development
  • Automatic packaging of dependencies from pyproject.toml or requirements.txt
  • Local testing capabilities before deployment
  • Direct integration with Fireworks datasets and evaluation jobs
  • Support for third-party libraries and complex evaluation logic

See our Developing Evaluators guide to get started with your first evaluator in minutes.

Added new Responses API for advanced conversational workflows and integrations

  • Continue conversations across multiple turns using the previous_response_id parameter to maintain context without resending full history
  • Stream responses in real time as they are generated for responsive applications
  • Control response storage with the store parameter—choose whether responses are retrievable by ID or ephemeral …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Supervised Fine-Tuning V2 und Reinforcement Fine-Tuning

Supervised Fine-Tuning V2 mit Unterstützung mehrerer Modellfamilien, längerem Kontextfenster, Multi-Turn-Function-Calling-Fine-Tuning und Quantization aware training sowie Reinforcement Fine-Tuning (RFT) wurden veröffentlicht.

<Badge color="purple">Training</Badge>

Supervised Fine-Tuning V2

Supervised Fine-Tuning V2 released.

Key features:

  • Supports Qwen 2/2.5/3 series, Phi 4, Gemma 3, the Llama 3 family, Deepseek V2, V3, R1
  • Longer context window up to full context length of the supported models
  • Multi-turn function calling fine-tuning
  • Quantization aware training

More details in the blogpost.

Reinforcement Fine-Tuning (RFT)

Reinforcement Fine-Tuning released. Train expert models to surpass closed source frontier models through verifiable reward. More details in blospost.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Diarization und Batch-Verarbeitung für Audio-Inferenz

Die Audio-Inferenz unterstützt jetzt Diarization und Batch-Verarbeitung.

<Badge color="blue">Inference</Badge>

Diarization and batch processing support added to audio inference

See our blog post for details.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

LoRA-Fine-Tunes mit einem Befehl und schneller deployen

Ein LoRA-Fine-Tune lässt sich nun mit einem einzigen firectl deployment create-Befehl deployen und erreicht ungefähr die Geschwindigkeit des Basismodells, was bisher zwei Schritte erforderte und langsamer war.

<Badge color="purple">Training</Badge>

🚀 Easier & faster LoRA fine-tune deployments on Fireworks

You can now deploy a LoRA fine-tune with a single command and get speeds that approximately match the base model:

firectl deployment create "accounts/fireworks/models/<MODEL_ID of lora model>"

Previously, this involved two distinct steps, and the resulting deployment was slower than the base model:

  1. Create a deployment using firectl deployment create "accounts/fireworks/models/<MODEL_ID of base model>" --enable-addons
  2. Then deploy the addon to the deployment: firectl load-lora <MODEL_ID> --deployment <DEPLOYMENT_ID>

For more information, see our deployment documentation.

<Note> This change is for dedicated deployments with a single LoRA. You can still deploy multiple LoRAs on a deployment as described in the documentation. </Note>

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Fireworks AI

Höhere Serverless-Rate-Limits für Small-Modelle

Die adaptiven Rate-Limit-Obergrenzen für Small-Serverless-Modelle steigen auf 108M Total Prompt TPM, 27M Uncached Prompt TPM und 1,08M Generated TPM, und die Small-Stufe umfasst nun Modelle mit weniger als 600B Gesamtparametern.

We increased the adaptive rate-limit ceilings for Small serverless models and expanded the Small tier to cover models with less than 600B total parameters.

The Small tier now has these ceilings:

  • Total Prompt TPM: 108M
  • Uncached Prompt TPM: 27M
  • Generated TPM: 1.08M

This applies to Small-tier serverless models such as GLM 5.3 Flash, DeepSeek V4.1 Flash, and OpenAI GPT OSS 120B. Medium and Large model ceilings are unchanged.

See Serverless rate limits for the full tier table. </Update>

<Update label="2026-10-01"> <Badge color="purple">Training</Badge>

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Fireworks AI

Managed RFT pausiert – RL über die Training API nutzen

Managed Reinforcement Fine-Tuning (RFT) ist pausiert und nimmt keine neuen Jobs mehr an, stattdessen wird RL über die Training API empfohlen, während Managed SFT und DPO unverändert bleiben und 70 Modelle nicht mehr feinabstimmbar sind.

Managed reinforcement fine-tuning (RFT) is paused. Managed Training no longer accepts new RFT jobs from the Fireworks UI, firectl, or the REST API. Existing jobs stay visible in your dashboard, and models you already trained with managed RFT keep serving.

Try RL on the Training API, where you write the rollout and training loop yourself and Fireworks runs the GPUs. Compared with managed RFT, you also get:

  • Full-parameter RL on most current models, not just LoRA
  • The training shape's full context length, up to 524K tokens, instead of managed RFT's fixed 32K limit
  • Other methods such as on-policy distillation (OPD) and custom objectives

Your evaluator logic carries over. Start with Cookbook: Reinforcement Learning.

Managed SFT and DPO are unaffected.

Models no longer available for fine-tuning

The 70 models below were tunable only through managed RFT. With managed RFT paused, none of them can be fine-tuned on Fireworks anymore, and they no longer appear on the Models page. Inference on these models is not affected by this change.

<Accordion title="Full list (70 models)"> …

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Fireworks AI

Neues Trainingsmodell: DeepSeek V4.1 Flash

DeepSeek V4.1 Flash steht auf der Dedicated Training API für LoRA-Training mit bis zu 262K Kontext bereit, wobei SFT von Beginn an und RL-Unterstützung bald folgt.

DeepSeek V4.1 Flash is now available for LoRA training on the Dedicated Training API, with up to 262K context. It's a strong base for agentic coding, terminal automation, and tool use. SFT is supported at launch, and RL support is coming soon.

See the training model catalog for current availability. </Update>

<Update label="2026-09-26"> <Badge color="blue">Inference</Badge>

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Fireworks AI

Serverless-Abschaltung: DeepSeek V4 Pro (0813), V4 Flash (0731) u. a.

Die angekündigte Serverless-Abschaltung ist in Kraft, sodass unter anderem DeepSeek V4 Pro (0813), DeepSeek V4 Flash (0731) und Kimi K2.6 nicht mehr auf öffentlichem Serverless verfügbar sind, während Dedicated Deployments unberührt bleiben und Migrationen etwa zu DeepSeek V4.1 Flash empfohlen werden.

The serverless deprecation announced for September 25, 2026 is now in effect. The models below are no longer available on public serverless, including Fast and US-only serverless endpoints where those existed. Dedicated deployments are unaffected.

Recommended migrations

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Fireworks AI

Neues Serverless-Training-Modell: GLM 5.3 Flash

GLM 5.3 Flash ist für LoRA-Workloads im gemeinsamen Serverless-Training-Pool verfügbar und unterstützt bis zu 200K Kontext sowie Text- und Bildeingaben.

GLM 5.3 Flash is now available for LoRA workloads on the shared Serverless Training pool, with up to 200K context and both text and vision inputs.

See the Serverless Training guide for setup and the training model catalog for current availability. </Update>

<Update label="2026-09-16"> <Badge color="purple">Training</Badge>

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Fireworks AI

Neues Serverless-Training-Modell: GLM 5.3

GLM 5.3 ist für LoRA-Workloads im gemeinsamen Serverless-Training-Pool mit bis zu 262K Kontext verfügbar, ohne Kapazitätsreservierung und mit Abrechnung pro Token.

GLM 5.3 is now available for LoRA workloads on the shared Serverless Training pool, with up to 262K context. There's no capacity to reserve and you pay per token. Move the same loop to Dedicated Training for your most demanding workloads.

See the Serverless Training guide for setup and the training model catalog for current availability. </Update>

<Update label="2026-09-16"> <Badge color="gray">Platform</Badge>

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Fireworks AI

Neue Deployment-Flags: deploymentShape "default" und acceptShapelessRisk

Die Create-Deployment-API, firectl und das Python SDK erhalten die Optionen deploymentShape "default" und acceptShapelessRisk, wobei das Erstellen von Deployments ohne Shape künftig ein ausdrückliches Opt-in erfordern wird.

Two new options are available on the Create Deployment API, in firectl (--deployment-shape default / --accept-shapeless-risk), and in the Python SDK (deployment_shape="default" / accept_shapeless_risk=True):

  • deploymentShape: "default" — Fireworks picks a validated deployment shape for the model and creates the deployment from it. If every compatible shape conflicts with fields in your request, the request fails with an error naming the conflicting fields and compatible shapes; the pick never silently overrides your settings or falls back to creating without a shape.
  • acceptShapelessRisk=true — an explicit opt-out that creates the deployment without a shape, preserving current behavior. It cannot be combined with a shape.

Deployments created without a shape skip shape validation and are the most common cause of failed deployment creations. Enforcement is coming soon: shapeless creation will then require the explicit opt-in, so start passing a shape (or default) now. The opt-out is for advanced users only. If you have a workload no existing shape covers, contact us and we'll help you find or add one. </Update>

<Update label="2026-09-13"> <Badge color="purple">Training</Badge>

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Fireworks AI

Neues Trainingsmodell: GLM 5.3 Flash

GLM 5.3 Flash ist auf der Dedicated Training API für LoRA-Training einschließlich Vision-Training mit bis zu 262K Kontext verfügbar.

GLM 5.3 Flash is now available for LoRA training on the Dedicated Training API, including vision training, with up to 262K context. It performs well on agentic coding, document analysis, and tool use, and is cost-efficient to serve.

See the training model catalog for current availability. </Update>

<Update label="2026-09-12"> <Badge color="blue">Inference</Badge>

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Fireworks AI

Serverless-Abkündigung älterer DeepSeek-, GLM-, Muse- und Kimi-Modelle

Mehrere ältere Serverless-Modelle werden am 25. September 2026 abgeschaltet, sodass Nutzer vorher auf empfohlene Nachfolger wie DeepSeek V4.1 Flash oder GLM 5.3 migrieren müssen, während Dedicated Deployments unberührt bleiben.

Several older Serverless models will be decommissioned on September 25, 2026 to better serve newer, higher-performance replacements. This applies only to serverless endpoints, including Fast and US-only Serverless endpoints for models that have those variants. Dedicated deployments are unaffected.

Action required

If you use any of the models below on serverless, migrate to a recommended replacement before September 25, 2026. After that date, they will no longer be available via serverless endpoints.

Recommended migrations

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Fireworks AI

Training Cost Estimator

Der neue Training Cost Estimator schätzt vor dem Start eines Trainingsjobs die Kosten, basierend auf veröffentlichten Token-Preisen bei Managed und Serverless bzw. GPU-Stundenpreisen bei Dedicated, wobei die Schätzungen keine verbindlichen Angebote sind.

The new training cost estimator helps you estimate what a training job will cost before you run it.

Managed and Serverless estimates use published per-token rates. Dedicated estimates use allocated GPU-hour rates. Planning estimates are not quotes. </Update>

<Update label="2026-09-09"> <Badge color="purple">Training</Badge>

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Fireworks AI

Neuer Fireworks Training Skill für Coding-Agents

Ein neuer Fireworks Training Skill für Claude Code, Cursor, Codex und weitere Coding-Agents plant aus einer Beschreibung in natürlicher Sprache einen Trainingslauf, schätzt die Kosten und wartet vor Ausgaben auf eine Freigabe.

A new Fireworks training skill is available for Claude Code, Cursor, Codex, and other compatible coding agents. Describe a training goal in plain language to plan a run, estimate cost, and wait for approval before spend.

See Agent Skills for install commands. </Update>

<Update label="2026-09-08"> <Badge color="gray">Platform</Badge>

Originalquelle(öffnet in neuem Tab)Problem melden