Zum Inhalt springen

Fireworks AI Release Notes

79 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Verdoppelte Serverless-Rate-Limits für Small-Modelle

Die adaptiven Rate-Limit-Obergrenzen für Small-Serverless-Modelle (unter 600B Parameter) wurden auf 216M Total Prompt TPM, 54M Uncached Prompt TPM und 2,16M Generated TPM verdoppelt, während Medium und Large unverändert bleiben.

<Badge color="blue">Inference</Badge>

Doubled serverless rate limits for Small models

We doubled the adaptive rate-limit ceilings for Small serverless models (less than 600B total parameters).

The Small tier now has these ceilings:

  • Total Prompt TPM: 216M
  • Uncached Prompt TPM: 54M
  • Generated TPM: 2.16M

Medium and Large model ceilings are unchanged. Serverless rate limits now also lists which tier each Serverless model falls in.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Höhere Serverless-Rate-Limits für Small-Modelle

Die Rate-Limit-Obergrenzen der Small-Stufe wurden auf 108M Total Prompt TPM, 27M Uncached Prompt TPM und 1,08M Generated TPM erhöht, und die Stufe umfasst nun Modelle mit weniger als 600B Parametern.

<Badge color="blue">Inference</Badge>

Higher serverless rate limits for Small models

We increased the adaptive rate-limit ceilings for Small serverless models and expanded the Small tier to cover models with less than 600B total parameters.

The Small tier now has these ceilings:

  • Total Prompt TPM: 108M
  • Uncached Prompt TPM: 27M
  • Generated TPM: 1.08M

This applies to Small-tier serverless models such as GLM 5.3 Flash, DeepSeek V4.1 Flash, and OpenAI GPT OSS 120B. Medium and Large model ceilings are unchanged.

See Serverless rate limits for the full tier table.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus dem Text des Eintrags.

Erstmals gesehen am .

Fireworks AI

Neue Serverless-Preise für DeepSeek V4.1 Flash ab 1. Oktober 2026

Ab dem 1. Oktober 2026 um 00:00 UTC ändern sich die Serverless-Preise für DeepSeek V4.1 Flash, etwa Standard-Output von 0,66 auf 1,20 US-Dollar pro 1M Tokens, während Dedicated Deployments und Reserved Throughput unberührt bleiben.

On October 1, 2026 at 00:00 UTC, serverless pricing for DeepSeek V4.1 Flash changes (uncached input / cached input / output price per 1M tokens):

  • Standard: $0.22 / $0.007 / $0.66 → $0.30 / $0.006 / $1.20
  • Priority: $0.275 / $0.00875 / $0.825 → $0.375 / $0.0075 / $1.50

This adjustment brings our pricing in line with current market rates for this model. It applies only to serverless usage. If you run DeepSeek V4.1 Flash on a dedicated deployment or use Reserved Throughput, your pricing is unaffected.

We are also rolling out infrastructure improvements designed to improve cache hit rate, minimize cost per task, and deliver a faster, more reliable experience across the board.

See Serverless pricing for the full rate card. </Update>

<Update label="2026-09-28"> <Badge color="purple">Training</Badge>

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Managed RFT pausiert, RL über die Training API

Managed Reinforcement Fine-Tuning (RFT) nimmt keine neuen Jobs mehr an, bestehende Modelle laufen weiter, und 70 nur darüber trainierbare Modelle lassen sich nicht mehr fine-tunen; als Alternative wird RL über die Training API empfohlen, Managed SFT und DPO bleiben unberührt.

<Badge color="purple">Training</Badge>

Managed RFT is paused; try RL on the Training API

Managed reinforcement fine-tuning (RFT) is paused. Managed Training no longer accepts new RFT jobs from the Fireworks UI, firectl, or the REST API. Existing jobs stay visible in your dashboard, and models you already trained with managed RFT keep serving.

Try RL on the Training API, where you write the rollout and training loop yourself and Fireworks runs the GPUs. Compared with managed RFT, you also get:

  • Full-parameter RL on most current models, not just LoRA
  • The training shape's full context length, up to 524K tokens, instead of managed RFT's fixed 32K limit
  • Other methods such as on-policy distillation (OPD) and custom objectives

Your evaluator logic carries over. Start with Cookbook: Reinforcement Learning.

Managed SFT and DPO are unaffected.

Models no longer available for fine-tuning

The 70 models below were tunable only through managed RFT. With managed RFT paused, none of them can be fine-tuned on Fireworks anymore, and they no longer appear on the Models page. Inference on these models is not affected by this change.

<Accordion title="Full list (70 models)"> …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Serverless-Preisänderung für DeepSeek V4.1 Flash

Ab 1. Oktober 2026 um 00:00 UTC steigen die Serverless-Preise für DeepSeek V4.1 Flash im Standard-Tarif auf $0.30 / $0.006 / $1.20 und im Priority-Tarif auf $0.375 / $0.0075 / $1.50 pro 1M Tokens, dedizierte Deployments und Reserved Throughput sind nicht betroffen.

<Badge color="blue">Inference</Badge>

Serverless pricing update: DeepSeek V4.1 Flash

On October 1, 2026 at 00:00 UTC, serverless pricing for DeepSeek V4.1 Flash changes (uncached input / cached input / output price per 1M tokens):

  • Standard: $0.22 / $0.007 / $0.66 → $0.30 / $0.006 / $1.20
  • Priority: $0.275 / $0.00875 / $0.825 → $0.375 / $0.0075 / $1.50

This adjustment brings our pricing in line with current market rates for this model. It applies only to serverless usage. If you run DeepSeek V4.1 Flash on a dedicated deployment or use Reserved Throughput, your pricing is unaffected.

We are also rolling out infrastructure improvements designed to improve cache hit rate, minimize cost per task, and deliver a faster, more reliable experience across the board.

See Serverless pricing for the full rate card.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Neues Trainingsmodell: DeepSeek V4.1 Flash

DeepSeek V4.1 Flash ist jetzt für LoRA-Training auf der Dedicated Training API mit bis zu 262K Kontext verfügbar, SFT wird unterstützt und RL-Support folgt bald.

<Badge color="purple">Training</Badge>

New training model: DeepSeek V4.1 Flash

DeepSeek V4.1 Flash is now available for LoRA training on the Dedicated Training API, with up to 262K context. It's a strong base for agentic coding, terminal automation, and tool use. SFT is supported at launch, and RL support is coming soon.

See the training model catalog for current availability.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Serverless-Abkündigung: DeepSeek V4 Pro (0813), V4 Flash (0731) und weitere

Die angekündigte Abschaltung mehrerer Modelle, darunter DeepSeek V4 Pro (0813), DeepSeek V4 Flash (0731) und Kimi K2.6, ist auf Public Serverless in Kraft getreten, dedizierte Deployments bleiben unberührt, und es werden Migrationsziele wie DeepSeek V4.1 Flash empfohlen.

<Badge color="blue">Inference</Badge>

Serverless deprecation: DeepSeek V4 Pro (0813), DeepSeek V4 Flash (0731), and related models

The serverless deprecation announced for September 25, 2026 is now in effect. The models below are no longer available on public serverless, including Fast and US-only serverless endpoints where those existed. Dedicated deployments are unaffected.

Recommended migrations

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Neues Serverless-Training-Modell: GLM 5.3 Flash

GLM 5.3 Flash ist jetzt für LoRA-Workloads im gemeinsamen Serverless-Training-Pool mit bis zu 200K Kontext sowie Text- und Vision-Eingaben verfügbar.

<Badge color="purple">Training</Badge>

New Serverless Training model: GLM 5.3 Flash

GLM 5.3 Flash is now available for LoRA workloads on the shared Serverless Training pool, with up to 200K context and both text and vision inputs.

See the Serverless Training guide for setup and the training model catalog for current availability.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Neues Serverless-Training-Modell: GLM 5.3

GLM 5.3 ist jetzt für LoRA-Workloads im gemeinsamen Serverless-Training-Pool mit bis zu 262K Kontext verfügbar, ohne Kapazitätsreservierung und mit Abrechnung pro Token.

<Badge color="purple">Training</Badge>

New Serverless Training model: GLM 5.3

GLM 5.3 is now available for LoRA workloads on the shared Serverless Training pool, with up to 262K context. There's no capacity to reserve and you pay per token. Move the same loop to Dedicated Training for your most demanding workloads.

See the Serverless Training guide for setup and the training model catalog for current availability.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Neue Deployment-Optionen: deploymentShape "default" und acceptShapelessRisk

Beim Erstellen von Deployments lässt sich per deploymentShape: "default" eine validierte Shape automatisch wählen oder per acceptShapelessRisk=true ausdrücklich ohne Shape erstellen, wobei die Shape-Pflicht künftig durchgesetzt werden soll.

<Badge color="gray">Platform</Badge>

New deployment creation flags: deploymentShape: "default" and acceptShapelessRisk

Two new options are available on the Create Deployment API, in firectl (--deployment-shape default / --accept-shapeless-risk), and in the Python SDK (deployment_shape="default" / accept_shapeless_risk=True):

  • deploymentShape: "default" — Fireworks picks a validated deployment shape for the model and creates the deployment from it. If every compatible shape conflicts with fields in your request, the request fails with an error naming the conflicting fields and compatible shapes; the pick never silently overrides your settings or falls back to creating without a shape.
  • acceptShapelessRisk=true — an explicit opt-out that creates the deployment without a shape, preserving current behavior. It cannot be combined with a shape.

Deployments created without a shape skip shape validation and are the most common cause of failed deployment creations. Enforcement is coming soon: shapeless creation will then require the explicit opt-in, so start passing a shape (or default) now. The opt-out is for advanced users only. If you have a workload no existing shape covers, contact us and we'll help you find or add one.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Neues Trainingsmodell: GLM 5.3 Flash

GLM 5.3 Flash ist jetzt für LoRA-Training auf der Dedicated Training API einschließlich Vision-Training mit bis zu 262K Kontext verfügbar.

<Badge color="purple">Training</Badge>

New training model: GLM 5.3 Flash

GLM 5.3 Flash is now available for LoRA training on the Dedicated Training API, including vision training, with up to 262K context. It performs well on agentic coding, document analysis, and tool use, and is cost-efficient to serve.

See the training model catalog for current availability.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Ankündigung: Serverless-Abkündigung älterer DeepSeek-, GLM-, Muse- und Kimi-Modelle

Mehrere ältere Serverless-Modelle werden am 25. September 2026 abgeschaltet, sodass Nutzer vorher auf empfohlene Nachfolger wie DeepSeek V4.1 Flash oder GLM 5.3 migrieren müssen, während dedizierte Deployments unberührt bleiben.

<Badge color="blue">Inference</Badge>

Upcoming Serverless deprecation: older DeepSeek, GLM, Muse, and Kimi models

Several older Serverless models will be decommissioned on September 25, 2026 to better serve newer, higher-performance replacements. This applies only to serverless endpoints, including Fast and US-only Serverless endpoints for models that have those variants. Dedicated deployments are unaffected.

Action required

If you use any of the models below on serverless, migrate to a recommended replacement before September 25, 2026. After that date, they will no longer be available via serverless endpoints.

Recommended migrations

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Neuer Trainingskosten-Schätzer

Ein neuer Trainingskosten-Schätzer hilft, die Kosten eines Trainingsjobs vor dem Start abzuschätzen, wobei die Schätzungen keine verbindlichen Angebote sind.

<Badge color="purple">Training</Badge>

Training cost estimator

The new training cost estimator helps you estimate what a training job will cost before you run it.

Managed and Serverless estimates use published per-token rates. Dedicated estimates use allocated GPU-hour rates. Planning estimates are not quotes.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Deployment-Tags und Änderungen an der Annotation-API

firectl 1.8.3 bietet Befehle zum Setzen, Entfernen und Auflisten von Deployment-Tags, während die REST-API für kundenverwaltete Annotationen das Präfix custom/ verlangt, bloße Schlüssel mit HTTP 403 ablehnt und beim Lesen nur custom/*-Einträge zurückgibt.

<Badge color="gray">Platform</Badge>

Deployment tags and annotation API changes

Deployment tags are customer-managed entries stored in a deployment's annotations map. firectl presents logical keys such as environment; the REST API represents the same key as custom/environment.

  • firectl: Version 1.8.3 adds deployment tag set, unset, and list, including atomic batch operations.
  • REST writes: Customer-managed annotation keys must begin with custom/. Bare keys now return HTTP 403 (PERMISSION_DENIED).
  • REST reads: For regular account users, GetDeployment and ListDeployments return only custom/* annotation entries. Keys outside that namespace are omitted without an error.

Existing stored annotations were not rewritten. Clients using a bare key such as environment should set custom/environment and update reads to use that canonical key.

See Deployment Tags for commands, REST examples, validation rules, and migration guidance.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Neues Trainingsmodell: GLM 5.3

GLM 5.3 ist jetzt für LoRA-Training auf der Dedicated Training API mit bis zu 204K Kontext verfügbar und für komplexes Coding und langfristige Agenten gedacht.

<Badge color="purple">Training</Badge>

New training model: GLM 5.3

GLM 5.3 is now available for LoRA training on the Dedicated Training API, with up to 204K context. It's built for complex coding and long-horizon agents.

See the training model catalog for current availability.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Serverless-Rate-Limit-Obergrenzen skalieren nach Modellgröße

Die adaptiven Serverless-Rate-Limit-Obergrenzen richten sich jetzt nach der Modellgrößenstufe, wobei kleinere Modelle höhere und große Modelle niedrigere Obergrenzen erhalten.

<Badge color="blue">Inference</Badge>

Serverless rate limit ceilings now scale by model size

Serverless adaptive rate limit ceilings now vary by model size tier. Smaller models get higher ceilings, medium models get intermediate ceilings, and large models use lower ceilings. See Serverless rate limits for current tier thresholds and ceiling values.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Neue Serverless-Training-Modelle: DeepSeek V4 Flash 0731, Qwen 3.8 27B, Muse Glimmer 30B

DeepSeek V4 Flash 0731 (bis 262K Kontext), Qwen 3.8 27B und Muse Glimmer 30B (je bis 128K Kontext) sind jetzt für LoRA-Workloads im gemeinsamen Serverless-Training-Pool verfügbar.

<Badge color="purple">Training</Badge>

New Serverless Training models: DeepSeek V4 Flash 0731, Qwen 3.8 27B, and Muse Glimmer 30B

The following models are now available for LoRA workloads on the shared Serverless Training pool:

See the Serverless Training guide for setup and the training model catalog for current availability.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Serverless-Abkündigung: MiniMax M2.7, GPT OSS 20B, Kimi- und DeepSeek-Modelle

MiniMax M2.7, GPT OSS 20B, Kimi K2.6 Turbo/Fast, Kimi K2.7 Code Fast und DeepSeek V4 Pro sind ab dem 27. August 2026 auf Serverless abgekündigt, und es werden jeweils Ersatzmodelle zur Migration empfohlen.

<Badge color="blue">Inference</Badge>

Serverless deprecation: MiniMax M2.7, GPT OSS 20B, Kimi K2.6 Turbo/Fast, Kimi K2.7 Code Fast, DeepSeek V4 Pro

The following models are deprecated from serverless effective August 27, 2026.

Recommended migrations

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Serverless-Training-Abkündigung: Qwen 3.5 9B und Qwen 3.6 27B

Qwen 3.5 9B und Qwen 3.6 27B sind ab dem 26. August 2026 im Serverless Training abgekündigt, und Workloads sollen auf Qwen 3.8 27B migriert werden.

<Badge color="purple">Training</Badge>

Serverless Training deprecation: Qwen 3.5 9B and Qwen 3.6 27B

Qwen 3.5 9B and Qwen 3.6 27B are deprecated from Serverless Training effective August 26, 2026.

Migrate new and existing Serverless Training workloads to Qwen 3.8 27B.

This change applies to the shared Serverless Training pool. Check the training model catalog for availability on other training surfaces.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fireworks AI

Serverless-Abkündigung: DeepSeek V4 Flash

DeepSeek V4 Flash ist auf Serverless abgekündigt, und Nutzer sollen auf DeepSeek V4 Flash (0731) migrieren.

<Badge color="blue">Inference</Badge>

Serverless deprecation: DeepSeek V4 Flash

DeepSeek V4 Flash is deprecated from serverless. Migrate to DeepSeek V4 Flash (0731).

Originalquelle(öffnet in neuem Tab)Problem melden