Zum Inhalt springen

KI- und ML-Infrastruktur: Release Notes

Plattformen und Dienste, auf denen KI-Modelle trainiert, betrieben und bereitgestellt werden. 11 Hersteller, 382 Einträge.

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Gemini Enterprise Agent Platform von Google

Gemini Nano Banana 2.1 ist allgemein verfügbar (GA)

Gemini Nano Banana 2.1 (gemini-nano-banana-2.1) ist jetzt allgemein verfügbar und bietet schnelle multimodale Bildgenerierung und -bearbeitung mit besserer Bildqualität, Prompt-Treue und Textdarstellung in den Ausgabeauflösungen 1K, 2K und 4K.

Feature

Gemini Nano Banana 2.1

Gemini Nano Banana 2.1 (gemini-nano-banana-2.1) is available in General Availability (GA). Gemini Nano Banana 2.1 is optimized for high-speed multimodal image generation and editing, offering improved visual quality, prompt adherence, and text rendering across 1K, 2K, and 4K output resolutions.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Kein Datum in der Quelle. Angezeigt ist der Tag, an dem wir den Eintrag aufgenommen haben.

Aufgenommen am .

Together AI

Rollouts für dedizierte Modell-Inferenz

Mit Rollouts lässt sich Live-Traffic per Canary-, Blue-Green- oder Rolling-Strategie ohne Änderung der Endpoint-URL von einem Deployment auf ein anderes verlagern, optional mit Metrik-Gates, und über tg beta endpoints rollout oder die Konsole steuern.

Rollouts shift live traffic from one deployment to another under the same endpoint, without changing the endpoint URL. Pick a canary, blue-green, or rolling strategy to determine how traffic moves, and optionally gate a canary rollout on live metrics so it pauses automatically if the new deployment regresses.

Start a rollout with the tg beta endpoints rollout CLI command or from the endpoint's Rollouts tab in the console, then pause, resume, promote, or cancel it at any point while it runs.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Kein Datum in der Quelle. Angezeigt ist der Tag, an dem wir den Eintrag aufgenommen haben.

Aufgenommen am .

Together AI

Automatische Leerlauf-Abschaltung für dedizierte Deployments

Dedizierte Deployments können mit --inactive-timeout (bzw. inactiveTimeout in der Management-API) nach der eingestellten Zeit ohne Inferenz-Anfragen automatisch auf null Replikate skalieren, wodurch Hardware freigegeben und die Abrechnung gestoppt wird.

Deployments can now stop themselves when they go unused. Set an inactivity timeout with --inactive-timeout (the inactiveTimeout field in the management API), and if the deployment serves no inference requests for that many minutes, it scales to zero replicas, releasing its hardware and stopping billing.

See Automatic idle shutdown. </Update>

<Update label="September 15, 2026" tags={["New releases", "Deprecations"]}>

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Kein Datum in der Quelle. Angezeigt ist der Tag, an dem wir den Eintrag aufgenommen haben.

Aufgenommen am .

Together AI

Preissenkung für Qwen3.7-Max und Qwen3.8-Flash

Ab dem 22. September 2026 sinken die Preise für Qwen/Qwen3.7-Max (1,50 $ Input / 4,50 $ Output) und Qwen/Qwen3.8-Flash (0,09 $ Input / 0,282 $ Output) pro 1M Tokens.

The following models have lower pricing, effective September 22, 2026. All usage from that date forward is billed at the new rates (per 1M tokens):

  • Qwen/Qwen3.7-Max: $2.50 → $1.50 (input), $7.50 → $4.50 (output).
  • Qwen/Qwen3.8-Flash: $0.15 → $0.09 (input), $0.47 → $0.282 (output).

See Serverless models for the full pricing catalog. </Update>

<Update label="September 16, 2026" tags={["New releases"]}>

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Kein Datum in der Quelle. Angezeigt ist der Tag, an dem wir den Eintrag aufgenommen haben.

Aufgenommen am .

Together AI

Neues serverloses Modell: Tev1-4B-experimental

Das Modell together/Tev1-4B-experimental mit 32.768 Kontextlänge ist jetzt serverless verfügbar, zu 0,042 $ Input und kostenlosem Output pro 1M Tokens.

The following models are now available on serverless:

  • together/Tev1-4B-experimental: 32,768 context length. Pricing: $0.042 input / free output (per 1M tokens).</Update>

<Update label="September 22, 2026" tags={["Pricing"]}>

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Kein Datum in der Quelle. Angezeigt ist der Tag, an dem wir den Eintrag aufgenommen haben.

Aufgenommen am .

Together AI

Fine-Tuned-Modell per Registry-Namen bereitstellen

Seit Together CLI Version 2.24.0 akzeptiert tg beta endpoints deploy anstelle der model_object_id auch den Registry-Namen model_object_name eines abgeschlossenen Fine-Tuning-Jobs.

Since Together CLI version 2.24.0, tg beta endpoints deploy accepts a completed fine-tuning job's model_object_name, the qualified <project_slug>/<model_name> registry name, in place of its model_object_id. The CLI resolves the name to the same model, so you can deploy straight from the name shown in the fine-tuning jobs dashboard. The SDK and API take model_object_id.

See Deploy a fine-tuned model. </Update>

<Update label="September 23, 2026" tags={["New models"]}>

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Together AI

Batch-API-Dateien werden jetzt 7 Tage aufbewahrt

Eingabe-, Ausgabe- und Fehlerdateien von Batch-Jobs werden jetzt 7 Tage aufbewahrt, danach sind sie nicht mehr abrufbar und eine hochgeladene Eingabedatei lässt sich nur in diesem Zeitraum für weitere Jobs wiederverwenden.

The input file you upload for a batch job, along with the output and error files the job produces, are now retained for 7 days. After that the files are no longer accessible, so download your results before the window closes. Reusing an uploaded input file across batch jobs also works only within that window.

See Batch inference. </Update>

<Update label="September 24, 2026" tags={["Improvements"]}>

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Kein Datum in der Quelle. Angezeigt ist der Tag, an dem wir den Eintrag aufgenommen haben.

Aufgenommen am .

Together AI

Längerer Fine-Tuning-Kontext für Qwen-27B-Modelle

Die Qwen-27B-Modelle unterstützen beim Fine-Tuning nun 131.072 Token Kontext für SFT und 65.536 für DPO, dafür sinkt die Batch-Größe bei LoRA-Jobs auf 2 und bei DPO mit Full Fine-Tuning auf maximal 8.

Qwen/Qwen3.8-27B, Qwen/Qwen3.6-27B, and Qwen/Qwen3.5-27B now support a 131,072-token context for SFT (up from 32,768) and 65,536 for DPO (up from 16,384), for both LoRA and full fine-tuning. Batch size limits have also changed: LoRA jobs on these models run at a batch size of 2, down from 16, and the maximum DPO batch size for full fine-tuning drops from 16 to 8.

See Supported models for each model's limits.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Kein Datum in der Quelle. Angezeigt ist der Tag, an dem wir den Eintrag aufgenommen haben.

Aufgenommen am .

Together AI

Höheres LoRA-Rang-Limit von 128 beim Fine-Tuning

LoRA-Adapter lassen sich für die meisten Modelle nun mit einem Rang von bis zu 128 (zuvor 64) trainieren, wobei der Standardrang 64 bleibt und frühere CLI- und SDK-Versionen ohne gesetztes lora_r jetzt den Maximalrang 128 verwenden.

You can now train LoRA adapters with a rank of up to 128 for the majority of models, up from 64. The default rank for these models stays at 64, so set lora_r to use a higher one.

The model limits response has a new lora_training.default_rank field next to lora_training.max_rank. Run tg fine-tuning model-limits <model> to see both values for a model.

Version 2.36.0 of the Together CLI and Python SDK uses the default rank when you don't set lora_r. Earlier versions, including the 1.x SDK, use the model's maximum rank instead, which is now 128 on these models. Upgrade to 2.36.0 or set lora_r yourself, especially if you continue training from a rank-64 adapter, where the rank has to match.

In the console, the rank field now starts at the model's default rank (64 on most models) instead of 8.

See Supported models for each model's default and maximum rank.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Kein Datum in der Quelle. Angezeigt ist der Tag, an dem wir den Eintrag aufgenommen haben.

Aufgenommen am .

Together AI

Neues Code-Sandbox-SDK und CLI: together-sandbox

Das neue together-sandbox SDK (Python, TypeScript und CLI) führt Befehle und Code in isolierten, aus Docker-Image-Snapshots erzeugten Umgebungen aus, authentifiziert sich mit dem Together API Key, setzt auf Snapshots statt Hibernate/Resume und ist für Organisationen auf einer Allowlist verfügbar.

The new together-sandbox SDK runs commands and code in isolated runtime environments built from Docker-image snapshots. It ships as a Python SDK, a TypeScript SDK, and a standalone CLI, and is available to organizations on an allowlist (contact us to request access).

What's changed from the legacy SDK (@codesandbox/sdk):

  • Together-native authentication: Clients authenticate with your Together API key (TOGETHER_API_KEY) instead of a CodeSandbox API token.
  • Python support: The legacy SDK was TypeScript-only. The new SDK is published on both PyPI and npm, and the CLI installs as a self-contained binary.
  • Docker-defined environments: Sandboxes boot from snapshots built from a Docker image or Dockerfile by Together's remote image builder, replacing templates built with the CodeSandbox CLI. No local Docker is required.
  • Snapshot-based persistence: Sandboxes are ephemeral by default and termination is permanent. To maintain state, snapshot the filesystem on termination and start a new sandbox from it. This replaces the legacy hibernate and resume model.

See Code sandbox for the new workflow. </Update>

<Update label="September 29, 2026" tags={["Improvements"]}>

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Kein Datum in der Quelle. Angezeigt ist der Tag, an dem wir den Eintrag aufgenommen haben.

Aufgenommen am .

Together AI

Together Link als Beta für macOS und Linux

Together Link ist als Beta für macOS und Linux verfügbar und startet sechs Coding-Agenten (Claude Code, Codex, OpenCode, Pi Code, Claude Desktop, ChatGPT Desktop) per Installationsbefehl auf bei Together AI gehosteten Modellen, ohne die normale Agent-Konfiguration zu verändern.

Together Link runs six coding agents on models hosted by Together AI: Claude Code, Codex, OpenCode, and Pi Code in the terminal, plus Claude Desktop (including Cowork) and ChatGPT Desktop. Install it with one command, launch your agent through it, and your normal agent configuration stays untouched. It's now in beta on macOS and Linux.

# Install Together Link
curl -fsSL https://link.together.ai/install | bash

# Open the interactive launcher
togetherlink

# Or launch an agent directly, pinned to one model
togetherlink --main moonshotai/Kimi-K3 claude

What's included:

  • Six agents: Launch Claude Code (tclaude), Codex (tcodex), OpenCode (topencode), or Pi Code (tpi) in your terminal, or switch Claude Desktop and ChatGPT Desktop to a reversible Together Link profile. OpenCode requires OpenCode 2, and Pi Code requires version 0.80.8 or newer.
  • Auto router: Sessions default to the auto model, which picks a Together AI model for each request. In Claude Code and Claude Desktop sessions with an Anthropic API key, it sends the most difficult requests to Claude Opus.
  • Models: moonshotai/Kimi-K3, zai-org/GLM-5.3, zai-org/GLM-5.3-Flash, and deepseek-ai/DeepSeek-V4.1-Flash, all with 1M context, billed at standard serverless rates. Pin one with --main, or switch with your agent's /model command. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Aufgenommen am .

Gemini Enterprise von Google

Admin-Steuerung für Antigravity-Funktionen Boost und Teamwork in Gemini Enterprise

Administratoren in Gemini Enterprise können nun steuern, ob Nutzer die Antigravity-Funktionen Boost (GA, per /boost) und Teamwork (Preview, per /teamwork-preview) verwenden dürfen.

Feature

Gemini Enterprise: Control access to Antigravity features

Administrators can control whether users in their organization have access to the following Google Antigravity features:

  • Boost: Uses a tiered multi-agent hierarchy to solve complex algorithmic challenges and deep debugging tasks with the /boost command. This feature is generally available (GA) in Antigravity.
  • Teamwork: Deploys a coordinated framework of autonomous subagents to execute large-scale, long-horizon projects using the /teamwork-preview command. This feature is in Preview in Antigravity.

These settings are generally available (GA) for administrators in Gemini Enterprise.

For more information, see Configure feature settings.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Kein Datum in der Quelle. Angezeigt ist der Tag, an dem wir den Eintrag aufgenommen haben.

Aufgenommen am .

Fireworks AI

Evaluator-Verbesserungen, Kimi K2 Thinking serverless und neue API-Endpunkte

Evaluatoren lassen sich über GitHub-Vorlagen erstellen und in einer sortierbaren Tabelle anzeigen, Kimi K2 Thinking ist serverless verfügbar, KAT Dev 32B und 72B Exp sind neue Modelle, es gibt Dokumentation zu W&B und MLflow, und neue REST-API-Endpunkte verwalten Reinforcement Fine-Tuning Steps und Deployments.

Improved Evaluator Creation Experience

The evaluator creation workflow has been significantly enhanced with GitHub template integration. You can now:

  • Fork evaluator templates directly from GitHub repositories
  • Browse and preview templates before using them
  • Create evaluators with a streamlined save dialog
  • View evaluators in a new sortable and paginated table

MLOps & Observability Integrations

New documentation for integrating Fireworks with MLOps and observability tools:

  • Weights & Biases (W&B) integration for experiment tracking during fine-tuning
  • MLflow integration for model management and experiment logging

✨ New Models

☁️ Serverless

📚 New REST API Endpoints

New REST API endpoints are now available for managing Reinforcement Fine-Tuning Steps and deployments:

  • Create Reinforcement Fine-Tuning Step …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Kein Datum in der Quelle. Angezeigt ist der Tag, an dem wir den Eintrag aufgenommen haben.

Aufgenommen am .

Fireworks AI

Audit Logs, Dataset-Download und gewichtetes Training für Reinforcement Fine-Tuning

Die Web-App bietet nun eine Audit-Logs-Seite mit Suche und Filter sowie den Download von Datasets, Reinforcement Fine-Tuning unterstützt Gewichtung pro Beispiel, KAT Coder ist neu in der Model Library und die Konsolen-Seitenleiste wurde neu gegliedert.

Audit Logs in Web App

You can now view and search audit logs directly from the Fireworks web app. The new Audit Logs page provides:

  • Search and filter logs by status and timeframe
  • Detailed view panel for individual log entries
  • Easy navigation from the console sidebar under Account settings

See the Audit Logs documentation for more information.

Dataset Download

You can now download datasets directly from the Fireworks web app. The new download functionality allows you to:

  • Download individual files from a dataset
  • Download all files at once with "Download All"
  • Access downloads from the Datasets table in the dashboard

Weighted Training for Reinforcement Fine-Tuning

Reinforcement Fine-Tuning now supports per-example weighting, giving you more control over which samples have greater influence during training. This feature mirrors the weighted training functionality already available in Supervised Fine-Tuning.

See the Weighted Training documentation for details on the weight field format.

✨ New Models

  • KAT Coder is now available in the Model Library
<Accordion title="Bug Fixes & Minor Improvements"> - **Console Navigation:** Redesigned sidebar with organized groups (CREATE, EXPLORE, MANAGE) for easier navigation (Web App) …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Kein Datum in der Quelle. Angezeigt ist der Tag, an dem wir den Eintrag aufgenommen haben.

Aufgenommen am .

Fireworks AI

DeepSeek V3.2 serverless, Cached-Token-Preise und neue Modelle

DeepSeek V3.2 ist serverless verfügbar, Model Library und Modellseiten zeigen Preise für gecachte und nicht gecachte Input-Tokens, das Evaluations-Dashboard erhält Statusspalte und Schnellfilter, und DeepSeek V3.2 sowie drei Ministral-3-Modelle (14B, 8B, 3B Instruct 2512) sind in der Model Library verfügbar.

☁️ Serverless

Cached Token Pricing Display

The Model Library and model detail pages now display cached and uncached input token pricing for serverless models that support prompt caching. This gives you better visibility into potential cost savings when using prompt caching with supported models.

Evaluations Dashboard Improvements

The Evaluations dashboard has been enhanced with new filtering and status tracking capabilities:

  • Status column showing evaluator build state (Active, Building, Failed)
  • Quick filters to filter evaluators and evaluation jobs by status
  • Improved table layout with actions integrated into the status column

✨ New Models

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Kein Datum in der Quelle. Angezeigt ist der Tag, an dem wir den Eintrag aufgenommen haben.

Aufgenommen am .

Fireworks AI

Reasoning-Guide, Prompt-Caching-Updates und neue Modelle

Es gibt einen neuen Reasoning-Guide, gecachte Prompt-Tokens kosten auf Serverless 50 % weniger, Session-Affinity-Routing ist per user-Feld oder x-session-affinity-Header möglich, und Devstral Small 2 24B Instruct 2512 sowie NVIDIA Nemotron Nano 3 30B A3B sind in der Model Library verfügbar.

Reasoning Guide

A new Reasoning guide is now available in the documentation. This comprehensive guide covers:

  • Accessing reasoning_content from thinking/reasoning models
  • Controlling reasoning effort with the reasoning_effort parameter
  • Streaming with reasoning content
  • Interleaved thinking for multi-step tool-calling workflows

The guide provides code examples using the Fireworks Python SDK and explains how to work with models that support extended reasoning capabilities.

Prompt Caching Updates

Prompt caching documentation has been updated with expanded guidance:

  • Cached prompt tokens on serverless now cost 50% less than uncached tokens
  • Session affinity routing via the user field or x-session-affinity header for improved cache hit rates
  • Prompt optimization techniques for maximizing cache efficiency

See the Prompt Caching guide for details.

✨ New Models

📚 Documentation Updates …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Kein Datum in der Quelle. Angezeigt ist der Tag, an dem wir den Eintrag aufgenommen haben.

Aufgenommen am .

Fireworks AI

Playground-Kategorien, neue Nutzerrollen und Fine-Tuning-Verbesserungen

Der Playground erhält Kategorie-Tabs (LLM, Image, TTS, STT), es gibt die neuen Rollen Contributor und Inference, und Fine-Tuning-Jobs lassen sich stoppen, fortsetzen und klonen, außerdem können Output-Datasets und Rollout-Logs von RFT-Jobs heruntergeladen werden.

Playground Categories

The Playground now features category tabs (LLM, Image, TTS, STT) in the header for easier switching between model types. The playground automatically detects the appropriate category based on the selected model and provides smart defaults for each category.

User Roles: Contributor and Inference

New user roles provide more granular access control for team collaboration:

  • Contributor: Read and write access to resources without administrative privileges
  • Inference: Read-only access with the ability to run inference on deployments

Assign these roles when inviting team members to provide appropriate access levels.

Fine-Tuning Improvements

Fine-tuning workflows have been enhanced with several new capabilities:

  • Stop and Resume Jobs: Stop running fine-tuning jobs and resume them later from where they left off. Available for Supervised Fine-Tuning and Reinforcement Fine-Tuning jobs.
  • Clone Jobs: Quickly create new fine-tuning jobs based on existing job configurations using the Clone action.
  • Download Output Datasets: Download output datasets from Reinforcement Fine-Tuning jobs, including individual files or bulk download as a ZIP archive.
  • Download Rollout Logs: Download rollout logs from Reinforcement Fine-Tuning jobs for offline analysis.

✨ New Models …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Kein Datum in der Quelle. Angezeigt ist der Tag, an dem wir den Eintrag aufgenommen haben.

Aufgenommen am .

Fireworks AI

Warm-Start-Training für RFT und Azure-Federated-Identity für Modell-Uploads

Reinforcement-Fine-Tuning-Jobs lassen sich per --warm-start-from aus SFT-Checkpoints starten, und Modell-Uploads aus Azure Blob Storage unterstützen alternativ zu SAS-Tokens die Azure-AD-Federated-Identity-Authentifizierung.

Warm-Start Training for Reinforcement Fine-Tuning

You can now warm-start Reinforcement Fine-Tuning jobs from previously supervised fine-tuned checkpoints using the --warm-start-from flag. This enables a streamlined SFT-to-RFT workflow where you first train a model with supervised fine-tuning, then continue training with reinforcement learning.

See the Warm-Start Training guide for details.

Azure Federated Identity for Model Uploads

Model uploads from Azure Blob Storage now support Azure AD federated identity authentication as an alternative to SAS tokens. This eliminates the need for credential rotation and enables secure, credential-less authentication.

See the Uploading Custom Models documentation for setup instructions.

📚 Documentation Updates

  • Warm-Start Training: New guide for SFT-to-RFT workflows (Warm-Start Training)
  • Azure Federated Identity: Setup instructions for Azure AD authentication (Uploading Custom Models)
  • Preserved Thinking: Multi-turn reasoning with preserved thinking context (Reasoning)
  • GLM 4.7: Added to models supporting reasoning_effort parameter

<Accordion title="Bug Fixes & Minor Improvements"> …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Kein Datum in der Quelle. Angezeigt ist der Tag, an dem wir den Eintrag aufgenommen haben.

Aufgenommen am .

Fireworks AI

Video- und Audio-Modelle, AWS-S3-Training und JIT-Provisioning für SSO

Multimodale Modelle wie Qwen3 Omni und Molmo2 verarbeiten nun Video- und Audio-Eingaben über die Chat Completions API, Trainingsdatasets können in eigenen AWS-S3-Buckets liegen, und Enterprise-SSO unterstützt Just-In-Time-Nutzerprovisionierung.

Video & Audio Input Models

You can now query multimodal models with video and audio inputs for video captioning, scene analysis, and multimodal question answering. Deploy models like Qwen3 Omni and Molmo2 to process video and audio content directly using the Chat Completions API.

See the Video & Audio Inputs guide for deployment instructions and code examples.

AWS S3 Integration for Training Datasets

Training datasets can now be stored in your own AWS S3 buckets using GCP-to-AWS OIDC federation. This Bring Your Own Bucket (BYOB) approach keeps your data private while enabling secure access during Supervised Fine-Tuning and Reinforcement Fine-Tuning jobs—no long-lived credentials required.

See the Secure Training (BYOB) documentation for IAM role setup and usage examples.

Just-In-Time (JIT) User Provisioning for SSO (Enterprise)

JIT user provisioning automatically creates user accounts when users sign in through SSO for the first time. Enable this when configuring your identity provider to eliminate manual user creation.

See the SSO documentation for setup instructions.

📚 Documentation Updates

  • Video & Audio Inputs: New guide for processing video and audio with Qwen3 Omni and Molmo2 models (Video & Audio Inputs) …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Kein Datum in der Quelle. Angezeigt ist der Tag, an dem wir den Eintrag aufgenommen haben.

Aufgenommen am .

Fireworks AI

Audio-Inferenz und Bildgenerierung abgekündigt

Audio-Inferenz und Bildgenerierung werden abgekündigt.

Audio inference and image generation are deprecated. </Update>

<Update label="2026-05-14"> <Badge color="blue">Inference</Badge>

Originalquelle(öffnet in neuem Tab)Problem melden