Zum Inhalt springen

Together AI Release Notes

10 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Together AI

Batch-API-Dateien werden jetzt 7 Tage aufbewahrt

Eingabe-, Ausgabe- und Fehlerdateien von Batch-Jobs werden jetzt 7 Tage aufbewahrt, danach sind sie nicht mehr abrufbar und eine hochgeladene Eingabedatei lässt sich nur in diesem Zeitraum für weitere Jobs wiederverwenden.

The input file you upload for a batch job, along with the output and error files the job produces, are now retained for 7 days. After that the files are no longer accessible, so download your results before the window closes. Reusing an uploaded input file across batch jobs also works only within that window.

See Batch inference. </Update>

<Update label="September 24, 2026" tags={["Improvements"]}>

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus dem Text des Eintrags.

Erstmals gesehen am .

Together AI

Neues serverloses Modell: Tev1-4B-experimental

Das Modell together/Tev1-4B-experimental mit 32.768 Kontextlänge ist jetzt serverless verfügbar, zu 0,042 $ Input und kostenlosem Output pro 1M Tokens.

The following models are now available on serverless:

  • together/Tev1-4B-experimental: 32,768 context length. Pricing: $0.042 input / free output (per 1M tokens).</Update>

<Update label="September 22, 2026" tags={["Pricing"]}>

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Together AI

Together Link als Beta für macOS und Linux

Together Link ist als Beta für macOS und Linux verfügbar und startet sechs Coding-Agenten (Claude Code, Codex, OpenCode, Pi Code, Claude Desktop, ChatGPT Desktop) per Installationsbefehl auf bei Together AI gehosteten Modellen, ohne die normale Agent-Konfiguration zu verändern.

Together Link runs six coding agents on models hosted by Together AI: Claude Code, Codex, OpenCode, and Pi Code in the terminal, plus Claude Desktop (including Cowork) and ChatGPT Desktop. Install it with one command, launch your agent through it, and your normal agent configuration stays untouched. It's now in beta on macOS and Linux.

# Install Together Link
curl -fsSL https://link.together.ai/install | bash

# Open the interactive launcher
togetherlink

# Or launch an agent directly, pinned to one model
togetherlink --main moonshotai/Kimi-K3 claude

What's included:

  • Six agents: Launch Claude Code (tclaude), Codex (tcodex), OpenCode (topencode), or Pi Code (tpi) in your terminal, or switch Claude Desktop and ChatGPT Desktop to a reversible Together Link profile. OpenCode requires OpenCode 2, and Pi Code requires version 0.80.8 or newer.
  • Auto router: Sessions default to the auto model, which picks a Together AI model for each request. In Claude Code and Claude Desktop sessions with an Anthropic API key, it sends the most difficult requests to Claude Opus.
  • Models: moonshotai/Kimi-K3, zai-org/GLM-5.3, zai-org/GLM-5.3-Flash, and deepseek-ai/DeepSeek-V4.1-Flash, all with 1M context, billed at standard serverless rates. Pin one with --main, or switch with your agent's /model command. …

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Together AI

Neues Code-Sandbox-SDK und CLI: together-sandbox

Das neue together-sandbox SDK (Python, TypeScript und CLI) führt Befehle und Code in isolierten, aus Docker-Image-Snapshots erzeugten Umgebungen aus, authentifiziert sich mit dem Together API Key, setzt auf Snapshots statt Hibernate/Resume und ist für Organisationen auf einer Allowlist verfügbar.

The new together-sandbox SDK runs commands and code in isolated runtime environments built from Docker-image snapshots. It ships as a Python SDK, a TypeScript SDK, and a standalone CLI, and is available to organizations on an allowlist (contact us to request access).

What's changed from the legacy SDK (@codesandbox/sdk):

  • Together-native authentication: Clients authenticate with your Together API key (TOGETHER_API_KEY) instead of a CodeSandbox API token.
  • Python support: The legacy SDK was TypeScript-only. The new SDK is published on both PyPI and npm, and the CLI installs as a self-contained binary.
  • Docker-defined environments: Sandboxes boot from snapshots built from a Docker image or Dockerfile by Together's remote image builder, replacing templates built with the CodeSandbox CLI. No local Docker is required.
  • Snapshot-based persistence: Sandboxes are ephemeral by default and termination is permanent. To maintain state, snapshot the filesystem on termination and start a new sandbox from it. This replaces the legacy hibernate and resume model.

See Code sandbox for the new workflow. </Update>

<Update label="September 29, 2026" tags={["Improvements"]}>

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Together AI

Höheres LoRA-Rang-Limit von 128 beim Fine-Tuning

LoRA-Adapter lassen sich für die meisten Modelle nun mit einem Rang von bis zu 128 (zuvor 64) trainieren, wobei der Standardrang 64 bleibt und frühere CLI- und SDK-Versionen ohne gesetztes lora_r jetzt den Maximalrang 128 verwenden.

You can now train LoRA adapters with a rank of up to 128 for the majority of models, up from 64. The default rank for these models stays at 64, so set lora_r to use a higher one.

The model limits response has a new lora_training.default_rank field next to lora_training.max_rank. Run tg fine-tuning model-limits <model> to see both values for a model.

Version 2.36.0 of the Together CLI and Python SDK uses the default rank when you don't set lora_r. Earlier versions, including the 1.x SDK, use the model's maximum rank instead, which is now 128 on these models. Upgrade to 2.36.0 or set lora_r yourself, especially if you continue training from a rank-64 adapter, where the rank has to match.

In the console, the rank field now starts at the model's default rank (64 on most models) instead of 8.

See Supported models for each model's default and maximum rank.

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Together AI

Längerer Fine-Tuning-Kontext für Qwen-27B-Modelle

Die Qwen-27B-Modelle unterstützen beim Fine-Tuning nun 131.072 Token Kontext für SFT und 65.536 für DPO, dafür sinkt die Batch-Größe bei LoRA-Jobs auf 2 und bei DPO mit Full Fine-Tuning auf maximal 8.

Qwen/Qwen3.8-27B, Qwen/Qwen3.6-27B, and Qwen/Qwen3.5-27B now support a 131,072-token context for SFT (up from 32,768) and 65,536 for DPO (up from 16,384), for both LoRA and full fine-tuning. Batch size limits have also changed: LoRA jobs on these models run at a batch size of 2, down from 16, and the maximum DPO batch size for full fine-tuning drops from 16 to 8.

See Supported models for each model's limits.

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Together AI

Fine-Tuned-Modell per Registry-Namen bereitstellen

Seit Together CLI Version 2.24.0 akzeptiert tg beta endpoints deploy anstelle der model_object_id auch den Registry-Namen model_object_name eines abgeschlossenen Fine-Tuning-Jobs.

Since Together CLI version 2.24.0, tg beta endpoints deploy accepts a completed fine-tuning job's model_object_name, the qualified <project_slug>/<model_name> registry name, in place of its model_object_id. The CLI resolves the name to the same model, so you can deploy straight from the name shown in the fine-tuning jobs dashboard. The SDK and API take model_object_id.

See Deploy a fine-tuned model. </Update>

<Update label="September 23, 2026" tags={["New models"]}>

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Together AI

Preissenkung für Qwen3.7-Max und Qwen3.8-Flash

Ab dem 22. September 2026 sinken die Preise für Qwen/Qwen3.7-Max (1,50 $ Input / 4,50 $ Output) und Qwen/Qwen3.8-Flash (0,09 $ Input / 0,282 $ Output) pro 1M Tokens.

The following models have lower pricing, effective September 22, 2026. All usage from that date forward is billed at the new rates (per 1M tokens):

  • Qwen/Qwen3.7-Max: $2.50 → $1.50 (input), $7.50 → $4.50 (output).
  • Qwen/Qwen3.8-Flash: $0.15 → $0.09 (input), $0.47 → $0.282 (output).

See Serverless models for the full pricing catalog. </Update>

<Update label="September 16, 2026" tags={["New releases"]}>

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Together AI

Automatische Leerlauf-Abschaltung für dedizierte Deployments

Dedizierte Deployments können mit --inactive-timeout (bzw. inactiveTimeout in der Management-API) nach der eingestellten Zeit ohne Inferenz-Anfragen automatisch auf null Replikate skalieren, wodurch Hardware freigegeben und die Abrechnung gestoppt wird.

Deployments can now stop themselves when they go unused. Set an inactivity timeout with --inactive-timeout (the inactiveTimeout field in the management API), and if the deployment serves no inference requests for that many minutes, it scales to zero replicas, releasing its hardware and stopping billing.

See Automatic idle shutdown. </Update>

<Update label="September 15, 2026" tags={["New releases", "Deprecations"]}>

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Together AI

Rollouts für dedizierte Modell-Inferenz

Mit Rollouts lässt sich Live-Traffic per Canary-, Blue-Green- oder Rolling-Strategie ohne Änderung der Endpoint-URL von einem Deployment auf ein anderes verlagern, optional mit Metrik-Gates, und über tg beta endpoints rollout oder die Konsole steuern.

Rollouts shift live traffic from one deployment to another under the same endpoint, without changing the endpoint URL. Pick a canary, blue-green, or rolling strategy to determine how traffic moves, and optionally gate a canary rollout on live metrics so it pauses automatically if the new deployment regresses.

Start a rollout with the tg beta endpoints rollout CLI command or from the endpoint's Rollouts tab in the console, then pause, resume, promote, or cancel it at any point while it runs.

Originalquelle(öffnet in neuem Tab)Problem melden