Zum Inhalt springen

Cartesia Release Notes

9 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Angaben zum Datum

Datum aus dem Text des Eintrags.

Erstmals gesehen am .

Cartesia

Cartesia März 2026: Breaking Changes, strukturierte API-Fehler

Der Text-to-Agent-Workflow (T2A) für Line ist veraltet, die API liefert ab Cartesia-Version: 2026-03-01 strukturierte JSON-Fehler, PVC-Stimmen benötigen eine datierte Modell-ID, und es gibt Verbesserungen bei Playground, Sprachsuche und Billing-Benachrichtigungen zur Concurrency.

Breaking

  • Text-to-Agent (T2A) API — Text-to-Agent workflow for Line is deprecated.

API

  • Error responses — For Cartesia-Version: 2026-03-01, we now return structured JSON. See API Errors.
    • API versions before 2026-03-01 continue to return legacy error formats (for example HTTP Title: Message).
    • Voices — PATCH /voices/{id}: voice owners can now update accent and gender. Voice creation validates language. Invalid voice UUIDs and pronunciation-dictionary IDs return 404 instead of ambiguous errors.
  • PVC model routing — PVC voices require a dated model ID (e.g. sonic-3-2026-01-12) instead of sonic-3. See Pro Voice Clone.
  • Voice search — Name and metadata search is diacritics-insensitive.

Playground

  • Pro voice clones
    • Clearer language mismatch messaging
    • Background noise removal is now a simple on/off control
    • Fine-tuning model support:
      • Removed support for older models
      • Now only sonic-3-2026-01-12 is supported
  • Multilingual agents — Multilingual agent configuration is now supported in the Playground.
  • Agents UI — Search by call ID and agent ID.

Billing

  • Concurrency — Organizations can receive notifications when concurrency nears configured limits.

Model / voice …

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Cartesia

Cartesia September 2026: Ink-2 mehrsprachig, Sonic 3.6 self-hosted

Ink-2 unterstützt jetzt zusätzlich Spanisch, Französisch, Hindi und Japanisch, Pronunciation Dictionaries auf Sonic 3.6 ignorieren standardmäßig die Groß-/Kleinschreibung, die Aussprache englischer Wörter in japanischen und chinesischen Texten wurde verbessert und Sonic 3.6 ist für Self-Hosted-Kunden verfügbar.

Speech-to-Text

  • Ink-2 is now multilingual — Ink 2 supports four new languages: Spanish, French, Hindi and Japanese with industry leading transcription accuracy and semantic endpointing. Try it in the playground or through the API.

Text-to-Speech

  • More control over pronunciation — Pronunciation dictionaries on Sonic 3.6 now ignore case by default, so "new york," "New York," and "NEW YORK" share one entry. For terms where capitalization changes the meaning, turn case sensitivity on for that entry in the Playground or API. See Case sensitivity.
  • Cleaner English inside Japanese and Chinese text — We improved spacing around Latin-script words in Japanese and Chinese transcripts, so Sonic 3.6 reads embedded English correctly.
  • Sonic 3.6 for self-hosted deployments — The September 17 release brings self-hosted customers Sonic 3.6's quality improvements, Odia and Urdu support, and locale-aware normalization. It also cuts stalls during voice updates. If an air-gapped license proxy goes down, you now have 24 hours to fix it instead of 5 minutes. See release details.

Voice Agents …

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Cartesia

Sonic 3.6 – Version 3.6

Sonic 3.6 ist ab dem 27. August als sonic-3.6 allgemein verfügbar und bietet natürlichere Sprache, bessere Mehrsprachigkeit in 44 Sprachen bzw. 61 Locales, die neuen Sprachen Odia und Urdu, ein neues normalization-Feld sowie bessere Aussprache von Alphanumerik, Disfluencies und englischen Heteronymen.

Sonic 3.6

Sonic 3.6 is an update to Sonic 3.5 — generally available from August 27 on sonic-3.6.

What's new

  • More natural speech, pacing, and emotional expression — the model adapts intonation, pacing, and emotiveness to the context of the transcript, with no SSML tags or explicit instructions. In blind head-to-head testing, listeners preferred Sonic 3.6 to Sonic 3.5 by nearly two to one on English transcripts.
  • Step-change multilingual naturalness — native-sounding pronunciation and accents across 44 languages / 61 locales.
  • Two new languages — Odia (or) and Urdu (ur), including voices and IVC support.
  • Hindi and Hinglish transcript following — romanized and Latin-script Hindi follows the transcript accurately. Set text normalization independently of the spoken language with the new normalization field — see Normalization.
  • Non-English alphanumerics — spelled-out letters, codes, and IDs use native letter names and accent, in all supported languages.
  • In-transcript disfluencies — written disfluencies ("It's on, uh, Fifth Street") now shift pacing and intonation to sound like natural hesitation.
  • English heteronyms — higher pronunciation accuracy on heteronyms like "read," "bass," and "live" in context.

How to try it …

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Cartesia

Cartesia Juli 2026: Keyterm Prompting und Zugriffskontrolle

Speech-to-Text unterstützt Keyterm Prompting und einen verbesserten Playground, Stimmen und Pronunciation Dictionaries lassen sich per Zugriffskontrolle verwalten, und Voice Agents erhalten Live-Events über die Agents WebSocket API.

Speech-to-Text

  • Keyterm prompting for better accuracy — Set domain-specific terms, brand or product names, and rare or invented words so the model transcribes them correctly. See the keyterms guide, and try it out on the Playground or API.
  • Improved Playground features — Watch transcription in real time on a sample clip or your own audio. Configure keyterms and adjust turn detection sensitivity right in the Speech-to-Text Playground.

Text-to-Speech

  • Access control for voices and pronunciation dictionaries — Manage access to your custom voices and pronunciation dictionaries via the Playground or API. Set access to public to use them on third-party platforms.

Voice Agents

  • Live agent experiences via API — Power live transcripts and keep your systems in sync with calls in real time via turn events, word-level assistant text, interruption state, tool calls, and call IDs. See the Agents WebSocket API. …

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Cartesia

Cartesia Juni 2026: Turn Detection, Wissensbasen und Batch-Anrufe

Neu sind einstellbare Schwellenwerte für die Turn Detection, Professional Voice Clones auf Sonic 3.5, frei wählbare Sampleraten im TTS Playground sowie für Voice Agents Wissensbasen, Batch-Outbound-Anrufe mit bis zu 5.000 Anrufen pro Anfrage und Zero Data Retention für Enterprise-Kunden.

Speech-to-Text

  • Turn detection controls — Adjust turn-start, turn-end, and eager-end thresholds to balance response speed against detection accuracy. See the Turn Detection guide for more details.

Text-to-Speech

  • Professional Voice Clones now available on Sonic 3.5 — They deliver better speaker similarity and more stable generation than on Sonic 3, especially for rare or non-native accents.
  • Test sample rates — Set any sample rate from 8 to 44.8 kHz based on your intended use case on the TTS Playground.

Voice Agents

  • Upload knowledge bases — Let agents access domain-specific information via docs including FAQs, pricing, and guides. Set it up with the Knowledge Base guide.
  • Batch outbound calling — Send up to 5,000 calls with a single API request. Control how many run at once, track status, retry failed calls, and schedule batches for later. See the batch calling guide.
  • Zero data retention — Enable ZDR so transcripts, audio recordings, and logs from agent calls are never stored. Now available for Enterprise customers. …

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Cartesia

Cartesia Mai 2026: Ink-2 und Sonic 3.5 allgemein verfügbar

Cartesia führt mit Ink-2 ein Streaming-STT-Modell mit integrierter Turn Detection ein (zunächst nur Englisch), und Sonic 3.5 ist nun allgemein verfügbar und produktionsreif über den Alias sonic-3.5.

Speech-to-Text

  • Ink-2, our state-of-the-art streaming STT model — Build responsive real-time voice experiences with built-in turn detection and accurate transcription even in noisy environments. It currently supports only English, with additional languages coming later.

Text-to-Speech

  • Sonic 3.5 is now generally available — Our most natural, expressive TTS model is out of preview and production-ready. Use the sonic-3.5 alias for the latest stable snapshot. See the Sonic 3.5 model overview.
    • Switching from Sonic 3? See Migrating from Sonic 3 to Sonic 3.5 for what's new and what to check before moving production traffic. …

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Cartesia

Sonic 3.5 – Version 3.5

Sonic 3.5 ist über sonic-3-latest verfügbar und bietet natürlichere Sprache, sauberere Audioqualität, bessere Ausgabe von Alphanumerik, verbesserte Mehrsprachigkeit und korrekte Aussprache englischer Heteronyme, wobei der -latest-Alias nicht für Produktivbetrieb empfohlen wird.

Sonic 3.5

Sonic 3.5 is now available on sonic-3-latest. We'd love for you to try it and tell us what you think.

Why you should try it

  • More natural speech, pacing, and emotional expression, especially noticeable on expressive, conversational, and support-style transcripts.
  • Cleaner audio quality across all languages and voices.
  • Better alphanumeric read-out — confirmation codes, order numbers, phone numbers, IDs, and emails sound meaningfully more natural, in all supported languages.
  • Step-change multilingual performance, particularly Hebrew, Japanese, Spanish, Hindi, German, Korean, and French.
  • English heteronyms — tricky English heteronyms like "read," "bass," and "bow" now pronounce correctly in context.

How to try it

  1. Point your API call or Playground request to the model ID sonic-3-latest.
  2. Keep your existing voice IDs, request shape, and prompting — no code changes required for most customers.
  3. Send us feedback on any voice or transcript that behaves differently than you expect.
<Note> As with any `-latest` alias, `sonic-3-latest` can be updated without notice and is not recommended for production. Pin to a dated snapshot (e.g. `sonic-3`) for production traffic. </Note>

What to know to be successful …

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Cartesia

Cartesia Februar 2026: Line-Neuerungen und bessere Aussprache

Line erhält eine History Management API, Custom User Events und unterbrechungsfreie Nachrichten, außerdem gibt es cartesia-python v3.0.0, neue Playground-Seiten und eine verbesserte Aussprache von Zahlen, Datumsangaben, Währungen und Maßeinheiten.

Line

  • History Management API: You can add or replace the history provided to your agent, for example, to summarize a long conversation.
  • Custom User Events: You can send bidirectional custom events between your client and the agent. You could use this, for example, if you have a web application with UI interactions.
  • Uninterruptible Messages: You can set messages as uninterruptible. A common use case is a legal disclaimer at the beginning of a call.
  • End Tool Call Improvements: The default end call tool call is more conservative to prevent calls from ending prematurely.

API

  • Increased reliability of API connections

Cartesia SDK

Playground

  • Shipped a new TTS page
  • Shipped a new Voice Creation page
  • Shipped a new Agents page

Model changes

  • Improved pronunciation of real-world text patterns across languages
    • Enhanced support for structured and formatted speech patterns: numbers, dates, times, currency, phone numbers, IDs, percentages, and amounts/measurements.
    • Support for various date formats (YYYY-MM-DD, YYYY/MM/DD, 年月日).
    • Support for measurement units (meters, kg, tablespoon, gigabytes, etc.) with locale awareness. …

Originalquelle(öffnet in neuem Tab)Problem melden

Datum unbekanntAngaben zum Datum

Kein Datum in der Quelle. Der Eintrag stammt aus dem ersten Abruf der Quelle, der Tag der Aufnahme sagt nichts über das Erscheinen.

Erstmals gesehen am .

Cartesia

Cartesia Januar 2026: Regionalisierung und Sonic-3-Versionierung

Anrufe werden nach Herkunft in die Regionen US, EU oder APAC geleitet, es gibt parametrisierte ausgehende Anrufe und Pronunciation Dictionaries, ein neues Versionierungsschema für Sonic-3 mit dem stabilen Checkpoint sonic-3-2026-01-12, Featured Voices und neue Stimmen in der Voice Library.

API

  • Regionalization — Calls routed to US, EU, APAC by origin.
  • Parameterized outbound calls — Docs
  • Pronunciation dictionaries — Docs

Model changes

  • Sonic-3 model versioning scheme introduced
    • New preview track: sonic-3-latest (continuous updates for early access and feedback).
    • Stable track: sonic-3 always points to the most recent stable release.
    • Immutable dated snapshots: sonic-3-YYYY-MM-DD never change.
    • Details: Continuous updates and model snapshots
  • Promotion to stable checkpoint: sonic-3-2026-01-12
    • Included improvements: consistent speed & volume, custom IPA pronunciations with stronger adherence, Hindi prosody improvements, Korean prosody/intonation improvements.

Voice changes

  • Featured Voices launched — Curated set of 30+ best-performing voices (e.g. Cathy, Henry).
  • Voice Library — December: 25 new voices across 6 languages.
  • Voice Library — January: 9 Spanish voices (Mexican, Colombian, Castilian).

Playground

  • Voice library usability improvements (test with your own scripts, call an agent per voice). …

Originalquelle(öffnet in neuem Tab)Problem melden