Zum Inhalt springen

AI Voice and Speech: Release Notes

Spracherkennung, Sprachausgabe und Stimmen per KI. 9 Hersteller, 158 Einträge.

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Procedures werden automatisch kompiliert, Merge Proposals in Alpha

Strukturierte Procedures werden beim Speichern und Veröffentlichen von Agenten automatisch kompiliert und validiert, der Compile-Endpunkt gilt als Legacy, und Merge Proposals sind in Alpha verfügbar.

ElevenAgents

  • Structured procedures compile on save and publish: Update agent and Create agent draft now compile and validate structured procedures automatically. Integrations no longer need to call the compile endpoint or manually update the generated workflow. Invalid procedures return per-procedure errors as 400 procedure_validation_failed. Agents without structured procedures are unaffected.
  • Compile procedures endpoint is legacy: Compile procedures is marked legacy and should not be used by new integrations. It remains available for existing callers.
  • Merge proposals: Merge proposals are available in alpha. Teams can open a proposal for an agent branch, review configuration changes and test results, leave comments or approval decisions, and merge approved changes into a target branch. See Merge proposals.

SDK releases

JavaScript SDK …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fish Audio

Drama 3 Preview

Fish Audio bietet mit Drama 3 Preview (Modell-ID drama-3-preview) ein Text-to-Speech-Vorschaumodell, das sich per natürlicher Sprache nach Tonfall, Tempo und Charakter steuern lässt, wobei sich Verhalten und Verfügbarkeit während der Preview ändern können.

Drama 3 Preview

Preview text-to-speech model you direct in plain language. Describe tone, pacing, and character without a fixed set of audio tags.

Use model ID drama-3-preview in the API. Behavior and availability may change while it is in preview.

Text to Speech API

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fish Audio

Transcribe 1 Pro

Fish Audio führt das Speech-to-Text-Modell Transcribe 1 Pro (transcribe-1-pro auf POST /v1/asr) für kurze Clips und lange Aufnahmen ein, dessen Transkripte Sprechermarker, Sprecherwechsel sowie Emotions- und Vokalereignis-Hinweise enthalten können.

Transcribe 1 Pro

Speech-to-text model for short clips and long recordings, including multi-speaker conversations. Transcripts can include inline speaker markers, speaker turns, and emotion and vocal-event cues.

Use model ID transcribe-1-pro on POST /v1/asr

Speech to Text | Models overview

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Fish Audio

Fish Audio S2.1 Pro

Fish Audio S2.1 Pro (s2.1-pro) ist ein Produktions-Text-to-Speech-Modell, das gegenüber S2-Pro bei Qualität, Latenz und Durchsatz verbessert ist und 83 Sprachen unterstützt, während s2.1-pro-free bis zum 30. November 2026 unter Fair-Use-Grenzen und ohne TTFA- und DPA-Garantien kostenlos nutzbar ist.

Fish Audio S2.1 Pro

Production text-to-speech model that improves on S2-Pro in quality, latency, and throughput. Supports 83 languages on one model, with the same [bracket] natural-language control and multi-speaker dialogue as S2.

Use model ID s2.1-pro for production workloads that need TTFA and DPA guarantees. s2.1-pro-free is the same model at no cost through November 30, 2026 for testing, prototyping, development, and smaller businesses, under fair-use limits and without those guarantees.

Read more about S2.1 Pro | Models overview

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Murf

Murf: Gen2 für Streaming API abgeschaltet, Wechsel zu Falcon 2

Das Gen2-Modell wurde für die Streaming API abgeschaltet, sodass Nutzer dort auf das neueste Falcon 2-Modell wechseln müssen, während Gen2 für nicht-streamingbasiertes Text-to-Speech über den Synthesize Speech-Endpunkt weiter verfügbar bleibt.

Gen2 Streaming Decommissioned

The Gen2 model has been decommissioned for the Streaming API. If you haven't migrated yet, use our latest Falcon 2 model on the Streaming API.

The Gen2 model remains available for non-streaming text-to-speech through the Synthesize Speech endpoint.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Hume

Hume stellt TTS- und EVI-APIs am 13. November 2026 ein

Hume stellt die TTS- und EVI-APIs ein: Der Zugriff endet am 13. November 2026 um 12:01 Uhr EST, bis dahin bleiben beide APIs voll unterstützt, und alle Kontodaten werden danach dauerhaft gelöscht.

TTS and EVI API sunset notice [#tts-evi-sunset-10-2-2026]

  • Hume is sunsetting the TTS and EVI APIs. Access ends November 13, 2026 at 12:01 a.m. EST, and both APIs remain fully supported until then. Any data in your account is permanently deleted after November 13, 2026. To compare alternative voice models, see Real World VoiceEQ. For questions, contact Hume support or email assist@hume.ai.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Letterly

Letterly Keyboard 2.7.6: Voice Rewrite per Sprachbefehl

Letterly Keyboard (Version 2.7.6) auf dem iPhone bietet unter Keyboard → Rewrites nun Voice Rewrite, mit dem sich diktierter Text per Sprachbefehl kürzen, im Ton ändern, übersetzen oder korrigieren lässt, inklusive Undo.

Letterly Keyboard Update

1 new feature

We’re continuing to improve Letterly Keyboard on iPhone. Now you can rewrite text by voice, right from the keyboard. You’ll find it in Keyboard → Rewrites.

Voice Rewrite in the keyboard

After dictation, say how you want to change the text. For example, make it shorter, change the tone, translate it, or fix a mistake.

  • If you haven’t edited the text, Voice Rewrite will apply to your last dictation.
  • Want to change a specific part? Select it first, then say how to change it.
  • Changed your mind? Tap Undo.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Eleven v4 und Eleven v4 Turbo sowie Transkriptbearbeitung per Anweisung

Eleven v4 und Eleven v4 Turbo sind verfügbar, dazu kommen Transkriptbearbeitung per Anweisung in Speech to Text sowie neue Agent-Modelle und Testkosten-Angaben in ElevenAgents.

Eleven v4 and Eleven v4 Turbo

Eleven v4 and Eleven v4 Turbo are now available. Eleven v4 provides higher-quality, more expressive speech and improved voice cloning across more than 90 languages. Eleven v4 Turbo provides the same model family for real-time applications, with median inference latency of approximately 100 ms.

Use eleven_v4 through the Text to Dialogue API for content creation and long-form audio. Use eleven_v4_turbo through the Text to Dialogue WebSocket for agents and interactive applications.

Speech to Text

  • Transcript editing: Batch and realtime transcription now accept a natural-language edit instruction of up to 2,000 characters. Batch responses return the result in edited_transcript, while realtime sessions emit an edited_transcript event. See the batch and realtime guides for incompatibilities, response handling and pricing.

ElevenAgents

  • Agent models: Added gpt-6-sol, gpt-6-luna, glm-52, deepseek-v41-flash, claude-opus-5 and claude-opus-5-5 to the supported agent LLM options.
  • Agent test usage: Test invocation summaries can include total credits and USD price. Individual test runs can include credits and a charging breakdown.

Flows …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Sora 2, Sora 2 Pro und Seedance 1.5 Pro werden eingestellt

Sora 2 und Sora 2 Pro werden zum 24. September 2026 eingestellt und Seedance 1.5 Pro wird am 11. November 2026 abgeschaltet, Nutzer müssen betroffene Flows auf andere Videomodelle umstellen.

Image & Video

  • Sora 2 and Sora 2 Pro retired: OpenAI is discontinuing the Sora API on September 24, 2026. Both models have been removed from the Image & Video model picker and stop generating on that date. Existing generations remain available in your history. Flows and templates that use a Sora node need to be switched to another video model, such as Gemini Omni 1.1 Flash, before they can run again.
  • Seedance 1.5 Pro deprecated: ByteDance retires Seedance 1.5 Pro on November 11, 2026. The model is no longer offered for new generations and will stop working on that date. Use a Seedance 2.0 model instead.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

ElevenAgents: Parallele Tool-Aufrufe, gpt-6-astra und Slack-Alerting

ElevenAgents unterstützt parallele Tool-Aufrufe, das Modell gpt-6-astra und Slack-Alerting, außerdem kommen die GPT-Image-2.5-Modelle, ein cascade_timeout_seconds-Parameter für Speech Engine und präzisere Turn-Marker bei Text to Speech hinzu.

ElevenAgents

  • Parallel tool calls: Agent prompt configuration now includes enable_parallel_tool_calls (boolean, default true). When enabled, supported models can execute multiple tools in one turn.
  • Agent models: Added gpt-6-astra to the agent LLM options.
  • Alerting: Agent alerting now supports Slack channels alongside PagerDuty and webhook notifiers.
  • Phone number search: Added a cursor-paginated phone number endpoint with filters for provider, outbound support, assigned agent and branch, label and phone number.

Image and video

  • GPT Image 2.5 models: Image generation now supports gpt-image-2.5-flare and gpt-image-2.5-sunburst. Both models accept up to 10 reference images, quality levels through max, 14 fixed aspect ratios plus auto, and 1K, 2K or 4K output.

Speech Engine

  • Cascade timeout: Speech Engine create and update requests now accept cascade_timeout_seconds (number, 2–15, default 4) to control how long ElevenLabs waits for an upstream speech engine before retrying.

Text to Speech

  • Multi-context turn boundaries: The is_final_audio_for_turn marker is now emitted after every buffered byte for that turn, including MP3 and Opus output. Clients can use the marker as an exact turn boundary without switching to PCM.

SDK releases

JavaScript SDK …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Letterly

Letterly für Windows 1.1.7: Start mit dem Computer

Letterly für Windows (Version 1.1.7) startet nun standardmäßig mit dem Computer und wartet im System Tray, was sich unter Preferences → Launch at startup abschalten lässt.

Launch at Startup on Windows

1 new feature

Letterly can now start with your computer, so it's ready the moment you sit down to work.

Launch at startup

Letterly starts together with Windows and waits in the system tray, next to the clock.

On by default. To turn it off: Preferences → Launch at startup.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Wispr Flow

Canto: das erste eigene Sprachmodell von Wispr Flow

Wispr Flow läuft jetzt für alle Nutzer auf Canto, dem ersten selbst entwickelten Sprachmodell, das laut Anbieter bei realen englischen Flow-Aufnahmen die niedrigste Wortfehlerrate der getesteten Modelle erreicht.

Flow now runs on Canto, the first speech model we've built ourselves. It's now rolled out for everyone, with nothing to switch on.

Most speech models are trained and tested on clean audio recorded in quiet rooms with good microphones. That isn't where people use Flow. You're on a train, in an open office, walking between meetings, or talking quietly so you don't disturb anyone. Canto was built and measured against those conditions instead.

On our evaluation of real-world English Flow recordings, 9.8 hours from more than 2,300 speakers, Canto had the lowest word error rate of every model we tested, including models from Google, OpenAI, AssemblyAI and Deepgram.

On a separate set built to stress-test the hardest audio, with background conversation, music, traffic, wind and whispered speech, Canto was the most accurate of the real-time models we tested.

We're already training its successor at more than ten times the scale, with better recognition in difficult conditions and broader language coverage ahead.

Read the full research post here.

‍

September 15, 2026

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Letterly

Letterly für Android 2.4.4: MCP-Anbindung und Abmelden

Letterly für Android (Version 2.4.4) verbindet sich nun per MCP mit Claude, ChatGPT, Cursor und anderen KI-Tools und ermöglicht zudem das Abmelden vom Konto, wobei vorher synchronisiert und dann lokale Notizen und Aufnahmen entfernt werden.

MCP and Log Out on Android

1 new feature · 1 improvement

Android's turn. Your Letterly notes are now available right inside the AI tools you already use. No copying, no switching apps.

MCP Connections

Letterly now connects to Claude, ChatGPT, Cursor, and other AI tools via MCP. You can:

  • ask about your notes
  • write and rewrite text
  • add custom rewrite options

Also in this update

Log out

You can now log out of your Letterly account on Android. Letterly will sync your data first, then remove local notes and recordings from this device. Internet connection required.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Wispr Flow

Wispr Flow Notetaker jetzt unter Windows verfügbar

Wispr Flow Notetaker ist jetzt auch unter Windows verfügbar und bietet Rückblicke während Anrufen, meetingübergreifende Fragen mit Quellenangaben, Entwürfe aus Notizen sowie die Ein-Klick-Anbindung an Gemini neben Claude, ChatGPT und anderen MCP-kompatiblen KI-Tools.

Notetaker is now available on Windows, alongside Mac. Here’s what it does and what we’ve added since the Mac launch.

Capture the details, then work from them

Notetaker detects meetings across the platforms you already use and offers to take notes. It uses your personal dictionary to help capture names, acronyms, and technical terms correctly so you can:

  • Catch up during a call: Ask “What did I miss?” for a recap of the discussion so far.
  • Find answers across meetings: Ask “What follow-ups do I owe people from my calls this week?” or “Based on the calls I’ve led this month, how could I ask better questions?” Answers include citations back to the meetings.
  • Draft from your notes: Turn conversations into follow-up emails, project updates, or briefs directly in Notetaker.

Find it: Open the Notetaker tab in Wispr Flow.

Connect your meetings to your other work

Notetaker just added one-click setup for Gemini, alongside connections to Claude and ChatGPT. You can also connect to any other MCP-compatible AI tool.

Once connected, your AI tool can grab your meeting notes and transcripts along with any other files and tools you’ve connected there. You’ll be able to dictate complex prompts:

  • For sales: “Compare the objections from my discovery calls this quarter with our objection-handling doc. Where are we not answering well?”…

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

ElevenAgents: Call Queueing, gemini-3.8-flash und music_v2_5

ElevenAgents erhält Call Queueing mit Warteschleifen-Audio, Versionsmetadaten für Testergebnisse und das Modell gemini-3.8-flash, und die Music-Endpunkte akzeptieren music_v2_5.

ElevenAgents

  • Call queueing and hold audio: Added call queueing for agents at their concurrency limit. Queued callers hear hold audio and receive queue_status events until they are admitted or time out. New API methods upload or delete custom hold audio.
  • Test version metadata: Agent test results now identify the agent version they ran against and whether they included draft changes. Invocation summaries also indicate when partial resubmissions ran against different versions.
  • Model and account metadata: Agent LLM configuration now accepts gemini-3.8-flash. WhatsApp account responses identify whether the account uses the Cloud API or coexistence signup flow.

Music

  • Music v2.5 API support: Music endpoints and SDKs now accept music_v2_5. Composition plan generation chunks support up to 6,132 characters, with up to 30 lines of 200 characters each.

SDK releases

JavaScript SDK

  • v2.68.0 - Added methods to upload and delete agent hold audio. Regenerated types for call queueing, procedure references, draft branch creation, RAG query limits, test version metadata, gemini-3.8-flash and music_v2_5.

Python SDK …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Scribe v2 Medical für klinische Audioinhalte allgemein verfügbar

Scribe v2 Medical, ein Batch-Spracherkennungsmodell für medizinische und klinische Audioinhalte, ist allgemein verfügbar und wird zum selben Preis wie Scribe v2 abgerechnet (model_id scribe_v2_medical).

Scribe v2 Medical

Scribe v2 Medical is now generally available. It is a batch speech recognition model specialized for medical and clinical audio, billed at the same rate as Scribe v2.

Pass scribe_v2_medical as model_id on Create transcript. Keyterm prompting, entity detection, speaker diarization, and no-verbatim mode work the same way as Scribe v2.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Tts von Inworld

OpenAI-kompatibler Speech-Endpunkt für Realtime TTS

Realtime TTS stellt jetzt den Endpunkt POST /v1/audio/speech im OpenAI-Format bereit, sodass Apps mit OpenAI-Text-to-Speech durch Umstellen des SDK auf Inworld mit einem Inworld-Modell und einer Inworld-Stimme laufen, inklusive Streaming, der Formate mp3, opus, flac, wav und pcm sowie SSE.

OpenAI-compatible speech endpoint

Realtime TTS now serves POST /v1/audio/speech in OpenAI's format, so apps built on OpenAI's text-to-speech can switch to Inworld by pointing the SDK at Inworld and choosing an Inworld model and voice — see OpenAI Compatibility.

  • Works with the OpenAI SDKs: The official Python and Node.js SDKs work unchanged, including streaming responses.
  • Formats: mp3, opus, flac, wav, and pcm, plus server-sent events with stream_format: "sse".
  • One base URL: https://api.inworld.ai/v1 also serves the Inworld Router, so one client configuration covers LLM and TTS.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Letterly

Letterly für Android 2.4.3: Notiz-Audio in der App abspielen

Letterly für Android (Version 2.4.3) enthält einen Audio-Player, mit dem sich die Aufnahme einer Notiz samt zusätzlicher Record-More-Aufnahmen direkt in der App abspielen lässt.

Listen to Note Audio in the App

1 new feature

A useful update for anyone who wants to return to the original recording, not just the transcript.

Audio player

You can now listen to a note’s recording directly in the app. Open a note and play the audio whenever you need.

If you used Record More, those added recordings are included too, so you can listen to all audio parts from the same note.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

ElevenAgents: Ticket-API, Variablenfilter und Twilio-Anrufbeantworter-Erkennung

ElevenAgents bietet eine API zum Auflisten von Workspace-Tickets, Filter für dynamische Variablen bei Konversationen, optionale Twilio-Anrufbeantworter-Erkennung (twilio_machine_detection) sowie allowed_values für Data-Collection-Eigenschaften und erweiterte Agent-Metadaten.

ElevenAgents

  • Workspace-wide conversation tickets: Added List workspace tickets to retrieve conversation triage tickets across accessible agents. Filter by status or assignee and use cursor pagination to retrieve up to 100 tickets per page.
  • Dynamic variable conversation filters: List conversations and Text search conversation messages now accept repeatable dynamic_variable_params filters. Use name:op:value with eq, gt, gte, lt or lte; comparison operators require numeric values.
  • Twilio answering machine detection: Outbound telephony and batch call configurations now support optional twilio_machine_detection. Choose enable for an early human-or-machine verdict or detect_message_end to wait for a voicemail greeting to finish. Results arrive through the new answering_machine_detection webhook event.
  • Data collection value constraints: Data collection properties can now use allowed_values to name a dynamic variable containing the permitted values. The previous allowed_values_dynamic_variable field is deprecated.
  • Agent metadata: List agent branches responses now indicate when a draft was created and whether it predates the branch tip. Agent and Speech Engine summaries now require voice_id. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Wispr Flow

Otter-Transkripte in den Notetaker importieren (Version v1.6.721)

Otter.ai-Transkripte lassen sich nun mit erhaltenen Sprechernamen in den Notetaker importieren, zudem können Admins neue Mitglieder beim Einladen Teams zuweisen, Wörterbucheinträge organisations- oder abteilungsweit teilen, und Updates warten, bis man nicht mehr aktiv arbeitet.

You can now bring your Otter.ai meeting transcripts straight into Wispr Flow. Just drop in your Otter export and your meetings appear alongside everything else, with speaker names preserved instead of collapsed into a single anonymous voice. Imports run in the background, so you can keep working while your history fills in.

Find it: Settings > Notetaker > Import, or the import picker where you'll see the Otter tile alongside Granola.

[Help Center: Import your Otter transcripts into Wispr Flow Notetaker]

For Teams & Enterprise

  • Assign teams at invite time: Admins can now pick which team a new member joins right from the invite form, for single invites or bulk batches, so new hires land in the right place from day one.
  • Org-wide and department-wide dictionary words: You can now choose to share dictionary entries to either your whole organization or your department.

[Help Center: Manage Users and Billing]

Quality-of-Life Improvements

  • Updates wait until you're done: Flow no longer restarts itself mid-task. It detects when you're actively working and holds pending updates until you step away.

‍

v2.3.2

Originalquelle(öffnet in neuem Tab)Problem melden