Zum Inhalt springen

Eleven Labs Release Notes

99 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Folge Eleven Labs, um die Release Notes in deinen Feed zu holen.

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Opus-Format, Twilio-Outbound und Actor Mode

Text to Speech unterstützt Opus mit 48 kHz und liefert genauere Websocket-Fehlercodes, die Agents Platform kann Twilio-Outbound-Anrufe nativ ausführen, ElevenCreative Studio bekommt den Actor Mode und Dubbing-Duplikation steht allen Nutzern offen.

Text to speech

  • Opus format support: Added support for Opus format with 48kHz sample rate across multiple bitrates (32-192 kbps).
  • Improved websocket error handling: Updated TTS websocket API to return more accurate error codes (1011 for internal errors instead of 1008) for better error identification and SLA monitoring.

Agents Platform

  • Twilio outbound: Added ability to natively run outbound calls.
  • Post-call webhook override: Added ability to override post-call webhook settings at the agent level, providing more flexible configurations.
  • Large knowledge base document viewing: Enhanced the knowledge base interface to allow viewing the entire content of large RAG documents.
  • Added call SID dynamic variable: Added system__call_sid as a system dynamic variable to allow referencing the call ID in prompts and tools.

ElevenCreative Studio

  • Actor Mode: Added Actor Mode in ElevenCreative Studio, allowing you to use your own voice recordings to direct the way speech should sound in ElevenCreative Studio projects.
  • Improved keyboard shortcuts: Updated keyboard shortcuts for viewing settings and editor shortcuts to avoid conflicts and simplified shortcuts for locking paragraphs.

Dubbing

  • Dubbing duplication: Made dubbing duplication feature available to all users.
  • Manual mode foreground generation: Added ability to generate foreground audio when using manual mode with a file and CSV. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Voices-Suche V2, automatische Spracherkennung und Outbound-Anrufe

Neu sind ein V2-Endpoint zur Stimmensuche, native Outbound-Anrufe für Twilio-Nummern, ein System-Tool zur automatischen Spracherkennung und anpassbare Widget-Steuerelemente, außerdem wurden Fehler bei Sound Effects und Phonem-Tags behoben und die Speech-to-Text-Wiederholungserkennung verbessert.

Voices

  • List Voices V2: Added a new V2 voice search endpoint with better search and additional filtering options

Agents Platform

  • Native outbound calling: Added native outbound calling for Twilio-configured numbers, eliminating the need for complex setup configurations. Outbound calls are now visible in the Call History page.
  • Automatic language detection: Added new system tool for automatic language detection that enables agents to switch languages based on both explicit user requests ("Let's talk in Spanish") and implicit language in user audio.
  • Pronunciation dictionary improvements: Fixed phoneme tags in pronunciation dictionaries to work correctly with Agents Platform.
  • Large RAG document viewing: Added ability to view the entire content of large RAG documents in the knowledge base.
  • Customizable widget controls: Updated UI to include an optional mute microphone button and made widget icons customizable via slots.

Sound Effects

  • Fractional duration support: Fixed an issue where users couldn't enter fractional values (like 0.5 seconds) for sound effect generation duration.

Speech to Text

  • Repetition handling: Improved detection and handling of repetitions in speech-to-text processing.

ElevenCreative Studio …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Gemini 2.0 Flash als Standard-LLM und günstigere Scribe-Preise

In der Agents Platform ist Gemini 2.0 Flash nun das Standard-LLM, die Wissensdatenbank wurde neu gestaltet und es gibt System-Variablen sowie RAG-Chunks im Verlauf, außerdem sinkt der Preis für Scribe und Voice Activity Detection sowie Diarisierung wurden verbessert.

Agents Platform

  • Default LLM update: Changed the default agent LLM from Gemini 1.5 Flash to Gemini 2.0 Flash for improved performance.
  • Fixed incorrect conversation abandons: Improved detection of conversation continuations, preventing premature abandons when users repeat themselves.
  • Twilio information in history: Added Twilio call details to conversation history for better tracking.
  • Knowledge base redesign: Redesigned the knowledge base interface.
  • System dynamic variables: Added system dynamic variables to use time, conversation id, caller id and other system values as dynamic variables in prompts and tools.
  • Twilio client initialisation: Adds an agent level override for conversation initiation client data twilio webhook.
  • RAG chunks in history: Added retrieved chunks by RAG to the call transcripts in the history view.

Speech to Text

  • Reduced pricing: Reduced the pricing of our Scribe model, see more here.
  • Improved VAD detection: Enhanced Voice Activity Detection with better pause detection at segment boundaries and improved handling of silent segments.
  • Enhanced diarization: Improved speaker clustering with a better ECAPA model, symmetric connectivity matrix, and more selective speaker embedding generation. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

HIPAA-Konformität, Cascade LLM und günstigeres Scribe

Agents Platform und Scribe sind unter bestimmten Voraussetzungen HIPAA-konform, die Agents Platform bekommt Cascade LLM, bessere Fehlermeldungen und Audio-Umschaltung, Scribe wird günstiger und unterstützt Diarization für Audiodateien bis 2 Stunden, und weitere Korrekturen betreffen Text to Speech und Dubbing.

Agents Platform

  • HIPAA compliance: Agents Platform is now HIPAA compliant on appropriate plans, when a BAA is signed, zero-retention mode is enabled and appropriate LLMs are used. For access please contact sales
  • Cascade LLM: Added dynamic dispatch during the LLM step to other LLMs if your default LLM fails. This results in higher latency but prevents the turn failing.
  • Better error messages: Added better error messages for websocket failures.
  • Audio toggling: Added ability to select only user or agent audio in the conversation playback.

Scribe

  • HIPAA compliance: Added a zero retention mode to Scribe to be HIPAA compliant.
  • Diarization: Increased time length of audio files that can be transcribed with diarization from 8 minutes to 2 hours.
  • Cheaper pricing: Updated Scribe's pricing to be cheaper, as low as $0.22 per hour for the Business tier.
  • Memory usage: Shipped improvements to Scribe's memory usage.
  • Fixed timestamps: Fixed an issue that was causing incorrect timestamps to be returned.

Text to Speech

  • Pronunciation dictionaries: Fixed pronunciation dictionary rule application for replacements that contain symbols.

Dubbing

  • Studio support: Added support for creating dubs with dubbing_studio enabled, allowing for more advanced dubbing workflows beyond one-off dubs.

Voices …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Scribe im Dubbing, Speed Control und Claude 3.7 Sonnet

Dubbing Studio nutzt standardmäßig Scribe zur Spracherkennung, Speech to Text wurde stabilisiert, und die Agents Platform erhält Speed Control, Post-Call-Webhooks, bessere Websocket-Fehlermeldungen und Claude 3.7 Sonnet als LLM-Option.

Dubbing

  • Scribe for speech recognition: Dubbing Studio now uses Scribe by default for speech recognition to improve accuracy.

Speech to Text

  • Fixes: Shipped several fixes improving the stability of Speech to Text.

Agents Platform

  • Speed control: Added speed control to an agent's settings in Agents Platform.
  • Post call webhook: Added the option of sending post-call webhooks after conversations are completed.
  • Improved error messages: Added better error messages to the Agents Platform websocket.
  • Claude 3.7 Sonnet: Added Claude 3.7 Sonnet as a new LLM option in Agents Platform.

API

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Neue Speech-to-Text-API, Geschwindigkeitssteuerung und Dark Mode

ElevenLabs startet eine Speech-to-Text-API für 99 Sprachen, Text to Speech erhält eine Geschwindigkeitssteuerung, die Web-App bietet einen Dark Mode und es gibt Verbesserungen bei ElevenCreative Studio, Dubbing, Instant Voice Cloning und der Agents Platform.

Speech to Text

  • ElevenLabs launched a new state of the art Speech to Text API available in 99 languages.

Text to Speech

  • Speed control: Added speed control to the Text to Speech API.

ElevenCreative Studio

  • Auto-assigned projects: Increased token limits for auto-assigned projects from 1 month to 3 months worth of tokens, addressing user feedback about working on longer projects.
  • Language detection: Added automatic language detection when generating audio for the first time, with suggestions to switch to Eleven Turbo v2.5 for languages not supported by Multilingual v2 (Hungarian, Norwegian, Vietnamese).
  • Project export: Enhanced project exporting in ElevenReader with better metadata tracking.

Dubbing

  • Clip overlap prevention: Added automatic trimming of overlapping clips in dubbing jobs to ensure clean audio tracks for each speaker and language.

Voice Management

  • Instant Voice Cloning: Improved preview generation for Instant Voice Cloning v2, making previews available immediately.

Agents Platform

  • Agent ownership: Added display of agent creators in the agent list, improving visibility and management of shared agents.

Web app

  • Dark mode: Added dark mode to the web app.

API

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Tool-Calling-Fixes und bessere Abrechnungsübersicht

In der Agents Platform funktioniert Tool Calling mit gpt-4o mini wieder und unterstützt dynamische Variablen in Objekten und Arrays, der Voice Isolator wurde repariert und die Abrechnung unterscheidet nun Rollover-, Zyklus-, geschenkte und nutzungsbasierte Credits.

Agents Platform

  • Tool calling fix: Fixed an issue where tool calling was not working with agents using gpt-4o mini. This was due to a breaking change in the OpenAI API.
  • Tool calling improvements: Added support for tool calling with dynamic variables inside objects and arrays.
  • Dynamic variables: Fixed an issue where dynamic variables of a conversation were not being displayed correctly.

Voice Isolator

  • Fixed: Fixed an issue that caused the voice isolator to not work correctly temporarily.

Workspace

  • Billing: Improved billing visibility by differentiating rollover, cycle, gifted, and usage-based credits.
  • Usage Analytics: Improved usage analytics load times and readability.
  • Fine grained fiat billing: Added support for customizable pricing based on several factors.

API

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Neue Preise, Wissensdatenbank-Seite und Aufbewahrungseinstellungen

Die Agents Platform bekommt gesenkte Self-Serve-Preise mit größerem Gratis-Kontingent, eine neue Wissensdatenbank-Seite, Anzeige laufender Anrufe, einstellbare Aufbewahrung von Transkripten und Audio sowie 8k-PCM-Unterstützung, zudem akzeptiert GenFM mehrere Eingabequellen und es gibt neue Endpoints für Workspace-Gruppen.

Agents Platform

  • Updated Pricing: Updated self-serve pricing for Agents Platform with reduced cost and a more generous free tier.
  • Knowledge Base UI: Created a new page to easily manage your knowledge base.
  • Live calls: Added number of live calls in progress in the user dashboard and as a new endpoint.
  • Retention: Added ability to customize transcripts and audio recordings retention settings.
  • Audio recording: Added a new option to disable audio recordings.
  • 8k PCM support: Added support for 8k PCM audio for both input and output.

ElevenCreative Studio

  • GenFM: Updated the create podcast endpoint to accept multiple input sources.
  • GenFM: Fixed an issue where GenFM was creating empty podcasts.

Enterprise

  • New workspace group endpoints: Added new endpoints to manage workspace groups.

API

Socials

  • ElevenLabs Developers: Follow our new developers account on X @ElevenLabsDevs

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Agent-Monitoring, ElevenCreative Studio und GenFM-API

Die Agents Platform erhält ein Monitoring-Dashboard und mehrere Fehlerbehebungen, Projects heißt jetzt ElevenCreative Studio und ist allgemein verfügbar, außerdem gibt es eine öffentliche GenFM-API und Bearbeitung von Kapitelinhalten per API.

Agents Platform

  • Agent monitoring: Added a new dashboard for monitoring ElevenLabs agents' activity. Check out your's here.
  • Proactive conversations: Enhanced capabilities with improved timeout retry logic. Learn more
  • Tool calls: Fixed timeout issues occurring during tool calls
  • Allowlist: Fixed implementation of allowlist functionality.
  • Content summarization: Added Gemini as a fallback model to ensure service reliability
  • Widget stability: Fixed issue with dynamic variables causing the Agents Platform widget to fail

Reader

  • Trending content: Added carousel showcasing popular articles and trending content
  • New publications: Introduced dedicated section for recent ElevenReader Publishing releases

ElevenCreative Studio (formerly Projects)

  • Projects is now ElevenCreative Studio and is now generally available to everyone
  • Chapter content editing: Added support for editing chapter content through the public API, enabling programmatic updates to chapter text and metadata
  • GenFM public API: Added public API support for podcast creation through GenFM. Key features include:
    • Conversation mode with configurable host and guest voices
    • URL-based content sourcing
    • Customizable duration and highlights
    • Webhook callbacks for status updates …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Neue Docs, dynamische Variablen und Bun-/Deno-Unterstützung

Die Dokumentation wurde neu veröffentlicht, die Agents Platform bietet dynamische Variablen, ignorierbare Unterbrechungen, bessere Twilio-Audioqualität und PCM 8000, Projects erhält Auto-regenerate ohne Zusatzkosten und die SDKs laufen mit Bun 1.1.45+ und Deno 2.1.7+.

Docs

  • Shipped our new docs: we're keen to hear your thoughts, you can reach out by opening an issue on GitHub or chatting with us on Discord

Agents Platform

  • Dynamic variables: Available in the dashboard and SDKs. Learn more
  • Interruption handling: Now possible to ignore user interruptions in Agents Platform. Learn more
  • Twilio integration: Shipped changes to increase audio quality when integrating with Twilio
  • Latency optimization: Published detailed blog post on latency optimizations. Read more
  • PCM 8000: Added support for PCM 8000 to ElevenLabs agents
  • Websocket improvements: Fixed unexpected websocket closures

Projects

  • Auto-regenerate: Auto-regeneration now available by default at no extra cost
  • Content management: Added updateContent method for dynamic content updates
  • Audio conversion: New auto-convert and auto-publish flags for seamless workflows

API

SDKs

  • Cross-Runtime Support: Now compatible with Bun 1.1.45+ and Deno 2.1.7+ …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Agents Platform: Sprachauswahl, End Call Tool und Flash als Standard

In der Agents Platform gibt es ein Sprach-Dropdown im Widget, ein neues "End Call"-Tool, Flash als Standardmodell für neue Agenten, zusätzliche Datenschutzoptionen und ein höheres Tool-Limit von 5 auf 15.

Product

Agents Platform

  • Additional languages: Add a language dropdown to your widget so customers can launch conversations in their preferred language. Learn more here.
  • End call tool: Let the agent automatically end the call with our new “End Call” tool. Learn more here
  • Flash default: Flash, our lowest latency model, is now the default for new agents. In your agent dashboard under “voice”, you can toggle between Turbo and Flash. Learn more about Flash here.
  • Privacy: Set concurrent call and daily call limits, turn off audio recordings, add feedback collection, and define customer terms & conditions.
  • Increased tool limits: Increase the number of tools available to your agent from 5 to 15. Learn more here.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Workspace-Gruppen und Berechtigungen eingeführt

Neue Funktionen zur Verwaltung von Workspace-Gruppen verbessern die Zugriffskontrolle innerhalb von Organisationen.

Product

  • Workspace Groups and Permissions: Introduced new workspace group management features to enhance access control within organizations. Learn more.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Flash: schnellstes Text-to-Speech-Modell mit 75 ms

Das neue Modell Flash erzeugt Sprache in nur 75 ms und ist per API über die Modell-IDs eleven_flash_v2 und eleven_flash_v2_5 verfügbar, zudem gibt es die Aktionen TalkToSanta.io und AI Engineer Pack.

Model

  • Introducing Flash: Our fastest text-to-speech model yet, generating speech in just 75ms. Access it via the API with model IDs eleven_flash_v2 and eleven_flash_v2_5. Perfect for low-latency Agents Platform applications. Try it now.

Launches

  • TalkToSanta.io: Experience Agents Platform in action by talking to Santa this holiday season. For every conversation with santa we donate 2 dollars to Bridging Voice (up to $11,000).

  • AI Engineer Pack: Get $50+ in credits from leading AI developer tools, including ElevenLabs.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

GenFM jetzt im Web verfügbar

GenFM lässt sich nun zusätzlich zur ElevenReader-App direkt über die Website nutzen.

Product

  • GenFM Now on Web: Access GenFM directly from the website in addition to the ElevenReader App, try it now.

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

API-Keys: Credit-Limits und Zugriffsrechte

Für API-Keys lassen sich nun Credit-Limits sowie Zugriffsrechte (z. B. "Access"/"No Access" für Dubbing oder Audio Native) festlegen, Workspace API Keys unterstützen Berechtigungen wie "Read" oder "Read and Write", und die Key-Verwaltung wurde mit eigenen Seiten statt Modals und mehr Detailinformationen überarbeitet.

API

  • Credit Usage Limits: Set specific credit limits for API keys to control costs and manage usage across different use cases by setting "Access" or "No Access" to features like Dubbing, Audio Native, and more. Check it out
  • Workspace API Keys: Now support access permissions, such as "Read" or "Read and Write" for User, Workspace, and History resources.
  • Improved Key Management:
    • Redesigned interface moving from modals to dedicated pages
    • Added detailed descriptions and key information
    • Enhanced visibility of key details and settings

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

GenFM, Agents Platform allgemein verfügbar und Projects-Überarbeitung

GenFM startet in der ElevenReader-App, die Agents Platform ist für alle Kunden verfügbar mit SDKs und APIs, das TTS-Redesign der Website wird ausgerollt und Projects erhält Auto-regenerate sowie weitere Bearbeitungsfunktionen.

Product

  • GenFM: Launched in the ElevenReader app. Learn more

  • Agents Platform: Now generally available to all customers. Try it now

  • TTS Redesign: The website TTS redesign is now rolled out to all customers.

  • Auto-regenerate: Now available in Projects. Learn more

  • Reader Platform Improvements:

    • Improved content sharing with enhanced landing pages and social media previews.
    • Added podcast rating system and improved voice synchronization.
  • Projects revamp:

    • Restore past generations, lock content, assign speakers to sentence fragments, and QC at 2x speed. Learn more
    • Auto-regeneration identifies mispronunciations and regenerates audio at no extra cost. Learn more

API

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

u-law-Formate, TTS-Websocket-Verbesserungen und TTS-Redesign

Die Convai API unterstützt u-law-Audioformate für Twilio, die TTS-Websockets wurden verbessert und erhalten einen Auto Mode für geringere Latenz, die Latenzkonsistenz aller Modelle wurde verbessert, und das TTS-Redesign der Website startet in der Alpha.

API

  • u-law Audio Formats: Added u-law audio formats to the Convai API for integrations with Twilio.
  • TTS Websocket Improvements: TTS websocket improvements, flushes and generation work more intuitively now.
  • TTS Websocket Auto Mode: A streamlined mode for using websockets. This setting reduces latency by disabling chunk scheduling and buffers. Note: Using partial sentences will result in significantly reduced quality.
  • Improvements to latency consistency: Improvements to latency consistency for all models.

Website

  • TTS Redesign: The website TTS redesign is now in alpha!

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Text-Normalisierung in der TTS-API und Voice Design Beta

Die TTS-API bietet den neuen Parameter apply_text_normalization zur Normalisierung des Eingabetexts (bei v2.5-Modellen nur mit Enterprise-Plänen), und Voice Design ist nun in der Beta.

API

  • Normalize Text with the API: Added the option normalize the input text in the TTS API. The new parameter is called apply_text_normalization and works on all models. For v2.5 models, this feature is available with Enterprise plans only.

Product

  • Voice Design: The Voice Design feature is now in beta!

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

Eleven Labs

Bessere Audiostabilität, geringere Latenz und Agents Platform Beta

Die Audiostabilität wurde über alle Modelle hinweg verbessert, die Latenz bis zum ersten Byte sinkt um etwa 20–30 ms, Hintergrundgeräusche lassen sich aus Stimmproben und STS-Eingaben entfernen und die Agents Platform befindet sich in der Beta.

Model

  • Stability Improvements: Significant audio stability improvements across all models, most noticeable on turbo_v2 and turbo_v2.5, when using:
    • Websockets
    • Projects
    • Reader app
    • TTS with request stitching
    • ConvAI
  • Latency Improvements: Reduced time to first byte latency by approximately 20-30ms for all models.

API

  • Remove Background Noise Voice Samples: Added the ability to remove background noise from voice samples using our audio isolation model to improve quality for IVCs and PVCs at no additional cost.
  • Remove Background Noise STS Input: Added the ability to remove background noise from STS audio input using our audio isolation model to improve quality at no additional cost.

Feature

  • Agents Platform Beta: Agents Platform is now in beta.

Originalquelle(öffnet in neuem Tab)Problem melden