Zum Inhalt springen

liteLLM Release Notes

16 Einträge aus 1 Quelle. Zuletzt aktualisiert:

Folge liteLLM, um die Release Notes in deinen Feed zu holen.

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

liteLLM

Day-0-Support für Claude Haiku 5.5 in LiteLLM

LiteLLM unterstützt Claude Haiku 5.5 ab Tag 0 über Anthropic, Bedrock, Gemini Enterprise Agent Platform und Azure im LiteLLM AI Gateway, mit Preisen ab $0.10 / MTok Input für Prompts bis 100K Tokens.

Claude Haiku 5.5 on LiteLLM

LiteLLM now supports Claude Haiku 5.5 on Day 0. Use it across Anthropic, Bedrock, Gemini Enterprise Agent Platform, and Azure through the LiteLLM AI Gateway, with spend, rate limits, and logging in one place.

Claude Haiku 5.5 pricing​

  • Input: $0.10 / MTok for prompts up to 100K tokens, $0.50 / MTok above
  • Output: $0.50 / MTok, or $2.50 / MTok above 100K
  • Cache read: $0.01 / MTok, or $0.05 / MTok above 100K
  • Batch: 50% off input and output

Prices on every provider: https://models.litellm.ai/models

What's new in Claude Haiku 5.5​

  • 90% cheaper than Haiku 4.5 for prompts up to 100K tokens, and around 75% less to run on average, per Anthropic
  • 72.4% on OSWorld 2.1, up from Haiku 4.5's 15.7%
  • 1M-token context, up to 128K output tokens

Use Claude Haiku 5.5 with LiteLLM​ …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

liteLLM

Day-0-Support für Nano Banana 2.1 in LiteLLM

LiteLLM unterstützt gemini-nano-banana-2.1 ab Tag 0 über Google AI Studio und die Gemini Enterprise Agent Platform für Bildgenerierung, Bildbearbeitung und Chat Completions und ersetzt damit Nano Banana 2, das am 29. Oktober 2026 abgeschaltet wird.

LiteLLM x Nano Banana 2.1

LiteLLM supports gemini-nano-banana-2.1 on day 0 on Google AI Studio (gemini/) and Gemini Enterprise Agent Platform, formerly Vertex AI (vertex_ai/), through /v1/images/generations, /v1/images/edits and /chat/completions. It replaces Nano Banana 2 (gemini-3.1-flash-image), which the Gemini API shuts down on October 29, 2026.

Per Google, Nano Banana 2.1 improves visual quality and text rendering at 1K, 2K and 4K, fixes tiling on wide aspect ratios, takes up to 14 reference images, and supports Search grounding and configurable thinking levels.

note

No Docker image upgrade needed. Hit Reload Model Cost Map in the Admin UI (or POST /reload/model_cost_map) to pull pricing, on v1.76.0 and above.

Pricing​

  • Image output: $30 / MTok ($0.0336 per 1K image)
  • Text input: $1.50 / MTok
  • Text output: $7.50 / MTok

Quick Start​

  • PROXY
  • SDK

1. Setup config.yaml

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

liteLLM

Time to first byte bei langen Prompts um 94 % gesenkt

LiteLLM überspringt die Token-Zählung der optionalen prompt_caching-Routingprüfung bei nur einem gesunden Deployment, wodurch die mediane Time to first byte im lokalen 440k-Token-Benchmark von 553 ms auf 35 ms sinkt.

LiteLLM's median time to first byte drops from 553 ms to 35 ms in a local 440k-token, single-deployment benchmark, a 94% reduction

LiteLLM was counting every token in a long conversation to answer a yes-or-no routing question.

Removing that unnecessary work took median time to first byte from 553 ms to 35 ms in our local 440k-token benchmark: 94% lower.

Counting 440,000 tokens to check for 1,024​

LiteLLM's optional prompt_caching routing check helps keep requests on deployments that can reuse their cached prompt. To check whether a prompt met a deployment's minimum size, often 1,024 tokens, it counted the entire conversation.

For a coding session with hundreds of turns and tool results, that could mean hundreds of milliseconds of Python work before the model call.

We skip the check for one healthy deployment. There is no routing choice to make, so we now skip the eligibility count, prefix hash, and cache-pin lookup entirely. This is the path measured below. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

liteLLM

Auto Router unterstützt selbst gehostete Klassifizierer Laya und Nimble

Der LiteLLM Auto Router unterstützt jetzt Laya und Bespoke Nimble als selbst gehostete Klassifizierer, die auf eigener Infrastruktur entscheiden, welches Completion-Modell eine Anfrage bearbeitet.

LiteLLM Auto Router now supports Laya and Bespoke Nimble as self-hosted classifiers. Run either model on your infrastructure to choose which completion model handles each request. Your application keeps calling the same router endpoint.

Why self-host a classifier?​

Customers have asked to keep classification local, avoid another hosted classifier vendor, and use clusters they already operate. Self-hosting gives you control over:

  • Prompt data. Process classification context on servers you control, with your own access and logging policies.
  • Vendor dependencies. Run the classifier without opening an account with a separate hosted classifier service.
  • Deployment location. Choose the network and region where classification runs, including your existing private cluster.
  • Capacity and latency tuning. Allocate compute for your traffic, keep the model warm, and place it near the gateway.
  • Model versions. Choose the supported checkpoint you serve and decide when to roll out upgrades.
  • Compute costs. Use your infrastructure budget and compare its cost against hosted inference on your workload.

You operate and pay for the classifier service. Measure latency, routing quality, and total cost on your own traffic with the evaluation guide. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

liteLLM

LiteLLM Lens startet

LiteLLM stellt Lens vor, eine Ebene am Gateway, die aus den Traces von Agenten Erkenntnisse gewinnt und an die Agenten zurückspielt, damit diese sich verbessern.

October 1, 2026

LiteLLM Lens

Agent swarms200K+ traces from every agent

LiteLLM gatewayevery call, one chokepoint

Lensinsight flows back to your agents

The gateway that helps your agents improve

  1. [1]The agentic swarm developer
  2. [2]The problem
  3. [3]Our solution
  4. [4]Launch partners
  5. [5]Get started

Published: Oct 1, 2026

Ishaan JafferCTO, LiteLLMMoe KhalilAI Product Engineer, LiteLLMTin LoAI Product Engineer, LiteLLMYujong LeeSenior SWE, LiteLLM

On this page …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

liteLLM

Redis-Roundtrips pro Request um 64 % reduziert

Ein Request an den LiteLLM-Proxy benötigt im Beispielszenario nur noch 8 statt 22 Redis-Roundtrips (5 vor und 3 nach dem Modellaufruf) bei gleichen Prüfungen und Schreibvorgängen.

LiteLLM's 22 Redis round trips merge into 8, a 64% reduction per request

This week, we cut the number of times a request to the LiteLLM proxy waits on Redis from 22 to 8.

A request from a key with a budget, in a team with a budget and TPM/RPM limits, against a model group with usage-based routing and a Redis response cache, made 22 Redis round trips: 12 before the model was called and 10 after. The same request, with the same checks and the same writes, now makes 5 before and 3 after.

Redis round trips per request

14 fewer Redis round trips per request

Before22 round trips

After8 round trips

Redis round trips

Before

After

Before the model call

12

5

After the model call

10

3

Streaming request

24

8

Response-cache hit

18

7

Once-a-minute auth refresh

46

16

One /v1/chat/completions request · key, team and end-user budgets · TPM and RPM limits · usage-based routing · response cache

Why it was slow​ …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

liteLLM

LiteLLM-Usage-Seite etwa 120x schneller

Die Usage-Seite lädt 30-Tage-Summen bei 5.000 API-Keys nun in 3,2 Sekunden statt in über sechs Minuten, weil Postgres die Aggregation übernimmt und der Browser weniger Daten lädt.

This week, we made the LiteLLM Usage page roughly 120x faster.

With 5,000 API keys, our Usage page took over six minutes to show 30 days of totals. We redesigned how it loads data and brought that down to 3.2 seconds in our benchmark. That's roughly 120x faster.

Time to usage totals

120× faster in this benchmark

Before386s

After3.2s

30-day page load

Before

After

Usage requests

315

6

Data transferred

1.3 GB

17 MB

Peak browser heap

2.0 GB

72 MB

30-day Usage view · 5,000 API keys · same database

Why it was slow​

The browser downloaded every key's daily usage, 1,000 rows at a time, then calculated totals and rankings in JavaScript.

More keys and more history meant more requests, more data, and more work in the browser. Our 30-day test needed 315 requests and 1.3 GB of data before the totals appeared.

What we changed​

Postgres now does the aggregation. The page receives calculated totals and usage breakdowns, instead of downloading every key's history to calculate them.

Before

Download, then calculate

  1. DatabaseDaily rows for every key
  2. TransferPage through the results
  3. BrowserSum totals and sort keys

After

Calculate, then download

  1. DatabaseAggregate across all keys …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

liteLLM

Day-0-Support für GPT-6.1 Sol in LiteLLM

LiteLLM unterstützt GPT-6.1 Sol ab Tag 0 über das AI Gateway mit der üblichen OpenAI-Konfiguration, zum Preis von $2 Input und $10 Output pro 1M Tokens bei auf $0.10 halbiertem Cached Input.

LiteLLM x GPT-6.1 Sol

LiteLLM now supports GPT-6.1 Sol. Route traffic to it through the LiteLLM AI Gateway with the same config you use for every other OpenAI model.

GPT-6.1 Sol is an upgrade to GPT-6 Sol at the same $2 input and $10 output per 1M tokens, with cached input cut in half to $0.10. Per OpenAI, it matches GPT-6 Astra on DeepSWE v1.1 at roughly a fifth of the cost, and on Terminal-Bench Science it averages $5.47 a task against $23.21 for Opus 5.5.

note

No image upgrade needed. Pricing landed in PR #43738; hit Reload Model Cost Map in the Admin UI (or POST /reload/model_cost_map) to pull it, on v1.76.0 and above.

Usage​

  • LiteLLM Proxy
  • LiteLLM Python SDK

1. Setup config.yaml

model_list:  - model_name: gpt-6.1-sol    litellm_params:      model: openai/gpt-6.1-sol      api_key: os.environ/OPENAI_API_KEY

2. Start the proxy

docker run -d \  -p 4000:4000 \  -e OPENAI_API_KEY=$OPENAI_API_KEY \  -v $(pwd)/config.yaml:/app/config.yaml \  ghcr.io/berriai/litellm:main-latest \  --config /app/config.yaml

3. Test it

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

liteLLM

Day-0-Support für Claude Sonnet 5.5 in LiteLLM

LiteLLM unterstützt Claude Sonnet 5.5 ab Tag 0 über Anthropic, Bedrock, Gemini Enterprise Agent Platform und Azure zum gleichen Preis wie Sonnet 5 ($2 / MTok Input, $10 / MTok Output).

LiteLLM x Claude Sonnet 5.5

LiteLLM now supports Claude Sonnet 5.5 on Day 0. Use it across Anthropic, Bedrock, Gemini Enterprise Agent Platform, and Azure through the LiteLLM AI Gateway, with spend, rate limits, and logging in one place.

What's new in Sonnet 5.5​

  • Same price as Sonnet 5: $2 / MTok input, $10 / MTok output, $0.20 / MTok cache reads
  • 30%+ faster, and up to 30% less per task, per Anthropic
  • 70.6% on Terminal-Bench 4.0, up from Sonnet 5's 10.3%
  • 1M-token context, up to 128K output tokens

Enabling Sonnet 5.5​

Pricing landed in PR #43586. No upgrade needed: open Price Data under Models + Endpoints in the UI and click Reload Price Data (or POST /reload/model_cost_map), on v1.76.0 and above.

Usage​

  • Anthropic
  • Bedrock
  • Gemini Enterprise Agent Platform
  • Azure

1. Setup config.yaml

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

liteLLM

LiteLLM ROI Calculator vorgestellt

Der neue LiteLLM ROI Calculator setzt die Ausgaben pro Nutzer aus dem Gateway mit dem geschätzten Aufwand gemergter GitHub-Pull-Requests ins Verhältnis und zeigt die Kosten pro geschätzter Engineering-Stunde.

LiteLLM ROI Calculator overview: $258 of matched gateway spend compared with 78 estimated engineering hours, or $3.31 per estimated hour, above a list of merged pull requests.

Your gateway tells you what your team spends on AI. It doesn't tell you what that spend produced.

The LiteLLM ROI Calculator compares gateway spend with the engineering work your team ships. It reads spend per user from LiteLLM, estimates the effort in each merged pull request with a model you choose, and matches people by email. The result is one number: spend per estimated engineering hour.

Connect your gateway and GitHub​

Run it locally with one command:

git clone https://github.com/BerriAI/litellm-roi-calculator.gitcd litellm-roi-calculatoruv run litellm-roi

The app opens at http://localhost:8787. There is no database server, Node installation, or frontend build.

Setup takes three steps:

  • Gateway: enter your gateway URL and an admin or read-only admin key that can read users and spend.
  • GitHub: click Connect GitHub. GitHub walks you through creating a read-only App and choosing its repositories. There are no client IDs or secrets to copy. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

liteLLM

LiteAgents: Agent-Harnesses per Feld wechseln

LiteAgents erlaubt es, den Agent-Harness (Deep Agents, Pydantic AI, Claude Agent SDK, Codex oder OpenCode) über ein Profilfeld zu wechseln, ohne Tools, MCP-Verbindungen, Modellkonfiguration oder Client-Code umzuschreiben.

LiteAgents: switch harnesses, keep your agent. A ProfileOptions example changes deepagents to claude-sdk while keeping the model, tools, and MCP connections.

You've built an agent with your own tools, prompts, and model configuration. Now you want to try a different harness on the same task.

LiteAgents lets you switch agent harnesses without rewriting your application. Its interface is modeled after the Claude Agent SDK, including the query() pattern and typed messages. Choose Deep Agents, Pydantic AI, Claude Agent SDK, Codex, or OpenCode while keeping your tools, MCP connections, model configuration, and client code. Each selected harness runs its own agent loop.

LiteLLM gives you a common interface to models. LiteAgents brings that approach to agent harnesses, with native controls and optional durability through Temporal.

Change one field to try another harness​

An agent harness runs the loop around a model: supplying context, calling tools, and deciding when to continue. Different harnesses approach that work differently. You should be able to compare them on your own tasks without rebuilding the surrounding application.

In LiteAgents, the harness is a field in your profile:

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

liteLLM

September-Townhall: 583 Bugfixes, OCR auf Rust, 83,5 % Coverage

Der Town-Hall-Rückblick nennt 52 Sicherheitsfixes, den Start-Abbruch von LiteLLM bei fehlendem oder Standard-Master-Key, OCR standardmäßig auf Rust und eine Testabdeckung über dem 80-%-Ziel.

Thank you to everyone who joined our September town hall. We covered security updates, stability updates, and new features in the product, including OCR running on Rust by default and test coverage past the 80% target we set in August.

Security​

We shipped 52 security fixes this month, broken down by category:

Category

Fixes

Credential / secret / PII hardening

27

General hardening (validation, fail-closed)

12

Quota / budget / rate-limit hardening

10

Access control / authz hardening

3

Major work in security went into authentication:

  • Configurable password policies
  • Blocking passwords found in known breaches
  • Enforced SSO login
  • The option to disable logins that use credentials held in environment variables

Master keys. Wiz found 1 in 10 of the LiteLLM instances they scanned running with a blank or default master key. LiteLLM now refuses to start when the master key is unset or still at the default, so a deployment leaning on it will fail to boot on upgrade.

A new disclosure practice, and more CVEs.

  • For anything an unauthenticated attacker can reach, we build the fix in a private fork and publish it at the same moment as the release, the backports, and the disclosure.
  • We're minting more CVEs than before, mostly low and medium severity.
  • The increase comes from a change in publishing policy. Our security posture hasn't changed. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

liteLLM

Day-0-Support für Gemini 3.8 Flash TTS und Flash-Lite TTS

LiteLLM unterstützt gemini-3.8-flash-tts und gemini-3.8-flash-lite-tts ab Tag 0 über /v1/audio/speech auf Google AI Studio und ersetzt damit die 3.1-Flash-TTS-Preview.

LiteLLM x Gemini 3.8 Flash TTS and Flash-Lite TTS

LiteLLM supports gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts on day 0 through /v1/audio/speech, on Google AI Studio (gemini/). Both are generally available and replace the 3.1 Flash TTS preview.

Per Google, Flash TTS is built for creative work like games, audiobooks and podcasts, and Flash-Lite TTS is the cheaper tier for high-volume dubbing and voice agents. Both draw on more than 2,000 prebuilt voices, and Flash TTS covers 130 languages with automatic detection.

note

No Docker image upgrade needed. Both models route through LiteLLM's existing Gemini speech path, so any recent version works for inference. For cost tracking, hit the Reload Model Cost Map button in the Admin UI (or POST /reload/model_cost_map) to pull pricing from PR #42752, on v1.76.0 and above.

Launch pricing​

Both models launch at promotional pricing through December 31, 2026, and Google's standard pricing applies from January 1, 2027. LiteLLM tracks cost at the launch rate. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

liteLLM

LiteAdmin MCP vorgestellt

LiteAdmin MCP verbindet MCP-Clients oder eigene Agenten mit der Management-API des Gateways, um Virtual Keys, Modell-Deployments, Teams, Budgets und Nutzung zu verwalten.

Introducing LiteAdmin MCP: your AI toolkit for gateway management. LiteAdmin for Slack is built on LiteAdmin MCP.Introducing LiteAdmin MCP: your AI toolkit for gateway management. LiteAdmin for Slack is built on LiteAdmin MCP.

An engineer asks for an API key for a new project. You need to choose its models, set a budget, and assign it to a team. As usage grows, you need to check spending and adjust those limits.

LiteAdmin MCP lets your agent handle these tasks through your gateway's management API. Connect it to an MCP client or a custom agent. We built LiteAdmin, our Slack admin agent, on the same connector.

Connect your agent to your gateway​

Connect LiteAdmin MCP to a client such as Claude Code or Cursor, or to your own agent. Your client provides the conversation and model; the connector calls your gateway's management API with your admin credential.

Through that connection, you can:

  • Create and manage virtual keys: Set model access and spending limits.
  • Add model deployments: Use credentials configured on your gateway.
  • Manage teams and budgets: Update team membership and budgets.
  • Inspect usage: Check spending and request logs. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

liteLLM

Day-0-Support für GPT-6 Sol und GPT-6 Luna

LiteLLM unterstützt GPT-6 Sol ($2 Input / $10 Output pro 1M Tokens) und GPT-6 Luna ($0.10 / $0.50) ab Tag 0 über das AI Gateway mit der üblichen OpenAI-Konfiguration.

LiteLLM x GPT-6 Sol and Luna

LiteLLM now supports GPT-6 Sol and GPT-6 Luna. Route traffic to them through the LiteLLM AI Gateway with the same config you use for every other OpenAI model.

Sol and Luna join GPT-6 Astra, trained with the same methods and priced for work at scale. gpt-6-sol is built for complex coding and agentic workflows at $2 input and $10 output per 1M tokens, and gpt-6-luna is the high-volume tier at $0.10 and $0.50. OpenAI prices both at half their GPT-5.6 counterparts' current rates. Per OpenAI, Sol at xhigh scores 33.2% on AutomationBench at $0.27 a task and 68.8% on DeepSWE v1.1 at max, while Luna reaches 66.6% on DeepSWE. Both keep a 1,050,000-token context window with 922K input and 128K output, take text and image input, and run reasoning effort from none to max, defaulting to medium. There is no GPT-6 Terra; Astra stays the top of the range.

note

No image upgrade needed. LiteLLM already treats GPT-6 names as the GPT-5 request family, so max_completion_tokens and the reasoning params are handled on any version from v1.101.0. Pricing landed in PR #42515; hit Reload Model Cost Map in the Admin UI (or POST /reload/model_cost_map) to pull it, on v1.76.0 and above. …

Originalquelle(öffnet in neuem Tab)Problem melden

Angaben zum Datum

Datum aus der Quelle.

Erstmals gesehen am .

liteLLM

Day-0-Support für Claude Opus 5.5 in LiteLLM

LiteLLM unterstützt Claude Opus 5.5 ab Tag 0 über Anthropic, Bedrock, Gemini Enterprise Agent Platform und Azure, mit $4 / MTok Input und $20 / MTok Output und damit 20 % günstiger als Opus 5.

LiteLLM x Claude Opus 5.5

LiteLLM now supports Claude Opus 5.5 on Day 0. Use it across Anthropic, Bedrock, Gemini Enterprise Agent Platform, and Azure through the LiteLLM AI Gateway. Call it with the same OpenAI-compatible request you already use, and track spend, rate limits, and logging in one place.

What's new in Opus 5.5​

Key changes (details from Anthropic):

  • 20% cheaper than Opus 5: $4 / MTok input and $20 / MTok output, down from $5 and $25
  • Cheaper cache reads: $0.20 / MTok, with 5-minute writes at $5 and 1-hour writes at $8
  • Stronger agentic coding: Terminal-Bench 4.0 at 66.4% and OSWorld 2.0 at 81.8%
  • Faster and leaner: Anthropic reports output more than 30% faster, using fewer tokens per task
  • Thinking is always on: effort is the only control, and its default is medium
  • Fast mode: about 2x the base price, down from Opus 5's $10 / $50

Context window and max output are unchanged from Opus 5 at 1M and 128K.

Enabling Opus 5.5​ …

Originalquelle(öffnet in neuem Tab)Problem melden