What BerriAI/litellm shipped
Written by FoxPlug from public releases; not affiliated with LiteLLM. An automatic summary of the public release, pull request and commit data of github.com/BerriAI/litellm. LiteLLM did not write it and does not use or endorse FoxPlug. Every line links to the public change it describes.
Get a weekly update like this for your product, free
Week of September 21, 2026
What shipped
- The CLI now reuses saved agent setup across reconnections and adds a `reconfigure` command to edit saved choices with prefilled defaults. Pull request #43392
- Router can now opt in to prompt-cache cost routing to factor in savings from warm caches when choosing between routes. Pull request #43232
- Fixed streaming text completions with `include_usage` to convert provider usage objects to LiteLLM format before serializing. Pull request #43047
- Fixed logging callback deletion on rc/1.104.0 so deleted callbacks are actually removed from the stored config instead of continuing to export. Pull request #43432
- Tool call arguments parsing now salvages concatenated JSON objects on the normalized response path instead of collapsing them to empty objects. Pull request #43260
- Added validation for `stream_chunk_size` before any provider call to catch bad values like zero or non-numeric strings early. Pull request #43222
- LiteLLM's internal `_litellm_*` kwargs are now filtered by construction to prevent them from leaking into provider request bodies. Pull request #43221
Why it matters
This week brought fixes for streaming completions with usage reporting, logging callback deletion, and tool parameter handling that improve reliability in production deployments. Router enhancements for cache-aware cost calculation and CLI improvements for agent reconfiguration make operational workflows smoother. Multiple backports to rc branches ensure the release candidate is stable before shipping.
Changelog entry
- CLI: Agent setup configuration reuse and reconfigure command Pull request #43392
- Router: Optional cache-aware cost routing Pull request #43232
- Streaming: Fix text-completion usage chunk serialization Pull request #43047
- Proxy: Logging callback deletion from stored config Pull request #43432
- Vertex AI: Skip vertexai SDK import in partner-model completion Pull request #42274
- Tools: Salvage concatenated JSON tool call arguments Pull request #43260
- Params: Validate stream_chunk_size before provider calls Pull request #43222
- Params: Filter _litellm_* kwargs from provider requests Pull request #43221
New in LiteLLM: prompt-cache cost routing, streaming usage fixes, callback deletion repair, and improved tool call parsing. v1.104.0-dev.2 available now.
LiteLLM this week: prompt-cache cost routing lets your router factor in warm cache savings when picking providers. Streaming completions with usage now work reliably. Logging callbacks delete cleanly. Tool calls with concatenated JSON get salvaged instead of dropped. CLI agents reuse setup across reconnects. Available in v1.104.0-dev.2.