Commit Graph

30712 Commits

Author SHA1 Message Date
Sameer Kankute
22158f8f03 Fix litellm_staging_01_20_2026 mypy issues 2026-01-21 17:35:28 +05:30
Sameer Kankute
c8d065656e Fix litellm_staging_01_20_2026 mypy issues 2026-01-21 17:33:53 +05:30
Sameer Kankute
37d27e67d5
Merge pull request #19489 from cluebbehusen/fix-gemini-2-5-flash-lite-pricing
fix: correct gemini-2.5-flash-lite audio input and cache read pricing
2026-01-21 16:41:22 +05:30
Sameer Kankute
a1aba2ed8d
Merge pull request #19491 from BerriAI/main
merge main 20 1 25
2026-01-21 16:40:48 +05:30
Connor Luebbehusen
b810a68f89
fix: correct gemini-2.5-flash-lite audio input and cache read pricing 2026-01-21 05:58:38 -05:00
Ishaan Jaff
a02c43d300
Litellm cc docs max (#19466)
* docs claude code max

* docs fix

* docs

* docs fix

* docs fix
2026-01-20 20:08:36 -08:00
Sameer Kankute
a5ea08a0bf Fix test_default_api_base failing because of chatgpt as provider 2026-01-21 09:32:38 +05:30
Sameer Kankute
a5b8da7f87
Merge pull request #19463 from BerriAI/litellm_2101_cicd_fixes
Fixes test_aaabasic_gcs_logger
2026-01-21 09:22:40 +05:30
Cesar Garcia
b4ed387d24
fix(vertex_ai): handle reasoning_effort as dict from OpenAI Agents SDK (#19419)
The OpenAI Agents SDK (v0.6.9+) now passes reasoning_effort as a dict
when summary is specified: {"effort": "high", "summary": "auto"}

This change extracts the "effort" value from the dict for Vertex AI,
which only supports thinkingLevel (not summary).

Before: reasoning_effort={"effort": "high"} was silently ignored
After: reasoning_effort={"effort": "high"} correctly maps to thinkingLevel

Fixes #19411
2026-01-20 19:31:25 -08:00
Sameer Kankute
0b9f6b543f Fixes test_aaabasic_gcs_logger 2026-01-21 08:58:16 +05:30
Harshit Jain
b36e704e06
fix: ensure auto-rotation updates existing AWS secret instead of creating new one (#19455) 2026-01-20 18:30:36 -08:00
Ishaan Jaffer
0d75e506c5 fix mock_handler 2026-01-20 18:01:22 -08:00
yuneng-jiang
6e41615a0b
Merge pull request #19460 from BerriAI/litellm_cicd_fix_yj_06
bump: version 1.81.0 → 1.81.1
2026-01-20 17:59:37 -08:00
yuneng-jiang
9df4bd2bbb bump: version 1.81.0 → 1.81.1 2026-01-20 17:58:10 -08:00
Ishaan Jaffer
b9de10bd27 test token ctr 2026-01-20 17:53:53 -08:00
yuneng-jiang
e57214dfab
Merge pull request #19457 from BerriAI/revert-19456-litellm_cicd_fix_yj_05
Revert "[Infra] Changing Google Tests to use Gemini 3 Flash Preview"
2026-01-20 17:32:29 -08:00
yuneng-jiang
96b2d134a7
Revert "[Infra] Changing Google Tests to use Gemini 3 Flash Preview" 2026-01-20 17:32:10 -08:00
Ishaan Jaff
af957a3a9e
[Fix] Claude Code - /messages/token_counter - ensure it works for Anthropic, Azure AI Anthropic on AI Gateway (#19432)
* fix count_tokens_with_anthropic_api

* remove outdated file

* fix ANTHROPIC_TOKEN_COUNTING_BETA_VERSION

* refactor: get_token_counter

* init test suite for token counter

* init token counters

* fix: fix pyrightI

* fix Code QA issues

* fix: return Ant response no transfrom

* fix optionally_handle_anthropic_oauth
2026-01-20 17:31:20 -08:00
Ishaan Jaff
ddebdd47bc
[Feat] Add Support for Claude Code Max/OAuth 2 on LiteLLM AI Gateway (#19453)
* fix count_tokens_with_anthropic_api

* remove outdated file

* fix ANTHROPIC_TOKEN_COUNTING_BETA_VERSION

* refactor: get_token_counter

* init test suite for token counter

* init token counters

* fix: fix pyrightI

* fix Code QA issues

* feat: add OAUTH handling ant

* feat: Oauth handling Ant

* test anthopic common utils

* fix code QA

* docs
2026-01-20 17:21:17 -08:00
yuneng-jiang
351e3a5f3c
Merge pull request #19456 from BerriAI/litellm_cicd_fix_yj_05
[Infra] Changing Google Tests to use Gemini 3 Flash Preview
2026-01-20 17:15:03 -08:00
yuneng-jiang
0dc92a6979 changing to gemini 3 flash preview 2026-01-20 17:14:06 -08:00
yuneng-jiang
498ee5a662
Merge pull request #19452 from BerriAI/litellm_cicd_fix_yj_04
[Infra] Increase Time to Wait for Spend Accuracy Tests
2026-01-20 16:34:05 -08:00
yuneng-jiang
65829411e0 increasing time for spend tracking 2026-01-20 16:31:25 -08:00
yuneng-jiang
d5305c6a61
Merge pull request #19450 from BerriAI/litellm_cicd_fix_yj_02
[Infra] Fix test_route_checks
2026-01-20 16:22:20 -08:00
yuneng-jiang
b37e42ac0f
Merge pull request #19451 from BerriAI/litellm_cicd_fix_yj_03
[Infra] Use mock db for claude code marketplace tests
2026-01-20 16:19:15 -08:00
yuneng-jiang
1a9a7df437 use mock db for cluade code marketplace 2026-01-20 16:18:21 -08:00
yuneng-jiang
232ae52b94 attempt test_route_checks fix 2026-01-20 15:55:44 -08:00
yuneng-jiang
e0811ad848
Merge pull request #19448 from BerriAI/litellm_cicd_fix_yj_01
[Infra[ Fixing dynamic_router_retry_policy CI
2026-01-20 15:52:40 -08:00
yuneng-jiang
dd9e8833db Fixing dynamic_router_retry_policy 2026-01-20 15:51:57 -08:00
yuneng-jiang
231023c422
Merge pull request #19446 from BerriAI/migration_fix_yj
[Infra] Fixing LiteLLM Proxy Extras
2026-01-20 15:24:15 -08:00
Cesar Garcia
94055741d4
docs: clarify Gemini vs Vertex AI model prefix behavior (#19443)
Add documentation explaining the difference between model formats:
- `gemini/model` → Gemini API (simple API key)
- `vertex_ai/model` → Vertex AI (GCP credentials)
- `model` (no prefix) → defaults to Vertex AI

This addresses user confusion when models without prefix require
GCP authentication instead of simple API key auth.

Ref #8424
2026-01-20 15:22:52 -08:00
yuneng-jiang
4b25ae6693 Adding build artifacts 2026-01-20 15:22:47 -08:00
yuneng-jiang
2c2e0649d9 bump: version 0.4.24 → 0.4.25 2026-01-20 15:22:18 -08:00
Cesar Garcia
2b44d02682
fix: add google-cloud-aiplatform as optional dependency with clear error message (#19437)
- Add google-cloud-aiplatform as optional dependency in pyproject.toml
- Add 'google' extra for easy installation: pip install litellm[google]
- Improve error messages when Google SDK is not installed to guide users

Fixes #5483
2026-01-20 15:22:13 -08:00
yuneng-jiang
0bcf7097d2 bump: version 0.4.23 → 0.4.24 2026-01-20 15:22:11 -08:00
yuneng-jiang
2b62e9fedf
Merge pull request #19440 from BerriAI/litellm_ui_chat-autofill
[Feature] UI - Playground: Button to Fill Custom API Base
2026-01-20 14:23:39 -08:00
yuneng-jiang
71a2fc5331 fix classnames 2026-01-20 14:16:03 -08:00
Cesar Garcia
7515f179e7
fix: sync Helm chart version with LiteLLM release version (#19438)
Replace independent auto-incrementing chart versioning with 1-1 sync
to LiteLLM version. This allows users to easily map Helm chart versions
to LiteLLM versions without needing to inspect appVersion.

Changes:
- Remove auto-increment logic that read from OCI registry
- Chart version now equals LiteLLM tag without 'v' prefix (v1.81.0 -> 1.81.0)
- appVersion equals full Docker tag (v1.81.0)
- Update both ghcr_deploy.yml and ghcr_helm_deploy.yml workflows

Before: helm chart 0.1.837 -> user has to guess LiteLLM version
After:  helm chart 1.81.0  -> matches LiteLLM v1.81.0

References:
- https://codefresh.io/docs/docs/ci-cd-guides/helm-best-practices/
2026-01-20 14:13:26 -08:00
yuneng-jiang
f434d1c847 Option to pre fill custom proxy base URL 2026-01-20 14:07:54 -08:00
Alexsander Hamir
7f81dea8b3
Add custom auth header support and increase default prompt size to 100k chars (#19436) 2026-01-20 13:25:12 -08:00
yuneng-jiang
e142474e0b
Merge pull request #19431 from BerriAI/litellm_ui_fix_build_002
[Infra] UI - Fixing UI Build
2026-01-20 13:17:25 -08:00
yuneng-jiang
3ae71bf49e fixing ui build 2026-01-20 13:03:27 -08:00
Harshit Jain
20323feecc
fix(prompts): fix prompt info lookup and delete using correct IDs (#19358)
* fix(prompts): fix prompt info lookup and delete using correct IDs

* add regression tests cases
2026-01-20 12:28:34 -08:00
yuneng-jiang
bfb94f56b7
Merge pull request #19276 from stiyyagura0901/litellm_fix_ui_auth_header_override
fix: UI dashboard respects custom authentication header override
2026-01-20 12:24:14 -08:00
Alexsander Hamir
5a06868652
Fix in-flight request termination on SIGTERM when health-check runs in a separate process (#19427) 2026-01-20 12:17:06 -08:00
Kris Xia
56bf6001e9
Supports setting media_resolution and fps parameters on each video file, when using Gemini video understanding. (#19273)
* feat: add gemini video metadata and detail support

Implement support for video_metadata and enhanced detail parameter
for Gemini 3.0+ models:

- Add video_metadata field to ChatCompletionFileObjectFile type
  - Supports fps, start_offset, and end_offset parameters
  - Properly converts snake_case to camelCase for Gemini API
- Extend detail parameter to support medium and ultra_high levels
  - Maps to MEDIA_RESOLUTION_MEDIUM and MEDIA_RESOLUTION_ULTRA_HIGH
- Update _process_gemini_image to handle video metadata transformation
- Add version gating to only apply features for Gemini 3+ models
- Add comprehensive test coverage (6 new test cases)
  - Test detail parameter with file objects
  - Test video_metadata fields (fps, start_offset, end_offset)
  - Test combined detail + video_metadata usage
  - Test new detail levels (medium, ultra_high)
  - Test version gating (Gemini 1.5 vs 3.0)

Note: video_metadata is only supported for video files but error
handling is delegated to Vertex AI for other media types.

* refactor: rename _process_gemini_image to _process_gemini_media

The function handles multiple media types (images, audio, video, PDF),
not just images. Renamed to better reflect its actual purpose.

- Update function name in transformation.py
- Update all function calls and references
- Update test names and imports to match
- Improve docstring to clarify it handles all media types

* docs: add video metadata and media resolution control documentation

Add comprehensive documentation for Gemini 3+ video processing features:
- Document media resolution control (detail parameter) for images and videos
- Add video_metadata field documentation (fps, start_offset, end_offset)
- Include usage examples with tabs for basic, combined, and proxy scenarios
- Update both Gemini and Vertex AI provider documentation
- Clarify snake_case to camelCase field conversion for Gemini API

Signed-off-by: Kris Xia <xiajiayi0506@gmail.com>

* refactor(gemini): extract metadata application into helper function

Extract duplicated Gemini 3+ media_resolution and video_metadata
application logic from _process_gemini_media into a dedicated
_apply_gemini_3_metadata helper function to improve code maintainability.

---------

Signed-off-by: Kris Xia <xiajiayi0506@gmail.com>
2026-01-20 11:36:55 -08:00
Krrish Dholakia
f95f5563ea docs: document input/output/total tokens behaviour
Closes https://github.com/BerriAI/litellm/issues/17480
2026-01-20 10:45:47 -08:00
Alexsander Hamir
1377721715
Fix: Handle PostgreSQL cached plan errors during rolling deployments (#19424) 2026-01-20 10:44:31 -08:00
Harshit Jain
1c8bf19f1e
fix(proxy_server): pass search_tools to Router during DB-triggered initialization (#19388) 2026-01-20 09:55:09 -08:00
Harshit Jain
75ee0d126c
Fix/prisma schema permission (#19391)
* fix: add prisma permission issue

* Add test case for prisma generate
2026-01-20 09:53:16 -08:00