Commit Graph

29996 Commits

Author SHA1 Message Date
Cesar Garcia
0ed261b34e
fix(gemini): fix negative text_tokens when using cache with images (#18768)
* fix(gemini): prevent negative text_tokens with explicit caching (#18750)

## Problem
When using Gemini with explicit caching (especially with images),
text_tokens would become negative (e.g., -3327) due to incorrectly
subtracting total cached_tokens from modality-specific text_tokens.

## Root Cause
The old code did:
```python
text_tokens = text_tokens - cached_tokens  # 737 - 4064 = -3327
```

This was wrong because:
- cached_tokens includes ALL modalities (text + image + audio + video)
- text_tokens only contains text
- Subtracting total from specific caused negative values

## Solution
Parse cacheTokensDetails to get per-modality cached token breakdown:
```python
if "cacheTokensDetails" in usage_metadata:
    cached_text_tokens = parse from cacheTokensDetails["TEXT"]
    text_tokens = text_tokens - cached_text_tokens  # Correct!
```

Now we subtract cached tokens per modality, preventing negatives.

## Changes
- Parse cacheTokensDetails field from Gemini response
- Calculate non-cached tokens per modality (text, image, audio)
- Remove incorrect global cached_tokens subtraction
- Add tests for explicit caching and implicit/no caching scenarios

## Testing
- Added test_gemini_cache_tokens_details_no_negative_values
- Added test_gemini_without_cache_tokens_details
- All existing Gemini caching tests pass

Fixes #18750

* feat: add cache_read_input_tokens to Usage object

Addresses reviewer feedback to include cached tokens at the top level
of the Usage object. This aligns with how Anthropic provider handles
cached tokens and ensures they are visible in the final usage response.

* fix: add cacheTokensDetails field to UsageMetadata TypedDict

Fixes mypy error where cacheTokensDetails was being accessed but not defined
in the UsageMetadata TypedDict type definition.
2026-01-12 17:04:33 +05:30
Cesar Garcia
932f06104d
fix: include IMAGE token count in cost calculation for Gemini models (#18876)
* fix: include IMAGE token count as separate usage count and pricing

* fix: remove duplicate TypedDict key and variable definitions

- Remove duplicate input_cost_per_image_token in ModelInfoBase TypedDict
- Remove duplicate image_tokens variable declaration in _calculate_usage()

Fixes MyPy errors:
- types/utils.py:146: Duplicate TypedDict key
- vertex_and_google_ai_studio_gemini.py:1541: Name already defined

---------

Co-authored-by: Thomas Rehn <271119+tremlin@users.noreply.github.com>
2026-01-12 17:03:42 +05:30
Ryan Malloy
c2366194d4
fix(anthropic): prevent dropping thinking when any message has thinking_blocks (#18929)
* fix(anthropic): prevent dropping thinking when any message has thinking_blocks

When Claude returns multiple assistant messages in a conversation, some may
have thinking_blocks while others may not (Claude's behavior varies). The
previous logic only checked the LAST assistant message with tool_calls,
dropping the thinking param if it had no thinking_blocks.

This caused errors when earlier messages still contained thinking_blocks:
"When thinking is disabled, an assistant message cannot contain thinking"

The fix adds a new check: only drop thinking if NO assistant messages
have thinking_blocks. If any message has thinking_blocks, we keep
thinking enabled.

Fixes #18926

* chore: re-trigger CI
2026-01-12 16:29:44 +05:30
Igal Boxerman
6bb63525db
fix(guardrails): fix SerializationIterator error and pass tools to guardrail (#18932)
* fix(generic-guardrail-api): fix SerializationIterator error on multimodal requests

When sending multimodal messages (with images) through the Generic Guardrail API,
the `model_dump()` call fails with "Object of type SerializationIterator is not
JSON serializable" error.

Root cause: The `ChatCompletionAssistantMessage` type defines `content` as an
`Iterable` (not just `List`), and Pydantic's `model_dump()` creates a
`SerializationIterator` for iterables which is not JSON serializable.

Fix: Use `model_dump(mode="json")` which properly converts all iterables to
lists and ensures all complex objects are JSON serializable.

* fix(guardrails): pass tools (function definitions) to guardrail inputs

The unified guardrail handler was not passing the `tools` parameter
(function definitions) from the request to the guardrail inputs.
This meant guardrails could not inspect or validate tool definitions.

Added extraction of `data.get("tools")` and inclusion in the
GenericGuardrailAPIInputs passed to `apply_guardrail()`.

* test(guardrails): add tests for tools passed to guardrail

Added tests verifying that tools (function definitions) are correctly
passed to guardrails in the unified guardrail handler:
- test_tools_passed_to_guardrail
- test_multiple_tools_passed_to_guardrail
- test_no_tools_in_request
- test_tools_and_tool_calls_both_passed
2026-01-12 16:27:54 +05:30
Rayan Pal
ec25e16784
fix: add created_at/updated_at fields to LiteLLM_ProxyModelTable (#18937)
The Pydantic model was missing these timestamp fields that exist in the
Prisma schema. This caused the UI to display "Unknown date" for models
added via the dashboard, even though the data was being stored in the DB.

The fields were being read in get_model_info_with_id() using getattr(),
but since they weren't defined in the Pydantic model, getattr() always
returned None.

Fixes #15487
2026-01-12 16:21:35 +05:30
YutaSaito
dd299c93c2
Merge pull request #18934 from BerriAI/litellm_fix_multiple_mcp_scheduler
[fix] prevent duplicate MCP reload scheduler registration
2026-01-12 07:51:29 +09:00
Yuta Saito
73f66b101c fix: prevent duplicate MCP reload scheduler registration 2026-01-12 07:10:22 +09:00
YutaSaito
c171954629
Merge pull request #18912 from BerriAI/litellm_fix_pangea_guardrail
[fix] respect pangea guardrail default_on during initialization
2026-01-12 05:43:54 +09:00
YutaSaito
a0b7e28fbc
Merge pull request #18882 from BerriAI/litellm_fix_db-migration-log-grep-flaky
[test] stabilize db_migration_disable_update_check test log check
2026-01-12 05:41:46 +09:00
Ishaan Jaffer
f768b698a9 docs fix 2026-01-11 11:57:23 -08:00
Ishaan Jaffer
1271044b06 1.80.15 2026-01-11 11:57:04 -08:00
Marco Vinciguerra
c82d74587d
Add ScrapeGraph MCP server configuration (#18923) 2026-01-11 21:57:46 +05:30
Ishaan Jaffer
7aee1fcb81 OCR test fixes 2026-01-11 08:00:31 -08:00
Ishaan Jaffer
080982e157 1.80.15 2026-01-10 17:50:13 -08:00
Ishaan Jaffer
f36e898cce bump: version 1.80.14 → 1.80.15 2026-01-10 17:49:46 -08:00
Alexsander Hamir
624c9420b9
perf release notes (#18915) 2026-01-10 16:41:54 -08:00
yuneng-jiang
9bcd6b7bc5
Merge pull request #18914 from BerriAI/ui_build_jan10
[Infra] Building UI
2026-01-10 15:49:15 -08:00
yuneng-jiang
c781bcfabb building ui 2026-01-10 15:46:55 -08:00
Ishaan Jaffer
34383ce97b fix route_request 2026-01-10 15:45:55 -08:00
yuneng-jiang
ecaa449f56
Merge pull request #18913 from BerriAI/endpoint_usage_docs
[Docs] Endpoint Usage Docs
2026-01-10 15:43:30 -08:00
yuneng-jiang
039e11cbe5 Endpoint activity writeup 2026-01-10 15:42:20 -08:00
yuneng-jiang
489a4730ab Merge remote-tracking branch 'origin' into endpoint_usage_docs 2026-01-10 15:41:46 -08:00
Yuta Saito
cef087f88b fix: respect pangea guardrail default_on during initialization 2026-01-11 08:33:31 +09:00
yuneng-jiang
728c5297cf
Merge pull request #18911 from BerriAI/litellm_ui_endpoint_usage_trend
[Fix] UI - Endpoint Activity Trend X-Axis and Time Range
2026-01-10 15:31:57 -08:00
Ishaan Jaffer
f1d516f0ae test fixes 2026-01-10 15:15:41 -08:00
yuneng-jiang
ef74f5edb1 fixing test 2026-01-10 15:15:27 -08:00
yuneng-jiang
0b34bfa14c Flipping x axis and removing time range picker 2026-01-10 15:14:17 -08:00
yuneng-jiang
dcae3d569f Flipping x axis and removing time range picker 2026-01-10 15:13:43 -08:00
Ishaan Jaffer
1ebe979eee test_manus_responses_api_with_file_upload 2026-01-10 15:12:00 -08:00
Ishaan Jaffer
58bc308d47 docs levo ai 2026-01-10 15:07:58 -08:00
Ishaan Jaffer
451d240a41 fix 2026-01-10 15:05:38 -08:00
Ishaan Jaffer
64d029bce7 docs fix 2026-01-10 15:04:06 -08:00
Ishaan Jaffer
a4758ab532 QA release notes 2026-01-10 14:57:16 -08:00
Ishaan Jaffer
17995c5192 rc1 writeup 2026-01-10 14:51:56 -08:00
Ishaan Jaffer
9b4f3308fc docs fix 2026-01-10 14:45:55 -08:00
Shivam Rawat
ab37a9872c
fixed _organization_max_budget_check so tests works (#18908) 2026-01-10 14:26:08 -08:00
akraines
7e6d0d7691
fix: Support batch requests with comma-separated models in validate_model_access (#18909)
This fixes a breaking change where batch completion requests with comma-separated
model strings (e.g., 'gpt-3.5-turbo,fake-openai-endpoint') were failing validation.

The validate_model_access function now:
- Detects comma-separated model strings
- Validates each model individually
- Provides clear error messages for inaccessible models in batch requests
- Maintains backward compatibility for single model validation

Fixes test_batch_chat_completions test failure.
2026-01-10 14:23:45 -08:00
Ishaan Jaffer
46cbf673b6 fix code QA 2026-01-10 14:16:55 -08:00
Ishaan Jaffer
3223554905 bump 2026-01-10 14:15:28 -08:00
Ishaan Jaffer
abeec38ada fix manus 2026-01-10 14:14:40 -08:00
Ishaan Jaffer
6a9041e67d Revert "aws fix base"
This reverts commit 225f411abc122087d484fb41d495d340ce5abb80.
2026-01-10 14:08:11 -08:00
Ishaan Jaffer
f37ccca19c docs fix 2026-01-10 14:04:47 -08:00
Ishaan Jaffer
69aaad5e69 Revert "added extraction of top level metadata for custom lables in prometheus callbacks (#18087)"
This reverts commit 14a4a9c031.
2026-01-10 14:00:54 -08:00
Ishaan Jaffer
6b7f114847 test_lists_with_sensitive_keys_are_masked 2026-01-10 13:55:11 -08:00
Ishaan Jaffer
35c636ba97 test_health_check_not_called_when_disabled 2026-01-10 13:55:11 -08:00
Ishaan Jaff
b1c122056f
Revert "fix: Improve error messages and validation for wildcard routing with multiple credentials" (#18907) 2026-01-10 13:49:43 -08:00
Ishaan Jaffer
ff8e9aeb5c Revert "Add support for Vertex AI API keys"
This reverts commit ad501048f3.
2026-01-10 13:39:49 -08:00
Ishaan Jaff
c0cf8bc27d
[Feat] Manus FILES API - Add File upload, get, delete, list (#18904)
* add MANUS get response

* init TwoStepFileUploadRequest

* init TwoStepFileUploadConfig

* add async_create_file to handle 2 step uploads

* init ManusFilesConfig

* add add get_provider_files_config MANUS

* fix validate_environment

* test_manus_files_api_e2e_all_methods

* aws fix base

* init files API MANUS

* test_manus_responses_api_with_file_upload

* mypy lint fixes

* fix BedrockFilesConfig

* manus docs

* docs manus

* mypy lint

* add add fix resposne api utils MANUS
2026-01-10 13:27:54 -08:00
Ishaan Jaffer
3c3ed3bcfb fix resposne api utils 2026-01-10 13:25:31 -08:00
Ishaan Jaffer
e3fe02148d _mask_sequence 2026-01-10 13:20:43 -08:00