Commit Graph

34378 Commits

Author SHA1 Message Date
Alexsander Hamir
c7847125c2
[Perf] Embeddings: Use router's O(1) lookup and shared sessions (#16344)
* Refactor proxy embeddings to use shared processor

- allow ProxyBaseLLMRequestProcessing to accept the aembedding route so embeddings requests reuse the base pipeline hooks

- route embeddings requests through base_process_llm_request, sharing logging, hook execution, retries, and header handling with chat/responses

- tighten token array decoding logic by using router deployment lookups and the unified error handler

* Fix: Correctly process embedding requests with token arrays

The `test_embedding_input_array_of_tokens` test was failing due to a regression that caused embedding requests with token arrays to be processed incorrectly. This prevented the `aembedding` function from being called as expected.

This was caused by a combination of three distinct issues:

1.  In `litellm/proxy/common_request_processing.py`, the `function_setup` utility was called with `aembedding` as the `original_function` for embedding routes. This has been corrected to `embedding` to ensure proper request setup.

2.  In `litellm/proxy/proxy_server.py`, a `TypeError` occurred because the `get_deployment` method was called with the `model_name` keyword argument instead of the expected `model_id`. This has been corrected. Additionally, the check for token arrays was improved to validate that all elements in the input subarray are integers.

3.  In `litellm/proxy/litellm_pre_call_utils.py`, the check for the `enforced_params` enterprise feature was too strict. It blocked valid requests even when the `enforced_params` list was empty. The condition has been adjusted to trigger the check only for non-empty lists.

Finally, the `test_embedding_input_array_of_tokens` assertion was updated to be more robust. The previous `assert_called_once_with` was overly strict, causing failures when unrelated internal parameters were added to the function call. The test now first asserts that `aembedding` is called and then separately verifies the `model` and `input` arguments. This makes the test more resilient to future changes without sacrificing its ability to catch regressions.

* test: align proxy embedding assertions

Update the embedding proxy test to match the new request pipeline: keep the data the proxy builds, expect the extra control kwargs, let the post-call hook return the actual response, and assert the normalized 'embeddings' hook type. This proves the refactor still forwards metadata and returns the mocked payload.

* Update proxy exception test

The proxy now forwards additional kwargs (request_timeout, litellm_call_id, litellm_logging_obj) to llm_router.aembedding. The test needs to accept these to match the real call signature and keep validating the error path instead of the kwargs list.

* testing: unsure of this change

I don't remember why I changed this, will revert and see if any tests fail since the manual test isn't failing without it.

* fix: remove unrelated change

This change was not related to the embeddings refactor and actually belonged to a different branch.
2025-11-14 09:21:45 -08:00
Sameer Kankute
52a42e1728
Add all imagen variants in fal ai in model map (#16579) 2025-11-13 22:31:49 -08:00
Sameer Kankute
13993d6ea3
Add fal-ai/flux/schnell support (#16580) 2025-11-13 22:31:31 -08:00
Krrish Dholakia
266744a5bd docs: add contribution guide for new guardrails 2025-11-13 22:29:42 -08:00
Dmitrii Tunikov
a22b2b0a67
fix(mcp): Fix Gemini conversation format issue with MCP auto-execution (#16592)
When using MCP tools with require_approval='never' and Gemini models,
the follow-up call after tool execution was failing with:

  'Please ensure that function call turn comes immediately after a user
   turn or after a function response turn.'

This was caused by adding an empty assistant message between the user
message and function calls, which violates Gemini's conversation format
requirements.

Changes:
- Only add assistant message to follow-up input if it contains actual content
- Allow function calls to come directly after user messages (as Gemini requires)
- Add explanatory comments about Gemini's format requirements

This fix allows MCP auto-execution to work correctly with Gemini models
while maintaining compatibility with other models.

Fixes: #[issue-number-if-any]
2025-11-13 22:20:50 -08:00
Tomáš Dvořák
eca226286a
fix: parse failed chunks for Groq (#16595)
* fix: parse failed chunks for Groq

Ref: #13960
Signed-off-by: Tomas Dvorak <toomas2d@gmail.com>

* chore: formatting

Signed-off-by: Tomas Dvorak <toomas2d@gmail.com>

---------

Signed-off-by: Tomas Dvorak <toomas2d@gmail.com>
2025-11-13 22:07:15 -08:00
Otavio Brito
aedfe8f7a1
remove generic exception handling (#16599) 2025-11-13 22:03:47 -08:00
yuneng-jiang
379aa7b79a
Pagination for /spend/logs/session/ui endpoint (#16603) 2025-11-13 22:03:00 -08:00
yuneng-jiang
01065a1284
Fixed inconsistent button sizes and variants (#16600) 2025-11-13 22:02:07 -08:00
YutaSaito
331be4f57b
fix: avoid crashing when MCP server record lacks credentials (#16601) 2025-11-13 22:01:11 -08:00
pnookala-godaddy
44bb18a6ba
fix: forward OpenAI organization for image generation (#16607) 2025-11-13 21:51:27 -08:00
yuneng-jiang
792339200a
Migrate Add Model Fields to backend (#16620) 2025-11-13 21:48:57 -08:00
Ishaan Jaff
21ba491656
[UI] Add RunwayML on Admin UI supported models/providers (#16606)
* add runway.png

* add gen4_turbo
2025-11-13 21:46:35 -08:00
sep-grindr
40a9d72be7
fix(ui): remove misleading 'Custom' option mention from OpenAI endpoint tooltips (#16622)
The tooltip for OpenAI api_base select fields incorrectly mentioned 'choose Custom to enter your own' but there was no Custom option available in the dropdown. This fix updates the tooltip text to accurately reflect the available options.

Affected providers:
- OpenAI
- OpenAI_Text
2025-11-13 21:43:57 -08:00
Ishaan Jaff
4a486dc669
[Bug fix] Fixes SambaNova API rejecting requests when message content is passed as a list format (#16612)
* add runwayml_models

* test_call_with_end_user_over_budget

* TestSambanovaContentListHandling

* add _transform_messages for sambanova
2025-11-13 17:03:14 -08:00
Ishaan Jaffer
3feae855bd fix mapped test 2025-11-13 17:00:09 -08:00
Ishaan Jaffer
3c662eadb4 add runwayml/eleven_multilingual_v2 pricing 2025-11-13 16:45:35 -08:00
Ishaan Jaffer
ee8b1cfabc test_call_with_end_user_over_budget 2025-11-13 16:26:02 -08:00
Ishaan Jaffer
32885087c2 add runwayml_models 2025-11-13 16:25:27 -08:00
Ishaan Jaffer
8be374c98d fix docker v 2025-11-13 16:12:48 -08:00
yuneng-jiang
cfc44ea279 Remove debug statements 2025-11-13 15:51:49 -08:00
yuneng-jiang
09afe14a5f Addressing linting issues 2025-11-13 15:49:35 -08:00
yuneng-jiang
d62dfd4be2 Merge remote-tracking branch 'origin' into litellm_org_usage 2025-11-13 15:26:59 -08:00
Ishaan Jaff
124ba463f8
[Feat] RunwayML - Add support for /audio/speech eleven_multilingual_v2 endpoint (#16604)
* init RunwayMLTextToSpeechConfig

* add RunwayMLTextToSpeechConfig

* add  RunwayMLTextToSpeechConfig

* test_runwayml_tts_async

* runway ml speech

* fix voices

* fix test

* docs runway lm

* add runwayml here

* fix RunwayMLTextToSpeechConfig

* test_openai_voice_mapping_to_runwayml
2025-11-13 14:32:09 -08:00
Nicholas Couture
4be372eb48
fix: support Anthropic tool_use and tool_result in token counter (#16351)
* fix: support Anthropic tool_use and tool_result in token counter

* refactor(token_counter): add dynamic field inference for Anthropic content blocks

* test: Add additional tests

* make format

* Fix lint error

* Fix mypy narrow type lint errors
2025-11-13 14:30:46 -08:00
Ishaan Jaff
911a009869
[Docs] LiteLLM Quick start - show how model resolution works (#16602)
* docs nderstanding Model Configuration

* docs fix
2025-11-13 13:28:01 -08:00
Ishaan Jaff
7133488282
[Feat] VertexAI - Add BGE Embeddings support (#16033)
* Support for Custom Vertex AI Models via PSC Endpoint with api_base (#15953)

* Support for Custom Vertex AI Models via PSC Endpoint with api_base

* Add docs related psc

* remove not needed files

* remove print statemnt

* fix mypy errors

* add TextEmbeddingBGEInput

* add VertexBGEConfig

* add BGE handling

* test_vertex_ai_bge_embedding_with_custom_api_base

* fix request transform vertex BGE

* test_vertex_ai_bge_embedding_with_custom_api_base

* tes BGE

* test_is_bge_model_detection

* docs cleanup

* handling BGE URL

* fix VertexBGEConfig

* test_vertex_ai_bge_with_endpoint_id_pattern

* docs vertex BGE

* docs

* docs fix

* fix VertexAIModelRoute

* from ..common_utils import VertexAIError, get_vertex_base_model_name
add

* fix VertexAIGemmaModels

* fix get_vertex_base_model_name

* test_vertex_ai_bge_psc_endpoint_url_construction

---------

Co-authored-by: Sameer Kankute <sameer@berri.ai>
2025-11-13 12:41:00 -08:00
YutaSaito
8e0b66a814
fix: exclude unauthorized MCP servers from allowed server list (#16551)
* fix: exclude unauthorized MCP servers from allowed server list

* fix: test after resolving merge conflicts
2025-11-13 12:33:54 -08:00
Sameer Kankute
ea80510f78
[Feat] Day-0, Add gpt-5.1 and gpt-5.1-codex family support (#16598)
* Add day 0 support for gpt-5.1 models

* Add gpt-5.1-codex day 0 support

* update pricing values
2025-11-13 10:55:54 -08:00
Jón Levy
555d7b8be8
feat(bedrock): Add bearer token authentication support for AgentCore (#16556) 2025-11-13 08:17:36 -08:00
Lucas Sugi
e1c607e22a
feat: Add headers to VLLM Passthrough requests [Log success Events] (#16532) 2025-11-12 19:46:34 -08:00
Cesar Garcia
491f57a349
feat: Add support for reasoning_effort="none" for Gemini models (#16548)
Implements support for reasoning_effort="none" parameter for Gemini models,
providing significant cost savings (up to 96% cheaper) by disabling thinking
budget while maintaining response quality.

Changes:
- Added "supports_reasoning": true to gemini-2.0-flash-thinking-exp-01-21 in model config
- Implemented mapping for reasoning_effort="none" to thinkingConfig {thinkingBudget: 0, includeThoughts: false}
- Added unit test to verify the mapping works correctly

Performance impact:
- Without reasoning_effort: ~313 tokens
- With reasoning_effort="none": ~12 tokens (96% cheaper)

Closes #16420

Co-authored-by: Krish Dholakia <krrishdholakia@gmail.com>
2025-11-12 19:41:07 -08:00
Cesar Garcia
c017f665e0
docs(openai): Document reasoning_effort summary field options (#16549)
Related to PR #16210 which fixed automatic summary field addition

Changes:
- Document reasoning_effort string vs dict formats
- Add summary field options (auto, detailed, concise)
- Add table of supported reasoning_effort values by GPT-5 model
- Clarify model-specific support and limitations
- Note that summary field requires org verification

The previous implementation automatically added summary field causing
400 errors for unverified orgs. Now users can opt-in by passing
reasoning_effort as dict with explicit summary field.
2025-11-12 19:40:11 -08:00
Cesar Garcia
049d45ea90
fix(gemini): Preserve non-ASCII characters in function call arguments (#16550)
Fixes #16533

Before this fix, non-ASCII characters (Japanese, Spanish, Chinese, etc.)
in function call arguments were being escaped as Unicode sequences.

Example:
- Before: "やあ" → "\u3084\u3042"
- After: "やあ" → "やあ" (preserved)

Changes:
- Add ensure_ascii=False to json.dumps() in _transform_parts()
- Add test for Japanese and Spanish Unicode character preservation

This is not a breaking change as both formats are equivalent in JSON.
The fix improves readability and aligns with OpenAI's behavior.
2025-11-12 19:01:38 -08:00
Sameer Kankute
018bd2e039
Add Gemini image edit support (#16430)
* Add gemini image edit support

* fix lint errors

* fix lint errors

* fix lint errors

* Add docs
2025-11-12 18:48:27 -08:00
yuneng-jiang
8bf491c939
[Fix] /spend/logs/ui Access Control (#16446)
* RBAC for /spend/logs/ui

* Addressing comments
2025-11-12 18:44:21 -08:00
yuneng-jiang
cb27d6c456
[Fix] UI - Delete Callbacks Failing (#16473)
* Temp commit for branch switching

* Created normalize callback name util function and tests
2025-11-12 18:43:37 -08:00
Sameer Kankute
92bd12c862
Fix raising wrong 429 error on wrong exception (#16482)
* fix raising wrong 429 error on wrong exception

* remove double re import
2025-11-12 18:41:12 -08:00
Ishaan Jaffer
78c169a524 docs fix 2025-11-12 18:30:01 -08:00
Sameer Kankute
394da34a0b
Add all gemini image models support in image generation (#16526) 2025-11-12 18:26:44 -08:00
yuneng-jiang
898f15c33c
Add Langfuse OTEL and SQS to health check (#16514) 2025-11-12 18:25:30 -08:00
yuneng-jiang
c23b2c5023
Config Guardrails should not be deletable from table (#16540) 2025-11-12 18:23:31 -08:00
yuneng-jiang
35eece2b7b Merge remote-tracking branch 'origin' into litellm_org_usage 2025-11-12 18:22:44 -08:00
yuneng-jiang
94da5c076c Add Daily Org Spend Table, read path, and write path 2025-11-12 18:22:26 -08:00
Ishaan Jaff
b30439257b
[Feat] Add RunwayML Img Gen API support (#16557)
* TestRunwaymlImageGeneration

* fix RUNWAYML

* rename

* fix rename

* get_runwayml_image_generation_config

* get_runwayml_image_generation_config

* TestRunwaymlImageGeneration

* add RUNWAYML_POLLING_TIMEOUT

* fix rnwayml transform img gen

* runwayml_image_cost_calculator

* runwayml_image_cost_calculator

* docs runwayml

* fix runwayML polling

* test_get_first_default_fallback
2025-11-12 18:20:14 -08:00
Benjamin Chrobot
1393900c22
[Docs] Fix code block indentation for fallbacks page (#16542) 2025-11-12 18:08:52 -08:00
yuneng-jiang
f0c83e183e
SSO Modal Cosmetic Changes (#16554) 2025-11-12 17:55:05 -08:00
Krrish Dholakia
a05ad46394 feat: add litellm cloud self-serve to docs 2025-11-12 17:13:14 -08:00
Ishaan Jaffer
d7824d076e docs project management 2025-11-12 16:24:23 -08:00
Ishaan Jaffer
5cf8f06b17 add to _get_spend_logs_metadata 2025-11-12 15:56:32 -08:00