* propagate model id on errors too
* make it work for messages and streaming
* fix
* cleanup
* cleanup
* final
* cleanup
* clean up method name and fix responses api streaming
* remove comment
* Add support for vector store files endpoints (#16490)
* Add base code for vector store integration
* fix azure related tests and linting error
* fix mypy errors
* Add vector store files documentation
* fix mapped tests
* Add bytedance and ideogram support in fal ai (#16636)
* Add fal ai flux pro v1.1 support (#16578)
* Add fal ai flux pro v1.1 support
* Add tests and docs
---------
Co-authored-by: Ishaan Jaffer <ishaanjaffer0324@gmail.com>
* Refactor proxy embeddings to use shared processor
- allow ProxyBaseLLMRequestProcessing to accept the aembedding route so embeddings requests reuse the base pipeline hooks
- route embeddings requests through base_process_llm_request, sharing logging, hook execution, retries, and header handling with chat/responses
- tighten token array decoding logic by using router deployment lookups and the unified error handler
* Fix: Correctly process embedding requests with token arrays
The `test_embedding_input_array_of_tokens` test was failing due to a regression that caused embedding requests with token arrays to be processed incorrectly. This prevented the `aembedding` function from being called as expected.
This was caused by a combination of three distinct issues:
1. In `litellm/proxy/common_request_processing.py`, the `function_setup` utility was called with `aembedding` as the `original_function` for embedding routes. This has been corrected to `embedding` to ensure proper request setup.
2. In `litellm/proxy/proxy_server.py`, a `TypeError` occurred because the `get_deployment` method was called with the `model_name` keyword argument instead of the expected `model_id`. This has been corrected. Additionally, the check for token arrays was improved to validate that all elements in the input subarray are integers.
3. In `litellm/proxy/litellm_pre_call_utils.py`, the check for the `enforced_params` enterprise feature was too strict. It blocked valid requests even when the `enforced_params` list was empty. The condition has been adjusted to trigger the check only for non-empty lists.
Finally, the `test_embedding_input_array_of_tokens` assertion was updated to be more robust. The previous `assert_called_once_with` was overly strict, causing failures when unrelated internal parameters were added to the function call. The test now first asserts that `aembedding` is called and then separately verifies the `model` and `input` arguments. This makes the test more resilient to future changes without sacrificing its ability to catch regressions.
* test: align proxy embedding assertions
Update the embedding proxy test to match the new request pipeline: keep the data the proxy builds, expect the extra control kwargs, let the post-call hook return the actual response, and assert the normalized 'embeddings' hook type. This proves the refactor still forwards metadata and returns the mocked payload.
* Update proxy exception test
The proxy now forwards additional kwargs (request_timeout, litellm_call_id, litellm_logging_obj) to llm_router.aembedding. The test needs to accept these to match the real call signature and keep validating the error path instead of the kwargs list.
* testing: unsure of this change
I don't remember why I changed this, will revert and see if any tests fail since the manual test isn't failing without it.
* fix: remove unrelated change
This change was not related to the embeddings refactor and actually belonged to a different branch.
* Add v1 cut of container api
* fix lint errors
* Add proxy support to container apis & logging support (#16049)
* Add proxy support to container apis
* Add logging support
* Add cost tracking support for containers and documentation
* Add new constant documentation
* Add container cost in model map
* fix failing azure tests
* Update tests based on model map changes
* fix model map tests
* fix model map tests
* Container modeshould be container
* Container tests fix
* Merge branch 'main' into litellm_sameer_oct_staging_2
---------
Co-authored-by: Ishaan Jaffer <ishaanjaffer0324@gmail.com>
* feat(responses_id_security.py): encrypt response.id - prevent user A from retrieving user B's response
additional security for retrievals on shared accounts
Closes LIT-1307
* feat(responses_id_security.py): allow admin to disable responses id security check
* test: add initial unit testing
* feat(responses_id_security.py): add streaming support
* docs: document new param
* docs: document new param
* feat(responses_id_security.py): add team id checks - ensure it works for service accounts
prevent service accounts keys from different teams from accessing each other's responses
more secure
* test: add unit testing
* fix: fix linting error
* Addd v2/chat support for cohere
* fix streaming
* Use v2_transformation for logging passthrough:
* Use v2_transformation for logging passthrough:
* Add test for checking if document and citation_options is getting passed
* Update the cohere model
* Add cost tracking for vertex ai passthrough batch jobs
* Add full passthrough support
* refactor code according to the comments
* Add passthrough handler
* remove invalid params
* Updated documentation
* Updated documentation
* Updated documentation
* Correct the import
* Add openai videos generation and retrieval support
* add retrieval endpoint
* Add docs
* Add imports
* remove orjson
* remove double import
* fix openai videos format
* remove mock code
* remove not required comments
* Add tests
* Add tests
* Add other video endpoints
* Fix cost calculation and transformation
* Fixed mypy tests
* remove not used imports
* fix documentation for get batch req (#15742)
* Add grounding info to responses API (#15737)
* Add grounding info to responses API
* fix lint errors
* Use typed objects for annotations
* Use typed objects for annotations
* fix mypy error
* Litellm fix json serialize alreting 2 (#15741)
* fix json serializable error for alerts
* Add test
* fix mypt errors
* fix mypt errors
* Add Qwen3 imported model support for AWS Bedrock (#15783)
* Add qwen imported model support
* fix mypy errors
* fix empty user message error (#15784)
* fix typed dict for list
* Add azure supported videos endpoint
* fix mapped tests
* add azure sora models to model map
* Add OpenAI video generation and content retrieval support (#15745)
* Add openai videos generation and retrieval support
* add retrieval endpoint
* Add docs
* Add imports
* remove orjson
* remove double import
* fix openai videos format
* remove mock code
* remove not required comments
* Add tests
* Add tests
* Add other video endpoints
* Fix cost calculation and transformation
* Fixed mypy tests
* remove not used imports
* fix typed dict for list
* fix mypy errors
* move directory
* make v2 chat default
* Fix mypy tests
* Fix mypy tests
* Fix mypy tests
* Fix mypy tests
* Revert "Add Azure Video Generation Support with Sora Integration"
* refactor videos repo
* add test
* Add azure openai videos support
* Add azure openai videos support
* Add router endpoint support for videos
* fix mypy error
* add azure models
* fix mapped test
* fix mypy error
* Add proxy router test
* Add proxy router test
* remove deprecated model name from tests
* fix import error
* fix import error
* Add gaurdrail integration in videos endpoint
* Add logging support for videos endpoint
* Add final documentation supporting videos integration
* fix model name and document input
* Update literals to avoid mypy errors
* Remove unused imports and print statements
* revert guardrail support for video generation and video remix
* revert guardrail support for video generation and video remix
* Fix failing mapped and llm translation tests
This commit implements a complete, end-to-end fix for the native Gemini API translation feature, allowing requests to be correctly routed to other model providers via `model_group_alias`.
The original implementation was broken, causing `systemInstruction` and `tools` to be dropped from requests. This was resolved by refactoring the Gemini endpoint to use a dedicated translation path, similar to the Anthropic adapter.
Additionally, this commit hardens the streaming response adapter to correctly handle tool calls generated by the newly-fixed request path. Key improvements to the response handling include:
- Replaced the fragile `id`-based tool call tracking with a robust `index`-based accumulation logic.
- Fixed a memory leak and improved logging in the stream finalization process.
- Prevented empty, non-compliant chunks from being sent to the client during tool call streaming.
- Optimized the accumulator to skip and log superfluous empty chunks sent by some models.
* cancel upstream on client disconnect
* add comments
* add test
* set timeout in constraints.py
* Guard against missing 'type' key
* update dependency to fix uvicorn bugs
* feat: initial commit with prompt management support on pre-call hooks
allows prompt templates to work before assigning specific models
* feat: initial logic for independent prompt management settings
* feat(proxy_server.py): working logic for loading in the prompt templates from config yaml
allows creating an independent 'prompts' section in the config yaml
* feat(prompt_registry.py): working e2e custom prompt templates with guardrails and models
* refactor(prompts/): move folder inside proxy folder
easier management for prompt endpoints
* feat(prompt_endpoints.py): working `/prompt/list` endpoint
returns all available prompts on proxy
* feat(key_management_endpoints.py): support storing 'prompts' in key metadata
allows giving keys access to specific prompts
* feat(prompt_endpoints.py): enable key-based access to /prompts/list
ensures key can only see prompts it has access to
* fix(init_prompts.py): fix linting error
* fix: fix ruff check
* fix(proxy/_types.py): add 'prompts' to newteamrequest
* fix(litellm_logging.py): update logged message with scrubbed value
* feat: initial commit with prompt management support on pre-call hooks
allows prompt templates to work before assigning specific models
* feat: initial logic for independent prompt management settings
* feat(proxy_server.py): working logic for loading in the prompt templates from config yaml
allows creating an independent 'prompts' section in the config yaml
* feat(prompt_registry.py): working e2e custom prompt templates with guardrails and models
* refactor(prompts/): move folder inside proxy folder
easier management for prompt endpoints
* fix: fix linting error
* fix: fix check
* fix(google_genai/adapters/transformation.py): enable calling non-googlegenai models via streaming
Fixes https://github.com/BerriAI/litellm/issues/12562
* test(test_openai.py): add unit test asserting streaming works as expected
* fix(streaming_handler.py): don't return stream options in clientside response
* test: add unit test
* fix(custom_guardrail.py): fix start time
don't start until function actually called
leading to incorrect high latency numbers
* fix(custom_guardrail.py): fix start time
* build(model_prices_and_context_window.json): remove 'supports_tool_choice' for specific mistral models
Closes https://github.com/BerriAI/litellm/issues/11750
* feat: initial commit adding cleaner ui for azure text moderation guardrails
* feat(guardrail_endpoints.py): add discoverable guardrail configs and improve converting base model to dict with types
* fix(guardrail_provider_fields.tsx): render from api endpoint correctly
* fix(guardrail_provider_fields.tsx): cleanup
* refactor(guardrail_endpoints.py): refactor to handle dictionaries with literal - allows multiselect
* feat(ui/): render dictionary with known keys correctly
* feat(ui/): render optional params on separate page
* style(ui/): style improvements to rendering optional params on the UI
* feat(azure/prompt_shield.py): add azure prompt shield back on UI
* fix(add_guardrail_form.tsx): fix form to handle updated api
* fix(guardrail_optional_params.tsx): ensure values are nested correctly for writing to api
* fix: fix linting error
* feat(text_moderation.py): handle str to int conversion
* fix(guardrail_info.tsx): only render pii settings if guardrail is presidio
* fix(guardrail_info.tsx): add guardrail specific fields to update settings
allows updating guardrail fields (e.g. severity threshold) post-create
* fix(guardrail_endpoints.py): set guardrail_id in guardrail object
ensures duplicate objects not created on guardrail update
* fix(guardrail_endpoints.py): allow provider specific fields to be updated on patch update
* refactor(guardrail_endpoints.py): remove duplicate info endpoint
* fix(guardrail_endpoints.py): mask sensitive keys on returning via guardrail `/info`
Prevent leaking keys
* fix(guardrail_optional_params.tsx): return numerical input when numerical component used
fixes issue where output was a str
* fix(guardrail_optional_params.tsx): render dict keys correctly
* fix(text_moderation.py): fix severity by category check
* fix(proxy/utils.py): check if guardrail should run for post call streaming hook
Prevents invalid guardrails from running if not requested
* test: fix import
* fix: fix linting error
* test: update test
* fix: fix tests
* fix: fix code qa errors
* fix(guardrail_endpoints.py): set max depth for function
* test: update recursive_detector.py
* test: update list
* build: merge main
* fix: fix ruff check errors
* refactor(passthrough_endpoints-success-handler): refactor llm passthrough logging logic
isolate the llm translation work to enable cost tracking on sdk
* feat: initial implementation of passthrough SDK cost calculation
enables bedrock passthrough cost tracking to work
* feat(cost_calculator.py): working cost calculation for bedrock passthrough
* feat(litellm_logging.py): consider allm_passthrough in cost tracking
allows async calls (e.g. via proxy) to work
* feat(bedrock/passthrough): working event stream decoding for bedrock passthrough calls + logging instrumentation for passthrough sdk calls (log on stream completion)
Enables bedrock streaming cost calculation
* feat(litellm_logging.py): support streaming passthrough cost tracking
* feat(passthrough/main.py): working async streaming cost calculation
Closes https://github.com/BerriAI/litellm/issues/11359
* feat(proxy_server.py): fix passthrough routing when llm router enabled
* feat: further fixes
* feat(bedrock/): working bedrock passthrough cost tracking (non-streaming)
* feat(litellm_logging.py): working usage tracking for bedrock passthrough calls
ensures tokens are logged
* feat(bedrock/passthrough): add converse passthrough cost tracking support
* feat(base_llm/passthrough): remove redundant function
* refactor(litellm_logging.py): refactor function to be below 50 LOC
* test: update test
* test: remove redundant test
* feat: initial commit adding bedrock support via the new sdk passthrough logic
ensures correct sequencing of tasks (pre call checks etc. can run before signing request)
* fix(route_llm_requests.py): passthrough to allm_passthrough_route if no model found
* feat(bedrock/passthrough): working bedrock passthrough via sdk support
* fix(passthrough/main.py): re-add data and json
* feat(passthrough/main): support async passthrough calls to bedrock
* feat(passthrough/main.py): async streaming + completion support
* feat(llm_passthrough_endpoints.py): migrate bedrock passthrough calls to to new bedrock passthrough sdk
Enables calls to work correctly
* fix: fix linting errors
* test: update test
* init litellm google gen ai methods
* feat init structure of functions for generate content
* add init
* add BaseGoogleGenAIGenerateContentConfig
* add generate_content_handler
* add get_provider_google_genai_generate_content_config
* fixes for generate content
* add get_vertex_ai_project etc to base
* use VertexBase
* fixes for BaseGoogleGenAIGenerateContentConfig
* working validate env for google gemini
* feat - add transform google response
* fixes for transform_generate_content_request
* fix get_supported_generate_content_optional_params
* add BaseGoogleGenAITest
* working e2e test
* fixes init config
* use correct types
* fix test for google gen ai
* fix types
* add sync_get_auth_token_and_url
* fixes for transform
* add llm http handler for google
* working non-streaming google endpoints
* add BaseGoogleGenAIGenerateContentStreamingIterator
* add GoogleGenAIGenerateContentStreamingIterator
* fix working sync stream
* fixes for litellm logging obj
* working async streaming
* add google gen ai types
* fix - required imports
* fix readme
* fix deps
* fix deps
* fix ruff code QA checks
* fix linting
* fixes TYPE_CHECKING
* fixes for typing
* add google gemini methods to litellm router
* [Feat] Add initial endpoints for using Gemini SDK (gemini-cli) with LiteLLM (#12040)
* init with google endpoints
* add Depends
* feat - add gemini endpoints
* google_generate_content
* fix init
* fixes import
* fixes for streaming
* fixes for sync/async
* working streaming with google gemini cli
* add google endpoints to llm api routes
* add VertexAIGoogleGenAIConfig
* use aiter_bytes
* use common request for streaming data
* re-use logic for anthropic streaming
* add GoogleAIStudioDataGenerator
* feat(codestral/completion): return litellm latency overhead for codestral
enables easier debugging of latency issues
* fix(types/utils.py): support _response_ms on hidden params model dump
Fixes issue where 'x-litellm-overhead-duration-ms' wasn't being returned on text c
ompletion calls
* fix(types/utils.py): add '__contains__' support for chatcompletiondeltatool call
Fixes https://github.com/BerriAI/litellm/issues/7099
* fix: fix linting error
* fix: fix linting error
For streaming requests, the remote request is triggered and
first chunk inspected for an error code as emitted by
async_data_generator.
Potential fix for #9035
* Add LiteLLM Managed file support for `retrieve`, `list` and `cancel` finetuning jobs (#11033)
* feat: initial commit adding managed file support to fine tuning endpoints
* feat(fine_tuning/endpoints.py): working call to openai finetuning route
Uses litellm managed files for finetuning api support
* feat(fine-tuning/main.py): refactor to use LiteLLMFineTuningJob pydantic object
includes 'hidden_params'
* fix: initial commit adding unified finetuning id support
return a unified finetuning id we can use to understand which deployment to route the ft request to
* test: fix test
* feat(managed_files.py): return unified finetuning job id on create finetuning job
enables retrieve, delete to work with litellm managed files
* feat(managed_files.py): support managed files for cancel ft job endpoint
* feat(managed_files.py): support managed files for cancel ft job endpoint
* feat(fine_tuning_endpoints/endpoints.py): add managed files support to list finetuning jobs
* feat(finetuning_endpoints/main): add managed files support for retrieving ft job
Makes it easier to control permissions for ft endpoint
* LiteLLM Managed Files - Enforce validation check if user can access finetuning job (#11034)
* feat: initial commit adding managed file support to fine tuning endpoints
* feat(fine_tuning/endpoints.py): working call to openai finetuning route
Uses litellm managed files for finetuning api support
* feat(fine-tuning/main.py): refactor to use LiteLLMFineTuningJob pydantic object
includes 'hidden_params'
* fix: initial commit adding unified finetuning id support
return a unified finetuning id we can use to understand which deployment to route the ft request to
* test: fix test
* feat(managed_files.py): return unified finetuning job id on create finetuning job
enables retrieve, delete to work with litellm managed files
* feat(managed_files.py): support managed files for cancel ft job endpoint
* feat(managed_files.py): support managed files for cancel ft job endpoint
* feat(fine_tuning_endpoints/endpoints.py): add managed files support to list finetuning jobs
* feat(finetuning_endpoints/main): add managed files support for retrieving ft job
Makes it easier to control permissions for ft endpoint
* feat(managed_files.py): store create fine-tune / batch response object in db
storing this allows us to filter files returned on list based on what user created
* feat(managed_files.py): Ensures users can't retrieve / modify each others jobs
* fix: fix check
* fix: fix ruff check errors
* test: update to handle testing
* fix: suppress linting warning - openai 'seed' is none on azure
* test: update tests
* test: update test