litellm/tests/test_litellm/proxy
Alexsander Hamir c7847125c2
[Perf] Embeddings: Use router's O(1) lookup and shared sessions (#16344)
* Refactor proxy embeddings to use shared processor

- allow ProxyBaseLLMRequestProcessing to accept the aembedding route so embeddings requests reuse the base pipeline hooks

- route embeddings requests through base_process_llm_request, sharing logging, hook execution, retries, and header handling with chat/responses

- tighten token array decoding logic by using router deployment lookups and the unified error handler

* Fix: Correctly process embedding requests with token arrays

The `test_embedding_input_array_of_tokens` test was failing due to a regression that caused embedding requests with token arrays to be processed incorrectly. This prevented the `aembedding` function from being called as expected.

This was caused by a combination of three distinct issues:

1.  In `litellm/proxy/common_request_processing.py`, the `function_setup` utility was called with `aembedding` as the `original_function` for embedding routes. This has been corrected to `embedding` to ensure proper request setup.

2.  In `litellm/proxy/proxy_server.py`, a `TypeError` occurred because the `get_deployment` method was called with the `model_name` keyword argument instead of the expected `model_id`. This has been corrected. Additionally, the check for token arrays was improved to validate that all elements in the input subarray are integers.

3.  In `litellm/proxy/litellm_pre_call_utils.py`, the check for the `enforced_params` enterprise feature was too strict. It blocked valid requests even when the `enforced_params` list was empty. The condition has been adjusted to trigger the check only for non-empty lists.

Finally, the `test_embedding_input_array_of_tokens` assertion was updated to be more robust. The previous `assert_called_once_with` was overly strict, causing failures when unrelated internal parameters were added to the function call. The test now first asserts that `aembedding` is called and then separately verifies the `model` and `input` arguments. This makes the test more resilient to future changes without sacrificing its ability to catch regressions.

* test: align proxy embedding assertions

Update the embedding proxy test to match the new request pipeline: keep the data the proxy builds, expect the extra control kwargs, let the post-call hook return the actual response, and assert the normalized 'embeddings' hook type. This proves the refactor still forwards metadata and returns the mocked payload.

* Update proxy exception test

The proxy now forwards additional kwargs (request_timeout, litellm_call_id, litellm_logging_obj) to llm_router.aembedding. The test needs to accept these to match the real call signature and keep validating the error path instead of the kwargs list.

* testing: unsure of this change

I don't remember why I changed this, will revert and see if any tests fail since the manual test isn't failing without it.

* fix: remove unrelated change

This change was not related to the embeddings refactor and actually belonged to a different branch.
2025-11-14 09:21:45 -08:00
..
_experimental/mcp_server fix: allow tool call even when server name prefix is missing (#16425) 2025-11-12 13:50:52 -08:00
anthropic_endpoints
auth fix: allow internal users to access video generation routes (#16472) 2025-11-10 17:44:16 -08:00
client fix delete callbacks 2025-11-06 17:06:34 -08:00
common_utils [Fix] UI - Delete Callbacks Failing (#16473) 2025-11-12 18:43:37 -08:00
db [Fix] Litellm tags usage add request_id (#16111) 2025-11-11 18:53:48 -08:00
experimental/mcp_server
google_endpoints test google endpoints 2025-10-31 20:50:31 -07:00
guardrails test_patch_guardrail_endpoint 2025-11-07 22:09:52 -08:00
health_endpoints Add Langfuse OTEL and SQS to health check (#16514) 2025-11-12 18:25:30 -08:00
hooks fix: Handle multiple rate limit types per descriptor and prevent IndexError (#16039) 2025-10-30 20:12:54 -07:00
image_endpoints test fix 2025-10-17 10:46:42 -07:00
management_endpoints fix: avoid crashing when MCP server record lacks credentials (#16601) 2025-11-13 22:01:11 -08:00
management_helpers [MCP Gateway] Litellm mcp fixes team control (#15304) 2025-10-07 16:48:00 -07:00
middleware
openai_files_endpoint fix issue from pr review 2025-10-01 11:57:10 +08:00
pass_through_endpoints Milvus - Passthrough API support - adds create + read vector store support via passthrough API's (#16170) 2025-11-02 09:47:58 -08:00
public_endpoints Migrate Add Model Fields to backend (#16620) 2025-11-13 21:48:57 -08:00
response_api_endpoints [Fix] - Responses API - add /openai routes for responses API. (Azure OpenAI SDK Compatibility) (#15988) 2025-10-27 19:12:13 -07:00
spend_tracking Pagination for /spend/logs/session/ui endpoint (#16603) 2025-11-13 22:03:00 -08:00
test_configs
ui_crud_endpoints [Infra] Litellm Backend SSO Changes (#16029) 2025-10-30 14:32:08 -07:00
vector_store_endpoints TestIsAllowedToCallVectorStoreEndpoint 2025-11-06 17:02:59 -08:00
__init__.py test fix 2025-10-17 10:46:42 -07:00
test_batch_metadata_none_fix.py
test_caching_routes.py
test_common_request_processing.py [Feat] Cost Tracking - specify a global vendor discount for costs. (#15546) 2025-10-14 20:07:04 -07:00
test_custom_proxy.py fix(ui/): fix routing for custom server root path (#15701) 2025-10-23 13:59:29 -07:00
test_fastapi_offline_routes.py
test_health_check_functions.py
test_litellm_pre_call_utils.py [Fix] Guardrails - Ensure Key Guardrails are applied (#16025) 2025-10-28 16:40:49 -07:00
test_proxy_cli.py feature/add max requests env var 2025-09-28 19:18:56 +01:00
test_proxy_server.py [Perf] Embeddings: Use router's O(1) lookup and shared sessions (#16344) 2025-11-14 09:21:45 -08:00
test_proxy_types.py
test_proxy_utils.py fix(ui/): fix routing for custom server root path (#15701) 2025-10-23 13:59:29 -07:00
test_route_llm_request.py
test_shared_health_check.py Add shared healthcheck 2025-10-09 22:18:05 +05:30
test_spend_log_cleanup.py
test_swagger_chat_completions.py
test_team_member_update.py