litellm/tests
Sameer Kankute 50df072d95
feat: add weighted-routing failover (#27980)
* Feat: Add Weighted-Routing Failover

* test(router): cover weighted failover helper functions

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(router): align weighted failover deployment list type with mypy

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(router): address greptile review on weighted failover

- Narrow exception swallowing in `_maybe_run_weighted_failover` to
  `openai.APIError` so model failures defer to the regular fallback
  while programming bugs (AttributeError/KeyError/TypeError) surface.
- Note async-only limitation of `enable_weighted_failover` in the
  Router constructor docstring.
- Make the weighted distribution test less flaky (1000 iterations,
  looser bound) and make the non-simple-shuffle test deterministic by
  failing both deployments instead of relying on the latency strategy's
  first pick.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(router): ensure weighted failover metadata persists in kwargs

The previous `kwargs.setdefault(metadata_variable_name, {}) or {}` returned
a brand-new dict whenever the existing metadata was falsy (empty dict or
None), so writes to `_failover_excluded_ids` never made it back into
`kwargs`. Multi-hop weighted failover then re-selected previously failed
deployments and exhausted `max_fallbacks` prematurely.

Explicitly assign a fresh dict into kwargs when metadata is missing so
mutations are visible to subsequent failover hops.

Co-authored-by: Yassin Kortam <yassin@berri.ai>

* test(router): regression for weighted failover metadata persistence

Asserts kwargs["metadata"]["_failover_excluded_ids"] is populated after
_maybe_run_weighted_failover, proving the metadata dict written by the
helper is the same object that lives in kwargs (no disconnected copy).
Pairs with the prior fix that replaced `setdefault(..., {}) or {}` with
an explicit get/assign so writes survive across hops.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(router): harden weighted failover error/state handling

- Catch RouterRateLimitError (ValueError) alongside openai.APIError in
  _maybe_run_weighted_failover so an exhausted intra-group retry falls
  through to the regular cross-group fallback path instead of bubbling
  out and bypassing configured fallbacks.
- Stop mutating the shared input_kwargs dict; build a local copy with
  the weighted-failover keys so the entry (with _excluded_deployment_ids)
  cannot leak into later fallback paths reading the same dict.
- _get_excluded_filtered_deployments now returns an empty list when the
  exclusion filter removes every healthy deployment, instead of falling
  back to the original list. The original-list behavior risked re-picking
  the just-failed deployment; callers already handle the empty case by
  raising their no-deployments error, which weighted failover now catches
  and converts into a normal cross-group fallback.

Co-authored-by: Yassin Kortam <yassin@berri.ai>

* fix(router): fall through to rpm/tpm when total weight is zero

When the weight metric's total is zero (e.g. after weighted-failover
exclusion leaves only zero-weight backups), continue to the next metric
(rpm/tpm) instead of returning a uniform random pick immediately. This
lets rpm/tpm still drive routing when present, and only falls back to
the uniform random pick at the end if no metric provides a positive
total weight.

Co-authored-by: Yassin Kortam <yassin@berri.ai>

* fix(router): skip weighted failover when remaining deployments are all in cooldown

_maybe_run_weighted_failover was computing 'remaining' from all_deployments
(every deployment in the model group, including those in cooldown). This meant
that when all non-excluded deployments were in cooldown the method still invoked
run_async_fallback unnecessarily, which propagated into async_get_healthy_deployments,
found no eligible deployments, and raised RouterRateLimitError — only safely
caught thanks to the earlier exception-broadening fix.

The fix: before computing 'remaining', fetch the current cooldown set via
_async_get_cooldown_deployments and subtract it from all_ids. This allows
_maybe_run_weighted_failover to return None immediately (skipping the
run_async_fallback call entirely) when every non-failed deployment is in cooldown,
letting the caller fall through to the correct cross-group fallback path without
the wasteful extra round-trip.

Tests added:
- unit: _maybe_run_weighted_failover returns None without calling run_async_fallback
  when all remaining deployments are in cooldown
- unit: _maybe_run_weighted_failover still calls run_async_fallback when at least
  one healthy (non-cooldown) deployment is available
- integration: end-to-end fallthrough to cross-group fallback when remaining
  deployments are in cooldown

Co-authored-by: Sameer Kankute <Sameerlite@users.noreply.github.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Yassin Kortam <yassin@berri.ai>
Co-authored-by: Sameer Kankute <Sameerlite@users.noreply.github.com>
2026-05-15 17:28:54 +00:00
..
agent_tests
audio_tests test(vcr): classify cache verdicts, detect live calls, surface cost leaks 2026-05-13 00:31:47 +00:00
basic_proxy_startup_tests
batches_tests fix(vertex-ai): fix zero cost/usage on completed Vertex AI batch jobs (#27912) 2026-05-15 04:47:02 -07:00
benchmarks
code_coverage_tests feat(audio_transcription): add NVIDIA Riva STT provider (#27185) 2026-05-05 17:17:51 -07:00
documentation_tests
enterprise fix(tests): use canonical litellm_enterprise import path (#27699) 2026-05-12 12:32:57 -07:00
guardrails_tests test(vcr): classify cache verdicts, detect live calls, surface cost leaks 2026-05-13 00:31:47 +00:00
image_gen_tests Merge pull request #27795 from BerriAI/litellm_vcr-cache-observability-and-fixes-c5bc 2026-05-14 13:51:16 -07:00
litellm Add new chat model metadata (#27313) 2026-05-06 15:15:21 -07:00
litellm_core_utils
litellm_utils_tests Merge pull request #27795 from BerriAI/litellm_vcr-cache-observability-and-fixes-c5bc 2026-05-14 13:51:16 -07:00
litellm-proxy-extras
llm_responses_api_testing test(vcr): classify cache verdicts, detect live calls, surface cost leaks 2026-05-13 00:31:47 +00:00
llm_translation fix(vcr): aggregate worker stats on the controller so the session summary actually renders under xdist 2026-05-13 07:24:32 +00:00
load_tests
local_testing test(vcr): drop dead 'from respx import MockRouter' imports 2026-05-13 00:32:03 +00:00
logging_callback_tests test(vcr): classify cache verdicts, detect live calls, surface cost leaks 2026-05-13 00:31:47 +00:00
mcp_tests feat: litellm shin agent oss staging 05 10 2026 (#27631) 2026-05-11 20:31:43 -07:00
multi_instance_e2e_tests
ocr_tests test(vcr): classify cache verdicts, detect live calls, surface cost leaks 2026-05-13 00:31:47 +00:00
old_proxy_tests/tests
openai_endpoints_tests
otel_tests fix(tests): swap dall-e to gpt-image-1 after openai deprecation 2026-05-12 16:55:18 -07:00
pass_through_tests
pass_through_unit_tests test(vcr): mark Bedrock prompt-caching cross-call tests VCR-incompatible 2026-05-13 01:19:03 +00:00
proxy_admin_ui_tests
proxy_e2e_anthropic_messages_tests
proxy_security_tests
proxy_unit_tests fix(managed_batches): convert raw output_file_id to managed ID in CheckBatchCost poller (#27984) 2026-05-15 04:41:38 -07:00
router_unit_tests Merge pull request #27795 from BerriAI/litellm_vcr-cache-observability-and-fixes-c5bc 2026-05-14 13:51:16 -07:00
scim_tests
search_tests test(vcr): classify cache verdicts, detect live calls, surface cost leaks 2026-05-13 00:31:47 +00:00
spend_tracking_tests
store_model_in_db_tests
test_litellm feat: add weighted-routing failover (#27980) 2026-05-15 17:28:54 +00:00
unified_google_tests test(vcr): classify cache verdicts, detect live calls, surface cost leaks 2026-05-13 00:31:47 +00:00
vector_store_tests
windows_tests
__init__.py
_flush_vcr_cache.py
_vcr_conftest_common.py fix(vcr): aggregate worker stats on the controller so the session summary actually renders under xdist 2026-05-13 07:24:32 +00:00
_vcr_redis_persister.py test: add 24hr Redis-backed VCR cache to additional test suites (#27159) 2026-05-05 15:13:31 -07:00
eval_swe_bench.py
gettysburg.wav
large_text.py
openai_batch_completions.jsonl
README.MD
test_budget_management.py
test_callbacks_on_proxy.py
test_config.py
test_debug_warning.py
test_default_encoding_non_root.py
test_end_users.py
test_entrypoint.py
test_fallbacks.py
test_gpt5_azure_temperature_support.py
test_health.py fix(tests): swap dall-e to gpt-image-1 after openai deprecation 2026-05-12 16:55:18 -07:00
test_keys.py fix(tests): swap dall-e to gpt-image-1 after openai deprecation 2026-05-12 16:55:18 -07:00
test_litellm_proxy_responses_config.py
test_logging.conf
test_models.py
test_new_vector_store_endpoints.py
test_openai_endpoints.py fix(tests): swap dall-e to gpt-image-1 after openai deprecation 2026-05-12 16:55:18 -07:00
test_organizations.py
test_otel_thread_leak.py
test_passthrough_endpoints.py
test_presidio_latency.py
test_proxy_server_non_root.py
test_ratelimit.py
test_resource_cleanup.py
test_service_logger_otel.py
test_spend_logs.py
test_team_logging.py
test_team_members.py
test_team.py
test_users.py Fix: tag budget reset must drop stale management-cache entry (#27568) 2026-05-10 00:18:55 +00:00

In total litellm runs 1000+ tests

[02/20/2025] Update:

To make it easier to contribute and map what behavior is tested,

we've started mapping the litellm directory in tests/test_litellm

This folder can only run mock tests.