2025-02-26 07:19:19 +08:00
|
|
|
import os
|
|
|
|
|
import sys
|
2026-05-22 10:08:37 +08:00
|
|
|
from pathlib import Path
|
|
|
|
|
from types import SimpleNamespace
|
2026-05-12 05:49:38 +08:00
|
|
|
from unittest.mock import AsyncMock, MagicMock, patch
|
2025-05-27 05:41:42 +08:00
|
|
|
|
2026-05-22 10:08:37 +08:00
|
|
|
import click
|
|
|
|
|
import fastapi
|
2025-08-13 13:03:39 +08:00
|
|
|
import pytest
|
2025-02-26 07:19:19 +08:00
|
|
|
|
|
|
|
|
sys.path.insert(
|
|
|
|
|
0, os.path.abspath("../../..")
|
|
|
|
|
) # Adds the parent directory to the system-path
|
|
|
|
|
|
2025-07-17 11:35:09 +08:00
|
|
|
import builtins
|
|
|
|
|
import types
|
2025-02-26 07:19:19 +08:00
|
|
|
|
2025-08-13 13:03:39 +08:00
|
|
|
from litellm.proxy.proxy_cli import ProxyInitializationHelpers
|
|
|
|
|
|
2025-02-26 07:19:19 +08:00
|
|
|
|
[Release Fix] (#22411)
* fix(lint): suppress PLR0915 for 3 complex methods that exceed 50-statement limit
- streaming_iterator.py: _process_event (84 statements)
- transformation.py: translate_messages_to_responses_input (51 statements)
- transformation.py: transform_realtime_response (54 statements)
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(mypy): resolve type errors in public_endpoints, user_api_key_auth, common_utils, transformation
- public_endpoints.py: fix _cached_endpoints type annotation
- user_api_key_auth.py: accept Optional[str] for end_user_id parameter
- common_utils.py: add NewProjectRequest/UpdateProjectRequest to Union type
- transformation.py: add ChatCompletionRedactedThinkingBlock and list[Any] to content type
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(proxy-extras): bump version to 0.4.50 and sync schema
- Bump litellm-proxy-extras from 0.4.49 to 0.4.50
- Sync schema.prisma with main proxy schema
- Includes new LiteLLM_ClaudeCodePluginTable model
- Includes new @@index([startTime, request_id]) on SpendLogs
- Update version references in requirements.txt and pyproject.toml
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(router): use string id in test_add_deployment and add defensive str() in register_model
- Change test to use string '100' instead of int 100 for model_info.id
- Add str() conversion in register_model to prevent AttributeError on non-string keys
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(security): update minimatch to 10.2.4 to fix CVE-2026-27903 and CVE-2026-27904
- Run npm audit fix in docs/my-website
- Updates minimatch from 10.2.1 to 10.2.4 (fixes HIGH severity ReDoS vulnerabilities)
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(test): update realtime guardrail test assertions to match actual guardrail behavior
- test_text_message_blocked_by_guardrail_no_ai_response: allow guardrail's own block
message text in response.done (previously expected empty content)
- test_voice_transcript_blocked_by_guardrail: allow guardrail to send response.cancel
+ block message + response.create flow (previously expected no response.create)
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix: revert proxy-extras version in requirements.txt and pyproject.toml
The litellm-proxy-extras 0.4.50 is not published to PyPI yet, so consumer
references must stay at 0.4.49. Only the source package pyproject.toml
should be bumped to 0.4.50 for the publish_proxy_extras CI job.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix: make transcript delta check optional in voice guardrail test
The guardrail sends an error event (guardrail_violation) when blocking
voice transcripts; it does not always produce transcript deltas. Remove
the assertion requiring response.audio_transcript.delta since the error
event is the primary signal that blocked content was handled.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* Add missing env keys to documentation: LITELLM_MAX_STREAMING_DURATION_SECONDS and LITELLM_USE_CHAT_COMPLETIONS_URL_FOR_ANTHROPIC_MESSAGES
These two environment variables were used in code but not documented in the
environment variables reference section of config_settings.md, causing the
test_env_keys.py CI test to fail.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* Fix 13 mypy type errors across 6 files
- in_flight_requests_middleware.py: Fix type: ignore error codes from
[union-attr] to [attr-defined], add [arg-type] for Gauge **kwargs
- transformation.py: Add [assignment] ignore for output_format reassignment,
add fallback empty string for tool use id to fix arg-type
- responses/main.py: Remove redundant type annotation on second
secret_fields assignment to fix no-redef
- streaming_iterator.py: Add [assignment] ignores for intermediate
cache token assignments
- handler.py: Add [typeddict-item] ignore for AnthropicMessagesRequest
construction from dict
- public_endpoints.py: Add [arg-type] ignore for _load_endpoints()
return type mismatch with SupportedEndpoint model
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix: add auth overrides to spend tracking tests, fix realtime guardrail assertion, update UI minimatch
- Add app.dependency_overrides for user_api_key_auth in 4 spend tracking tests
that were returning 401 Unauthorized (error_code, error_message,
error_code_and_key_alias, key_hash)
- Fix realtime guardrail test to check ANY error event for guardrail_violation
instead of just the first (OpenAI may send its own errors first)
- Update ui/litellm-dashboard/package-lock.json to fix minimatch vulnerability
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* Fix failing MCP e2e and create_mcp_server UI tests
Test 1 (test_independent_clients_no_shared_session):
- Add allow_all_keys: true to MCP servers in test config. With master_key
and no DB, get_allowed_mcp_servers returned empty, causing 0 tools and
403 on tool calls. allow_all_keys bypasses per-key restrictions.
- Add asyncio.sleep(0.5) between client connections to allow MCP SDK
TaskGroup cleanup and avoid ExceptionGroup on connection close (MCP #915).
Test 2 (create_mcp_server 'auth value is provided'):
- Use userEvent.setup({ delay: null }) for instant keystrokes to avoid
timeout from default typing delay on CI.
- Increase per-test timeout to 15000ms for CI environments.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix: stabilize proxy unit tests for parallel execution
- test_response_polling_handler: add xdist_group to prevent heavy import OOM
- test_db_schema_migration: use temp dir for worker isolation, sync schema.prisma index
- test_custom_tokenizer_bug: use lighter tokenizer to prevent OOM in parallel
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix: add auth overrides to more spend tracking and model info tests
- Fix test_ui_view_spend_logs_pagination missing auth override (401)
- Fix test_view_spend_tags missing auth override (401)
- Fix test_view_spend_tags_no_database missing auth override (401)
- Fix test_empty_model_list.py to use app.dependency_overrides instead of patch()
for FastAPI dependency injection auth
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(test): use patch.object for aiohttp transport test to work in parallel execution
The @patch decorator was not intercepting the static method call in parallel
xdist workers. Using patch.object on the directly-imported class is more reliable.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(security): update minimatch from 10.2.1 to 10.2.4 in Dockerfile
The Docker image was explicitly pinning minimatch@10.2.1 which has HIGH
severity ReDoS vulnerabilities (GHSA-7r86-cg39-jmmj, GHSA-23c5-xmqv-rm74).
Update to 10.2.4 which includes fixes for both CVEs.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(ui): prevent MCP and TeamInfo test timeouts on CI
- Add userEvent.setup({ delay: null }) to all tests using userEvent in both files
- Add timeout: 15000 to tests with significant user interaction (typing, multiple clicks)
- Fixes: create_mcp_server Bearer Token test, TeamInfo cancel button test
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix: stabilize parallel test execution and aiohttp transport test
- test_aiohttp_handler: rewrite transport test to not rely on static method mock
(consistently fails in parallel xdist workers)
- test_proxy_cli: add xdist_group to prevent timeout during heavy imports
- test_swagger_chat_completions: add xdist_group to prevent timeout
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(security): add serialize-javascript override to fix GHSA-5c6j-r48x-rmvq
Add npm override for serialize-javascript>=7.0.3 in docs/my-website
to fix HIGH severity RCE vulnerability via RegExp.flags.
Also bump minimatch override to >=10.2.4.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* Fix flaky tests: remove broken Vertex model, add retries for Anthropic
- Remove vertex_ai/meta/llama-4-scout-17b-16e-instruct-maas from
test_partner_models_httpx_streaming - consistently returns 400 BadRequest
- Add @pytest.mark.flaky(retries=6, delay=10) to test_function_call_parsing
for transient Anthropic API overload errors
- Add @pytest.mark.flaky(retries=6, delay=10) to test_openai_stream_options_call
for transient Anthropic InternalServerError
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(ci): add xdist_group(proxy_heavy) to prevent OOM in parallel proxy tests
- Add pytestmark = pytest.mark.xdist_group('proxy_heavy') to test_proxy_utils.py
- Change test_db_schema_migration.py from schema_migration to proxy_heavy group
- Add @pytest.mark.xdist_group('proxy_heavy') to test_proxy_server.py::test_health
Groups heavy proxy tests to run on same worker, avoiding worker OOM crashes.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* Fix vertex AI qwen global endpoint test to mock vertexai module import
The test_vertex_ai_qwen_global_endpoint_url test was failing because the
VertexAIPartnerModels.completion() method tries to 'import vertexai' before
any of the mocked code runs. In environments without google-cloud-aiplatform
installed, this import fails with a VertexAIError(status_code=400).
Fix by:
- Adding patch.dict('sys.modules', {'vertexai': MagicMock()}) to mock the
vertexai module import
- Adding vertex_ai_location parameter to the acompletion call for completeness
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(ci): add xdist_group to health endpoint and watsonx tests for parallel stability
- test_health_liveliness_endpoint: add xdist_group('proxy_health') to prevent timeout
- test_watsonx_gpt_oss tests: add xdist_group('watsonx_heavy') to prevent mock interference
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(test): pre-populate WatsonX IAM token cache to prevent parallel test interference
The watsonx prompt transformation test was failing in parallel execution because
litellm.module_level_client.post mock was being interfered with by other tests.
Pre-populating the IAM token cache avoids the HTTP call entirely.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(test): add spend data polling with retries for e2e pass-through tests
- test_vertex_with_spend.test.js: Replace 15s fixed wait with polling loop
(up to 6 attempts, 10s apart) for spend data to appear in DB
- Increase test timeout from 25s to 90s to accommodate polling
- base_anthropic_messages_tool_search_test.py: Add flaky(retries=3) for
streaming test that depends on live Anthropic API
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(ci): reduce parallel workers from 8 to 4 for proxy tests to prevent OOM
- litellm_proxy_unit_testing_part2: -n 8 -> -n 4
- litellm_mapped_tests_proxy_part2: -n 8 -> -n 4, timeout 60 -> 120
- Worker crashes consistently caused by too many parallel proxy tests
each loading the full FastAPI app and heavy dependency tree
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(db): add migration for SpendLogs composite index (startTime, request_id)
The @@index([startTime, request_id]) was added to schema.prisma but had no
corresponding migration. This caused test_aaaasschema_migration_check to fail
because prisma migrate diff detected the missing index.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(db): add migration for MCP available_on_public_internet default change to true
The schema.prisma changed the default for available_on_public_internet from
false to true, but no migration was created. This caused the schema migration
test to detect drift.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(test): increase server wait time and add retry to flaky external API tests
- test_basic_python_version.py: increase server startup wait from 60s to 90s
for slower CI environments (fixes installing_litellm_on_python_3_13)
- test_a2a_agent.py: add flaky(retries=3, delay=5) for non-streaming test
that depends on live A2A agent endpoint
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(test): add flaky retries to all intermittent external API tests for 0-fail CI
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(test): add auth overrides to file endpoint tests that return 500
The test_target_storage tests were getting 500 because the FastAPI auth
dependency wasn't overridden. Added app.dependency_overrides for proper
auth bypass in test environment.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
2026-03-01 01:46:35 +08:00
|
|
|
@pytest.mark.xdist_group("proxy_cli")
|
2025-02-26 07:19:19 +08:00
|
|
|
class TestProxyInitializationHelpers:
|
|
|
|
|
@patch("importlib.metadata.version")
|
2025-06-07 08:55:45 +08:00
|
|
|
@patch("click.echo")
|
|
|
|
|
def test_echo_litellm_version(self, mock_echo, mock_version):
|
2025-02-26 07:19:19 +08:00
|
|
|
# Setup
|
|
|
|
|
mock_version.return_value = "1.0.0"
|
|
|
|
|
|
|
|
|
|
# Execute
|
|
|
|
|
ProxyInitializationHelpers._echo_litellm_version()
|
|
|
|
|
|
|
|
|
|
# Assert
|
|
|
|
|
mock_version.assert_called_once_with("litellm")
|
2025-06-07 08:55:45 +08:00
|
|
|
mock_echo.assert_called_once_with("\nLiteLLM: Current Version = 1.0.0\n")
|
2025-02-26 07:19:19 +08:00
|
|
|
|
|
|
|
|
@patch("httpx.get")
|
2025-06-07 08:55:45 +08:00
|
|
|
@patch("builtins.print")
|
|
|
|
|
@patch("json.dumps")
|
|
|
|
|
def test_run_health_check(self, mock_dumps, mock_print, mock_get):
|
2025-02-26 07:19:19 +08:00
|
|
|
# Setup
|
|
|
|
|
mock_response = MagicMock()
|
2025-06-07 08:55:45 +08:00
|
|
|
mock_response.json.return_value = {"status": "healthy"}
|
2025-02-26 07:19:19 +08:00
|
|
|
mock_get.return_value = mock_response
|
2025-06-07 08:55:45 +08:00
|
|
|
mock_dumps.return_value = '{"status": "healthy"}'
|
2025-02-26 07:19:19 +08:00
|
|
|
|
|
|
|
|
# Execute
|
|
|
|
|
ProxyInitializationHelpers._run_health_check("localhost", 8000)
|
|
|
|
|
|
|
|
|
|
# Assert
|
|
|
|
|
mock_get.assert_called_once_with(url="http://localhost:8000/health")
|
|
|
|
|
mock_response.json.assert_called_once()
|
2025-06-07 08:55:45 +08:00
|
|
|
mock_dumps.assert_called_once_with({"status": "healthy"}, indent=4)
|
2025-02-26 07:19:19 +08:00
|
|
|
|
|
|
|
|
@patch("openai.OpenAI")
|
|
|
|
|
@patch("click.echo")
|
|
|
|
|
@patch("builtins.print")
|
|
|
|
|
def test_run_test_chat_completion(self, mock_print, mock_echo, mock_openai):
|
|
|
|
|
# Setup
|
|
|
|
|
mock_client = MagicMock()
|
|
|
|
|
mock_openai.return_value = mock_client
|
|
|
|
|
|
|
|
|
|
mock_response = MagicMock()
|
|
|
|
|
mock_client.chat.completions.create.return_value = mock_response
|
|
|
|
|
|
|
|
|
|
mock_stream_response = MagicMock()
|
|
|
|
|
mock_stream_response.__iter__.return_value = [MagicMock(), MagicMock()]
|
|
|
|
|
mock_client.chat.completions.create.side_effect = [
|
|
|
|
|
mock_response,
|
|
|
|
|
mock_stream_response,
|
|
|
|
|
]
|
|
|
|
|
|
|
|
|
|
# Execute
|
|
|
|
|
with pytest.raises(ValueError, match="Invalid test value"):
|
|
|
|
|
ProxyInitializationHelpers._run_test_chat_completion(
|
|
|
|
|
"localhost", 8000, "gpt-3.5-turbo", True
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
# Test with valid string test value
|
|
|
|
|
ProxyInitializationHelpers._run_test_chat_completion(
|
|
|
|
|
"localhost", 8000, "gpt-3.5-turbo", "http://test-url"
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
# Assert
|
|
|
|
|
mock_openai.assert_called_once_with(
|
|
|
|
|
api_key="My API Key", base_url="http://test-url"
|
|
|
|
|
)
|
|
|
|
|
mock_client.chat.completions.create.assert_called()
|
|
|
|
|
|
|
|
|
|
def test_get_default_unvicorn_init_args(self):
|
|
|
|
|
# Test without log_config
|
|
|
|
|
args = ProxyInitializationHelpers._get_default_unvicorn_init_args(
|
|
|
|
|
"localhost", 8000
|
|
|
|
|
)
|
|
|
|
|
assert args["app"] == "litellm.proxy.proxy_server:app"
|
|
|
|
|
assert args["host"] == "localhost"
|
|
|
|
|
assert args["port"] == 8000
|
|
|
|
|
|
|
|
|
|
# Test with log_config
|
|
|
|
|
args = ProxyInitializationHelpers._get_default_unvicorn_init_args(
|
|
|
|
|
"localhost", 8000, "log_config.json"
|
|
|
|
|
)
|
|
|
|
|
assert args["log_config"] == "log_config.json"
|
|
|
|
|
|
|
|
|
|
# Test with json_logs=True
|
|
|
|
|
with patch("litellm.json_logs", True):
|
|
|
|
|
args = ProxyInitializationHelpers._get_default_unvicorn_init_args(
|
|
|
|
|
"localhost", 8000
|
|
|
|
|
)
|
2026-01-25 04:59:51 +08:00
|
|
|
# When json_logs is True, log_config should be set to the JSON log config dict
|
|
|
|
|
assert args["log_config"] is not None
|
|
|
|
|
assert isinstance(args["log_config"], dict)
|
|
|
|
|
assert "version" in args["log_config"]
|
|
|
|
|
assert "formatters" in args["log_config"]
|
2025-02-26 07:19:19 +08:00
|
|
|
|
2025-06-11 04:30:19 +08:00
|
|
|
# Test with keepalive_timeout
|
|
|
|
|
args = ProxyInitializationHelpers._get_default_unvicorn_init_args(
|
|
|
|
|
"localhost", 8000, None, 60
|
|
|
|
|
)
|
|
|
|
|
assert args["timeout_keep_alive"] == 60
|
|
|
|
|
|
|
|
|
|
# Test with both log_config and keepalive_timeout
|
|
|
|
|
args = ProxyInitializationHelpers._get_default_unvicorn_init_args(
|
|
|
|
|
"localhost", 8000, "log_config.json", 120
|
|
|
|
|
)
|
|
|
|
|
assert args["log_config"] == "log_config.json"
|
|
|
|
|
assert args["timeout_keep_alive"] == 120
|
|
|
|
|
|
2026-04-28 02:06:56 +08:00
|
|
|
class _FakeUvicornConfig:
|
|
|
|
|
def __init__(self, timeout_worker_healthcheck=None):
|
|
|
|
|
pass
|
|
|
|
|
|
|
|
|
|
with patch("uvicorn.Config", _FakeUvicornConfig):
|
|
|
|
|
args = ProxyInitializationHelpers._get_default_unvicorn_init_args(
|
|
|
|
|
"localhost", 8000, timeout_worker_healthcheck=15
|
|
|
|
|
)
|
|
|
|
|
assert args["timeout_worker_healthcheck"] == 15
|
|
|
|
|
|
feat(proxy): hot-reload .env in dev when running with --reload (#29783)
* feat(proxy): hot-reload .env in dev when running with --reload
The --reload watcher already restarts the worker on *.py and --config YAML
edits, but .env was unwatched, so changing a key there did nothing until a
manual restart. Add .env to the uvicorn reload_includes (and to the
StatReload monkeypatch, which ignores reload_includes) so an edit triggers a
worker restart.
A reloaded worker is a fresh process that inherits the reloader's
environment, so load_dotenv(override=False) would keep serving the stale
inherited value for any key already in the environment. The CLI now exports
LITELLM_DEV_ENV_HOT_RELOAD when --reload is set, and litellm/__init__.py
reads it to load .env with override=True only on that dev path, leaving
normal startup precedence untouched.
* feat(proxy): warn that --reload makes .env override shell env vars
When --reload is active, worker processes re-read .env with override=True, so
.env values win over shell-exported environment variables. Surface this dotenv
precedence change with a startup warning so a developer who relies on a
shell-exported override is not silently surprised.
* fix(proxy): type reload helper paths as Optional[str] to satisfy mypy
* fix(proxy): watch the cwd .env in both reload backends for parity
WatchFiles only watches cwd (and the --config dir) for .env, while the
StatReload fallback used find_dotenv(usecwd=True), which walks up to a
parent-dir .env that WatchFiles never sees. Point StatReload at the same
cwd .env so the two reload backends react to the same file.
2026-06-07 00:39:21 +08:00
|
|
|
def test_get_reload_options_no_config_still_watches_env(self):
|
2026-05-07 00:06:58 +08:00
|
|
|
opts = ProxyInitializationHelpers._get_reload_options(None)
|
feat(proxy): hot-reload .env in dev when running with --reload (#29783)
* feat(proxy): hot-reload .env in dev when running with --reload
The --reload watcher already restarts the worker on *.py and --config YAML
edits, but .env was unwatched, so changing a key there did nothing until a
manual restart. Add .env to the uvicorn reload_includes (and to the
StatReload monkeypatch, which ignores reload_includes) so an edit triggers a
worker restart.
A reloaded worker is a fresh process that inherits the reloader's
environment, so load_dotenv(override=False) would keep serving the stale
inherited value for any key already in the environment. The CLI now exports
LITELLM_DEV_ENV_HOT_RELOAD when --reload is set, and litellm/__init__.py
reads it to load .env with override=True only on that dev path, leaving
normal startup precedence untouched.
* feat(proxy): warn that --reload makes .env override shell env vars
When --reload is active, worker processes re-read .env with override=True, so
.env values win over shell-exported environment variables. Surface this dotenv
precedence change with a startup warning so a developer who relies on a
shell-exported override is not silently surprised.
* fix(proxy): type reload helper paths as Optional[str] to satisfy mypy
* fix(proxy): watch the cwd .env in both reload backends for parity
WatchFiles only watches cwd (and the --config dir) for .env, while the
StatReload fallback used find_dotenv(usecwd=True), which walks up to a
parent-dir .env that WatchFiles never sees. Point StatReload at the same
cwd .env so the two reload backends react to the same file.
2026-06-07 00:39:21 +08:00
|
|
|
assert opts["reload"] is True
|
|
|
|
|
assert opts["reload_dirs"] == [os.path.abspath(os.getcwd())]
|
|
|
|
|
assert opts["reload_includes"] == ["*.py", ".env"]
|
2026-05-07 00:06:58 +08:00
|
|
|
|
|
|
|
|
def test_get_reload_options_with_config_in_cwd(self, tmp_path, monkeypatch):
|
|
|
|
|
config_file = tmp_path / "config.yaml"
|
|
|
|
|
config_file.write_text("model_list: []\n")
|
|
|
|
|
monkeypatch.chdir(tmp_path)
|
|
|
|
|
|
|
|
|
|
opts = ProxyInitializationHelpers._get_reload_options("config.yaml")
|
|
|
|
|
|
|
|
|
|
assert opts["reload"] is True
|
|
|
|
|
assert opts["reload_dirs"] == [str(tmp_path)]
|
feat(proxy): hot-reload .env in dev when running with --reload (#29783)
* feat(proxy): hot-reload .env in dev when running with --reload
The --reload watcher already restarts the worker on *.py and --config YAML
edits, but .env was unwatched, so changing a key there did nothing until a
manual restart. Add .env to the uvicorn reload_includes (and to the
StatReload monkeypatch, which ignores reload_includes) so an edit triggers a
worker restart.
A reloaded worker is a fresh process that inherits the reloader's
environment, so load_dotenv(override=False) would keep serving the stale
inherited value for any key already in the environment. The CLI now exports
LITELLM_DEV_ENV_HOT_RELOAD when --reload is set, and litellm/__init__.py
reads it to load .env with override=True only on that dev path, leaving
normal startup precedence untouched.
* feat(proxy): warn that --reload makes .env override shell env vars
When --reload is active, worker processes re-read .env with override=True, so
.env values win over shell-exported environment variables. Surface this dotenv
precedence change with a startup warning so a developer who relies on a
shell-exported override is not silently surprised.
* fix(proxy): type reload helper paths as Optional[str] to satisfy mypy
* fix(proxy): watch the cwd .env in both reload backends for parity
WatchFiles only watches cwd (and the --config dir) for .env, while the
StatReload fallback used find_dotenv(usecwd=True), which walks up to a
parent-dir .env that WatchFiles never sees. Point StatReload at the same
cwd .env so the two reload backends react to the same file.
2026-06-07 00:39:21 +08:00
|
|
|
assert opts["reload_includes"] == ["*.py", ".env", "config.yaml"]
|
2026-05-07 00:06:58 +08:00
|
|
|
|
|
|
|
|
def test_get_reload_options_with_config_outside_cwd(self, tmp_path, monkeypatch):
|
|
|
|
|
cwd_dir = tmp_path / "work"
|
|
|
|
|
cwd_dir.mkdir()
|
|
|
|
|
elsewhere = tmp_path / "configs"
|
|
|
|
|
elsewhere.mkdir()
|
|
|
|
|
config_file = elsewhere / "proxy.yaml"
|
|
|
|
|
config_file.write_text("model_list: []\n")
|
|
|
|
|
monkeypatch.chdir(cwd_dir)
|
|
|
|
|
|
|
|
|
|
opts = ProxyInitializationHelpers._get_reload_options(str(config_file))
|
|
|
|
|
|
|
|
|
|
assert opts["reload"] is True
|
|
|
|
|
assert opts["reload_dirs"] == [str(cwd_dir), str(elsewhere)]
|
feat(proxy): hot-reload .env in dev when running with --reload (#29783)
* feat(proxy): hot-reload .env in dev when running with --reload
The --reload watcher already restarts the worker on *.py and --config YAML
edits, but .env was unwatched, so changing a key there did nothing until a
manual restart. Add .env to the uvicorn reload_includes (and to the
StatReload monkeypatch, which ignores reload_includes) so an edit triggers a
worker restart.
A reloaded worker is a fresh process that inherits the reloader's
environment, so load_dotenv(override=False) would keep serving the stale
inherited value for any key already in the environment. The CLI now exports
LITELLM_DEV_ENV_HOT_RELOAD when --reload is set, and litellm/__init__.py
reads it to load .env with override=True only on that dev path, leaving
normal startup precedence untouched.
* feat(proxy): warn that --reload makes .env override shell env vars
When --reload is active, worker processes re-read .env with override=True, so
.env values win over shell-exported environment variables. Surface this dotenv
precedence change with a startup warning so a developer who relies on a
shell-exported override is not silently surprised.
* fix(proxy): type reload helper paths as Optional[str] to satisfy mypy
* fix(proxy): watch the cwd .env in both reload backends for parity
WatchFiles only watches cwd (and the --config dir) for .env, while the
StatReload fallback used find_dotenv(usecwd=True), which walks up to a
parent-dir .env that WatchFiles never sees. Point StatReload at the same
cwd .env so the two reload backends react to the same file.
2026-06-07 00:39:21 +08:00
|
|
|
assert opts["reload_includes"] == ["*.py", ".env", "proxy.yaml"]
|
2026-05-07 00:06:58 +08:00
|
|
|
|
feat(proxy): hot-reload .env in dev when running with --reload (#29783)
* feat(proxy): hot-reload .env in dev when running with --reload
The --reload watcher already restarts the worker on *.py and --config YAML
edits, but .env was unwatched, so changing a key there did nothing until a
manual restart. Add .env to the uvicorn reload_includes (and to the
StatReload monkeypatch, which ignores reload_includes) so an edit triggers a
worker restart.
A reloaded worker is a fresh process that inherits the reloader's
environment, so load_dotenv(override=False) would keep serving the stale
inherited value for any key already in the environment. The CLI now exports
LITELLM_DEV_ENV_HOT_RELOAD when --reload is set, and litellm/__init__.py
reads it to load .env with override=True only on that dev path, leaving
normal startup precedence untouched.
* feat(proxy): warn that --reload makes .env override shell env vars
When --reload is active, worker processes re-read .env with override=True, so
.env values win over shell-exported environment variables. Surface this dotenv
precedence change with a startup warning so a developer who relies on a
shell-exported override is not silently surprised.
* fix(proxy): type reload helper paths as Optional[str] to satisfy mypy
* fix(proxy): watch the cwd .env in both reload backends for parity
WatchFiles only watches cwd (and the --config dir) for .env, while the
StatReload fallback used find_dotenv(usecwd=True), which walks up to a
parent-dir .env that WatchFiles never sees. Point StatReload at the same
cwd .env so the two reload backends react to the same file.
2026-06-07 00:39:21 +08:00
|
|
|
def test_patch_statreload_extra_paths_yields_config_and_py(self, tmp_path):
|
2026-05-07 00:06:58 +08:00
|
|
|
from pathlib import Path
|
|
|
|
|
|
|
|
|
|
from uvicorn.supervisors.statreload import StatReload
|
|
|
|
|
|
|
|
|
|
if hasattr(StatReload, "_litellm_patched_config_paths"):
|
|
|
|
|
StatReload._litellm_patched_config_paths.clear()
|
|
|
|
|
|
|
|
|
|
config_file = tmp_path / "config.yaml"
|
|
|
|
|
config_file.write_text("model_list: []\n")
|
|
|
|
|
py_file = tmp_path / "module.py"
|
|
|
|
|
py_file.write_text("x = 1\n")
|
|
|
|
|
|
feat(proxy): hot-reload .env in dev when running with --reload (#29783)
* feat(proxy): hot-reload .env in dev when running with --reload
The --reload watcher already restarts the worker on *.py and --config YAML
edits, but .env was unwatched, so changing a key there did nothing until a
manual restart. Add .env to the uvicorn reload_includes (and to the
StatReload monkeypatch, which ignores reload_includes) so an edit triggers a
worker restart.
A reloaded worker is a fresh process that inherits the reloader's
environment, so load_dotenv(override=False) would keep serving the stale
inherited value for any key already in the environment. The CLI now exports
LITELLM_DEV_ENV_HOT_RELOAD when --reload is set, and litellm/__init__.py
reads it to load .env with override=True only on that dev path, leaving
normal startup precedence untouched.
* feat(proxy): warn that --reload makes .env override shell env vars
When --reload is active, worker processes re-read .env with override=True, so
.env values win over shell-exported environment variables. Surface this dotenv
precedence change with a startup warning so a developer who relies on a
shell-exported override is not silently surprised.
* fix(proxy): type reload helper paths as Optional[str] to satisfy mypy
* fix(proxy): watch the cwd .env in both reload backends for parity
WatchFiles only watches cwd (and the --config dir) for .env, while the
StatReload fallback used find_dotenv(usecwd=True), which walks up to a
parent-dir .env that WatchFiles never sees. Point StatReload at the same
cwd .env so the two reload backends react to the same file.
2026-06-07 00:39:21 +08:00
|
|
|
applied = ProxyInitializationHelpers._patch_statreload_extra_paths(
|
|
|
|
|
[str(config_file)]
|
2026-05-07 00:06:58 +08:00
|
|
|
)
|
|
|
|
|
assert applied is True
|
|
|
|
|
|
|
|
|
|
fake_self = types.SimpleNamespace(
|
|
|
|
|
config=types.SimpleNamespace(reload_dirs=[tmp_path])
|
|
|
|
|
)
|
|
|
|
|
yielded_paths = {Path(p).resolve() for p in StatReload.iter_py_files(fake_self)}
|
|
|
|
|
|
|
|
|
|
assert config_file.resolve() in yielded_paths
|
|
|
|
|
assert py_file.resolve() in yielded_paths
|
|
|
|
|
|
feat(proxy): hot-reload .env in dev when running with --reload (#29783)
* feat(proxy): hot-reload .env in dev when running with --reload
The --reload watcher already restarts the worker on *.py and --config YAML
edits, but .env was unwatched, so changing a key there did nothing until a
manual restart. Add .env to the uvicorn reload_includes (and to the
StatReload monkeypatch, which ignores reload_includes) so an edit triggers a
worker restart.
A reloaded worker is a fresh process that inherits the reloader's
environment, so load_dotenv(override=False) would keep serving the stale
inherited value for any key already in the environment. The CLI now exports
LITELLM_DEV_ENV_HOT_RELOAD when --reload is set, and litellm/__init__.py
reads it to load .env with override=True only on that dev path, leaving
normal startup precedence untouched.
* feat(proxy): warn that --reload makes .env override shell env vars
When --reload is active, worker processes re-read .env with override=True, so
.env values win over shell-exported environment variables. Surface this dotenv
precedence change with a startup warning so a developer who relies on a
shell-exported override is not silently surprised.
* fix(proxy): type reload helper paths as Optional[str] to satisfy mypy
* fix(proxy): watch the cwd .env in both reload backends for parity
WatchFiles only watches cwd (and the --config dir) for .env, while the
StatReload fallback used find_dotenv(usecwd=True), which walks up to a
parent-dir .env that WatchFiles never sees. Point StatReload at the same
cwd .env so the two reload backends react to the same file.
2026-06-07 00:39:21 +08:00
|
|
|
def test_patch_statreload_extra_paths_yields_env(self, tmp_path):
|
|
|
|
|
from pathlib import Path
|
|
|
|
|
|
|
|
|
|
from uvicorn.supervisors.statreload import StatReload
|
|
|
|
|
|
|
|
|
|
if hasattr(StatReload, "_litellm_patched_config_paths"):
|
|
|
|
|
StatReload._litellm_patched_config_paths.clear()
|
|
|
|
|
|
|
|
|
|
env_file = tmp_path / ".env"
|
|
|
|
|
env_file.write_text("FOO=bar\n")
|
|
|
|
|
|
|
|
|
|
applied = ProxyInitializationHelpers._patch_statreload_extra_paths(
|
|
|
|
|
[str(env_file)]
|
|
|
|
|
)
|
|
|
|
|
assert applied is True
|
|
|
|
|
|
|
|
|
|
fake_self = types.SimpleNamespace(
|
|
|
|
|
config=types.SimpleNamespace(reload_dirs=[tmp_path])
|
|
|
|
|
)
|
|
|
|
|
yielded_paths = {Path(p).resolve() for p in StatReload.iter_py_files(fake_self)}
|
|
|
|
|
|
|
|
|
|
assert env_file.resolve() in yielded_paths
|
|
|
|
|
|
|
|
|
|
def test_patch_statreload_extra_paths_skips_falsy(self, tmp_path):
|
|
|
|
|
from uvicorn.supervisors.statreload import StatReload
|
|
|
|
|
|
|
|
|
|
if hasattr(StatReload, "_litellm_patched_config_paths"):
|
|
|
|
|
StatReload._litellm_patched_config_paths.clear()
|
|
|
|
|
|
|
|
|
|
assert ProxyInitializationHelpers._patch_statreload_extra_paths([]) is False
|
|
|
|
|
assert (
|
|
|
|
|
ProxyInitializationHelpers._patch_statreload_extra_paths([None, ""])
|
|
|
|
|
is False
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
def test_patch_statreload_extra_paths_is_idempotent(self, tmp_path):
|
2026-05-07 00:06:58 +08:00
|
|
|
from pathlib import Path
|
|
|
|
|
|
|
|
|
|
from uvicorn.supervisors.statreload import StatReload
|
|
|
|
|
|
|
|
|
|
if hasattr(StatReload, "_litellm_patched_config_paths"):
|
|
|
|
|
StatReload._litellm_patched_config_paths.clear()
|
|
|
|
|
|
|
|
|
|
config_file = tmp_path / "config.yaml"
|
|
|
|
|
config_file.write_text("model_list: []\n")
|
|
|
|
|
py_file = tmp_path / "only.py"
|
|
|
|
|
py_file.write_text("x = 1\n")
|
|
|
|
|
|
|
|
|
|
for _ in range(3):
|
feat(proxy): hot-reload .env in dev when running with --reload (#29783)
* feat(proxy): hot-reload .env in dev when running with --reload
The --reload watcher already restarts the worker on *.py and --config YAML
edits, but .env was unwatched, so changing a key there did nothing until a
manual restart. Add .env to the uvicorn reload_includes (and to the
StatReload monkeypatch, which ignores reload_includes) so an edit triggers a
worker restart.
A reloaded worker is a fresh process that inherits the reloader's
environment, so load_dotenv(override=False) would keep serving the stale
inherited value for any key already in the environment. The CLI now exports
LITELLM_DEV_ENV_HOT_RELOAD when --reload is set, and litellm/__init__.py
reads it to load .env with override=True only on that dev path, leaving
normal startup precedence untouched.
* feat(proxy): warn that --reload makes .env override shell env vars
When --reload is active, worker processes re-read .env with override=True, so
.env values win over shell-exported environment variables. Surface this dotenv
precedence change with a startup warning so a developer who relies on a
shell-exported override is not silently surprised.
* fix(proxy): type reload helper paths as Optional[str] to satisfy mypy
* fix(proxy): watch the cwd .env in both reload backends for parity
WatchFiles only watches cwd (and the --config dir) for .env, while the
StatReload fallback used find_dotenv(usecwd=True), which walks up to a
parent-dir .env that WatchFiles never sees. Point StatReload at the same
cwd .env so the two reload backends react to the same file.
2026-06-07 00:39:21 +08:00
|
|
|
ProxyInitializationHelpers._patch_statreload_extra_paths([str(config_file)])
|
2026-05-07 00:06:58 +08:00
|
|
|
|
|
|
|
|
fake_self = types.SimpleNamespace(
|
|
|
|
|
config=types.SimpleNamespace(reload_dirs=[tmp_path])
|
|
|
|
|
)
|
|
|
|
|
yielded = list(StatReload.iter_py_files(fake_self))
|
|
|
|
|
assert len(yielded) == len(set(map(str, yielded)))
|
|
|
|
|
yielded_paths = {Path(p).resolve() for p in yielded}
|
|
|
|
|
assert config_file.resolve() in yielded_paths
|
|
|
|
|
assert py_file.resolve() in yielded_paths
|
|
|
|
|
|
feat(proxy): hot-reload .env in dev when running with --reload (#29783)
* feat(proxy): hot-reload .env in dev when running with --reload
The --reload watcher already restarts the worker on *.py and --config YAML
edits, but .env was unwatched, so changing a key there did nothing until a
manual restart. Add .env to the uvicorn reload_includes (and to the
StatReload monkeypatch, which ignores reload_includes) so an edit triggers a
worker restart.
A reloaded worker is a fresh process that inherits the reloader's
environment, so load_dotenv(override=False) would keep serving the stale
inherited value for any key already in the environment. The CLI now exports
LITELLM_DEV_ENV_HOT_RELOAD when --reload is set, and litellm/__init__.py
reads it to load .env with override=True only on that dev path, leaving
normal startup precedence untouched.
* feat(proxy): warn that --reload makes .env override shell env vars
When --reload is active, worker processes re-read .env with override=True, so
.env values win over shell-exported environment variables. Surface this dotenv
precedence change with a startup warning so a developer who relies on a
shell-exported override is not silently surprised.
* fix(proxy): type reload helper paths as Optional[str] to satisfy mypy
* fix(proxy): watch the cwd .env in both reload backends for parity
WatchFiles only watches cwd (and the --config dir) for .env, while the
StatReload fallback used find_dotenv(usecwd=True), which walks up to a
parent-dir .env that WatchFiles never sees. Point StatReload at the same
cwd .env so the two reload backends react to the same file.
2026-06-07 00:39:21 +08:00
|
|
|
def test_configure_dev_reload_watches_env_and_sets_override_flag(
|
|
|
|
|
self, tmp_path, monkeypatch
|
|
|
|
|
):
|
|
|
|
|
from pathlib import Path
|
|
|
|
|
|
|
|
|
|
from uvicorn.supervisors.statreload import StatReload
|
|
|
|
|
|
|
|
|
|
if hasattr(StatReload, "_litellm_patched_config_paths"):
|
|
|
|
|
StatReload._litellm_patched_config_paths.clear()
|
|
|
|
|
monkeypatch.delenv("LITELLM_DEV_ENV_HOT_RELOAD", raising=False)
|
|
|
|
|
|
|
|
|
|
config_file = tmp_path / "config.yaml"
|
|
|
|
|
config_file.write_text("model_list: []\n")
|
|
|
|
|
env_file = tmp_path / ".env"
|
|
|
|
|
env_file.write_text("FOO=bar\n")
|
|
|
|
|
monkeypatch.chdir(tmp_path)
|
|
|
|
|
|
|
|
|
|
uvicorn_args: dict = {}
|
|
|
|
|
with patch("litellm._logging.verbose_proxy_logger.warning") as mock_warning:
|
|
|
|
|
ProxyInitializationHelpers._configure_dev_reload(
|
|
|
|
|
uvicorn_args, str(config_file)
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
assert os.environ["LITELLM_DEV_ENV_HOT_RELOAD"] == "True"
|
|
|
|
|
assert uvicorn_args["reload"] is True
|
|
|
|
|
assert ".env" in uvicorn_args["reload_includes"]
|
|
|
|
|
|
|
|
|
|
mock_warning.assert_called_once()
|
|
|
|
|
warning_text = mock_warning.call_args.args[0].lower()
|
|
|
|
|
assert "override" in warning_text
|
|
|
|
|
assert ".env" in warning_text
|
|
|
|
|
|
|
|
|
|
fake_self = types.SimpleNamespace(
|
|
|
|
|
config=types.SimpleNamespace(reload_dirs=[tmp_path])
|
|
|
|
|
)
|
|
|
|
|
yielded_paths = {Path(p).resolve() for p in StatReload.iter_py_files(fake_self)}
|
|
|
|
|
assert env_file.resolve() in yielded_paths
|
|
|
|
|
assert config_file.resolve() in yielded_paths
|
|
|
|
|
|
|
|
|
|
def test_dev_env_hot_reload_enabled_reads_flag(self, monkeypatch):
|
|
|
|
|
import litellm
|
|
|
|
|
|
|
|
|
|
monkeypatch.setenv("LITELLM_DEV_ENV_HOT_RELOAD", "True")
|
|
|
|
|
assert litellm._dev_env_hot_reload_enabled() is True
|
|
|
|
|
|
|
|
|
|
monkeypatch.setenv("LITELLM_DEV_ENV_HOT_RELOAD", "false")
|
|
|
|
|
assert litellm._dev_env_hot_reload_enabled() is False
|
|
|
|
|
|
|
|
|
|
monkeypatch.delenv("LITELLM_DEV_ENV_HOT_RELOAD", raising=False)
|
|
|
|
|
assert litellm._dev_env_hot_reload_enabled() is False
|
|
|
|
|
|
2025-02-26 07:19:19 +08:00
|
|
|
@patch("asyncio.run")
|
|
|
|
|
@patch("builtins.print")
|
|
|
|
|
def test_init_hypercorn_server(self, mock_print, mock_asyncio_run):
|
|
|
|
|
# Setup
|
|
|
|
|
mock_app = MagicMock()
|
|
|
|
|
|
|
|
|
|
# Execute
|
|
|
|
|
ProxyInitializationHelpers._init_hypercorn_server(
|
2025-06-21 05:45:48 +08:00
|
|
|
mock_app, "localhost", 8000, None, None, None
|
2025-02-26 07:19:19 +08:00
|
|
|
)
|
|
|
|
|
|
|
|
|
|
# Assert
|
|
|
|
|
mock_asyncio_run.assert_called_once()
|
|
|
|
|
|
|
|
|
|
# Test with SSL
|
|
|
|
|
ProxyInitializationHelpers._init_hypercorn_server(
|
2025-06-21 05:45:48 +08:00
|
|
|
mock_app, "localhost", 8000, "cert.pem", "key.pem", "ECDHE"
|
2025-02-26 07:19:19 +08:00
|
|
|
)
|
|
|
|
|
|
2026-05-22 10:08:37 +08:00
|
|
|
@patch("granian.Granian")
|
|
|
|
|
@patch("builtins.print")
|
|
|
|
|
def test_init_granian_server(self, mock_print, mock_granian_cls):
|
|
|
|
|
pytest.importorskip("granian")
|
|
|
|
|
mock_server = MagicMock()
|
|
|
|
|
mock_granian_cls.return_value = mock_server
|
|
|
|
|
fake_interfaces = SimpleNamespace(ASGI="asgi")
|
|
|
|
|
with patch("granian.constants.Interfaces", fake_interfaces):
|
|
|
|
|
ProxyInitializationHelpers._init_granian_server(
|
|
|
|
|
host="0.0.0.0",
|
|
|
|
|
port=4000,
|
|
|
|
|
num_workers=2,
|
|
|
|
|
ssl_certfile_path=None,
|
|
|
|
|
ssl_keyfile_path=None,
|
|
|
|
|
max_requests_before_restart=None,
|
|
|
|
|
ciphers=None,
|
|
|
|
|
granian_runtime_threads=None,
|
|
|
|
|
)
|
|
|
|
|
mock_granian_cls.assert_called_once()
|
|
|
|
|
call_kwargs = mock_granian_cls.call_args.kwargs
|
|
|
|
|
assert call_kwargs["target"] == "litellm.proxy.proxy_server:app"
|
|
|
|
|
assert call_kwargs["address"] == "0.0.0.0"
|
|
|
|
|
assert call_kwargs["port"] == 4000
|
|
|
|
|
assert call_kwargs["workers"] == 2
|
|
|
|
|
assert call_kwargs["interface"] == "asgi"
|
|
|
|
|
assert call_kwargs["websockets"] is True
|
|
|
|
|
assert "runtime_threads" not in call_kwargs
|
|
|
|
|
mock_server.serve.assert_called_once()
|
|
|
|
|
|
|
|
|
|
@patch("granian.Granian")
|
|
|
|
|
@patch("builtins.print")
|
|
|
|
|
def test_init_granian_server_runtime_threads(self, mock_print, mock_granian_cls):
|
|
|
|
|
pytest.importorskip("granian")
|
|
|
|
|
mock_server = MagicMock()
|
|
|
|
|
mock_granian_cls.return_value = mock_server
|
|
|
|
|
fake_interfaces = SimpleNamespace(ASGI="asgi")
|
|
|
|
|
with patch("granian.constants.Interfaces", fake_interfaces):
|
|
|
|
|
ProxyInitializationHelpers._init_granian_server(
|
|
|
|
|
host="0.0.0.0",
|
|
|
|
|
port=4000,
|
|
|
|
|
num_workers=1,
|
|
|
|
|
ssl_certfile_path=None,
|
|
|
|
|
ssl_keyfile_path=None,
|
|
|
|
|
max_requests_before_restart=None,
|
|
|
|
|
ciphers=None,
|
|
|
|
|
granian_runtime_threads=4,
|
|
|
|
|
)
|
|
|
|
|
assert mock_granian_cls.call_args.kwargs["runtime_threads"] == 4
|
|
|
|
|
|
|
|
|
|
@patch("granian.Granian")
|
|
|
|
|
@patch("builtins.print")
|
|
|
|
|
def test_init_granian_server_ssl(self, mock_print, mock_granian_cls):
|
|
|
|
|
pytest.importorskip("granian")
|
|
|
|
|
mock_server = MagicMock()
|
|
|
|
|
mock_granian_cls.return_value = mock_server
|
|
|
|
|
fake_interfaces = SimpleNamespace(ASGI="asgi")
|
|
|
|
|
with patch("granian.constants.Interfaces", fake_interfaces):
|
|
|
|
|
ProxyInitializationHelpers._init_granian_server(
|
|
|
|
|
host="0.0.0.0",
|
|
|
|
|
port=4000,
|
|
|
|
|
num_workers=1,
|
|
|
|
|
ssl_certfile_path="/path/to/cert.pem",
|
|
|
|
|
ssl_keyfile_path="/path/to/key.pem",
|
|
|
|
|
max_requests_before_restart=None,
|
|
|
|
|
ciphers=None,
|
|
|
|
|
granian_runtime_threads=None,
|
|
|
|
|
)
|
|
|
|
|
call_kwargs = mock_granian_cls.call_args.kwargs
|
|
|
|
|
assert call_kwargs["ssl_cert"] == Path("/path/to/cert.pem")
|
|
|
|
|
assert call_kwargs["ssl_key"] == Path("/path/to/key.pem")
|
|
|
|
|
mock_server.serve.assert_called_once()
|
|
|
|
|
|
|
|
|
|
@patch("granian.Granian")
|
|
|
|
|
def test_init_granian_server_ssl_requires_cert_and_key(self, mock_granian_cls):
|
|
|
|
|
pytest.importorskip("granian")
|
|
|
|
|
fake_interfaces = SimpleNamespace(ASGI="asgi")
|
|
|
|
|
with patch("granian.constants.Interfaces", fake_interfaces):
|
|
|
|
|
with pytest.raises(click.ClickException, match="Both --ssl_certfile_path"):
|
|
|
|
|
ProxyInitializationHelpers._init_granian_server(
|
|
|
|
|
host="0.0.0.0",
|
|
|
|
|
port=4000,
|
|
|
|
|
num_workers=1,
|
|
|
|
|
ssl_certfile_path="/path/to/cert.pem",
|
|
|
|
|
ssl_keyfile_path=None,
|
|
|
|
|
max_requests_before_restart=None,
|
|
|
|
|
ciphers=None,
|
|
|
|
|
granian_runtime_threads=None,
|
|
|
|
|
)
|
|
|
|
|
mock_granian_cls.assert_not_called()
|
|
|
|
|
|
2025-02-26 07:19:19 +08:00
|
|
|
@patch("subprocess.Popen")
|
|
|
|
|
def test_run_ollama_serve(self, mock_popen):
|
|
|
|
|
# Execute
|
|
|
|
|
ProxyInitializationHelpers._run_ollama_serve()
|
|
|
|
|
|
|
|
|
|
# Assert
|
|
|
|
|
mock_popen.assert_called_once()
|
|
|
|
|
|
|
|
|
|
# Test exception handling
|
|
|
|
|
mock_popen.side_effect = Exception("Test exception")
|
|
|
|
|
ProxyInitializationHelpers._run_ollama_serve() # Should not raise
|
|
|
|
|
|
|
|
|
|
@patch("socket.socket")
|
|
|
|
|
def test_is_port_in_use(self, mock_socket):
|
|
|
|
|
# Setup for port in use
|
|
|
|
|
mock_socket_instance = MagicMock()
|
|
|
|
|
mock_socket_instance.connect_ex.return_value = 0
|
|
|
|
|
mock_socket.return_value.__enter__.return_value = mock_socket_instance
|
|
|
|
|
|
|
|
|
|
# Execute and Assert
|
|
|
|
|
assert ProxyInitializationHelpers._is_port_in_use(8000) is True
|
|
|
|
|
|
|
|
|
|
# Setup for port not in use
|
|
|
|
|
mock_socket_instance.connect_ex.return_value = 1
|
|
|
|
|
|
|
|
|
|
# Execute and Assert
|
|
|
|
|
assert ProxyInitializationHelpers._is_port_in_use(8000) is False
|
|
|
|
|
|
|
|
|
|
def test_get_loop_type(self):
|
|
|
|
|
# Test on Windows
|
|
|
|
|
with patch("sys.platform", "win32"):
|
|
|
|
|
assert ProxyInitializationHelpers._get_loop_type() is None
|
|
|
|
|
|
|
|
|
|
# Test on Linux
|
|
|
|
|
with patch("sys.platform", "linux"):
|
|
|
|
|
assert ProxyInitializationHelpers._get_loop_type() == "uvloop"
|
2025-05-19 23:00:41 +08:00
|
|
|
|
2025-05-20 00:29:12 +08:00
|
|
|
@patch.dict(os.environ, {}, clear=True)
|
|
|
|
|
def test_database_url_construction_with_special_characters(self):
|
|
|
|
|
# Setup environment variables with special characters that need escaping
|
|
|
|
|
test_env = {
|
|
|
|
|
"DATABASE_HOST": "localhost:5432",
|
|
|
|
|
"DATABASE_USERNAME": "user@with+special",
|
2025-12-23 09:03:53 +08:00
|
|
|
"DATABASE_PASSWORD": "test-password-special-chars",
|
2025-05-27 05:41:42 +08:00
|
|
|
"DATABASE_NAME": "db_name/test",
|
2025-05-20 00:29:12 +08:00
|
|
|
}
|
|
|
|
|
|
|
|
|
|
with patch.dict(os.environ, test_env):
|
|
|
|
|
# Call the relevant function - we'll need to extract the database URL construction logic
|
|
|
|
|
# This is simulating what happens in the run_server function when database_url is None
|
|
|
|
|
import urllib.parse
|
|
|
|
|
|
2025-05-27 05:41:42 +08:00
|
|
|
from litellm.proxy.proxy_cli import append_query_params
|
|
|
|
|
|
2025-05-20 00:29:12 +08:00
|
|
|
database_host = os.environ["DATABASE_HOST"]
|
|
|
|
|
database_username = os.environ["DATABASE_USERNAME"]
|
|
|
|
|
database_password = os.environ["DATABASE_PASSWORD"]
|
|
|
|
|
database_name = os.environ["DATABASE_NAME"]
|
|
|
|
|
|
|
|
|
|
# Test the URL encoding part
|
|
|
|
|
database_username_enc = urllib.parse.quote_plus(database_username)
|
|
|
|
|
database_password_enc = urllib.parse.quote_plus(database_password)
|
|
|
|
|
database_name_enc = urllib.parse.quote_plus(database_name)
|
|
|
|
|
|
|
|
|
|
# Construct DATABASE_URL from the provided variables
|
|
|
|
|
database_url = f"postgresql://{database_username_enc}:{database_password_enc}@{database_host}/{database_name_enc}"
|
|
|
|
|
|
|
|
|
|
# Assert the correct URL was constructed with properly escaped characters
|
2025-12-23 09:03:53 +08:00
|
|
|
expected_url = "postgresql://user%40with%2Bspecial:test-password-special-chars@localhost:5432/db_name%2Ftest"
|
2025-05-20 00:29:12 +08:00
|
|
|
assert database_url == expected_url
|
|
|
|
|
|
|
|
|
|
# Test appending query parameters
|
2025-05-27 05:41:42 +08:00
|
|
|
params = {"connection_limit": 10, "pool_timeout": 60}
|
2025-05-20 00:29:12 +08:00
|
|
|
modified_url = append_query_params(database_url, params)
|
|
|
|
|
assert "connection_limit=10" in modified_url
|
|
|
|
|
assert "pool_timeout=60" in modified_url
|
|
|
|
|
|
2026-02-17 01:03:10 +08:00
|
|
|
def test_append_query_params_handles_missing_url(self):
|
|
|
|
|
from litellm.proxy.proxy_cli import append_query_params
|
|
|
|
|
|
|
|
|
|
modified_url = append_query_params(None, {"connection_limit": 10})
|
|
|
|
|
assert modified_url == ""
|
|
|
|
|
|
2025-05-19 23:00:41 +08:00
|
|
|
@patch("uvicorn.run")
|
2026-02-20 06:16:07 +08:00
|
|
|
@patch("atexit.register") # critical
|
|
|
|
|
@patch("litellm.proxy.db.prisma_client.PrismaManager.setup_database")
|
2026-04-18 04:02:59 +08:00
|
|
|
@patch(
|
|
|
|
|
"litellm.proxy.db.prisma_client.should_update_prisma_schema", return_value=False
|
|
|
|
|
)
|
|
|
|
|
def test_skip_server_startup(
|
|
|
|
|
self, mock_should_update, mock_setup_db, mock_atexit_register, mock_uvicorn_run
|
|
|
|
|
):
|
2025-05-19 23:00:41 +08:00
|
|
|
from click.testing import CliRunner
|
2026-01-25 04:59:51 +08:00
|
|
|
|
2025-05-19 23:00:41 +08:00
|
|
|
from litellm.proxy.proxy_cli import run_server
|
2025-05-27 05:41:42 +08:00
|
|
|
|
2026-02-22 02:02:24 +08:00
|
|
|
runner = CliRunner()
|
2025-05-19 23:00:41 +08:00
|
|
|
|
2026-02-19 10:20:32 +08:00
|
|
|
mock_proxy_module = MagicMock(
|
|
|
|
|
app=MagicMock(),
|
|
|
|
|
ProxyConfig=MagicMock(),
|
|
|
|
|
KeyManagementSettings=MagicMock(),
|
|
|
|
|
save_worker_config=MagicMock(),
|
|
|
|
|
)
|
2026-02-22 02:02:24 +08:00
|
|
|
# Remove DATABASE_URL/DIRECT_URL so the CLI doesn't attempt
|
|
|
|
|
# real prisma operations when these are set in CI.
|
2026-04-18 04:02:59 +08:00
|
|
|
clean_env = {
|
|
|
|
|
k: v
|
|
|
|
|
for k, v in os.environ.items()
|
|
|
|
|
if k not in ("DATABASE_URL", "DIRECT_URL")
|
|
|
|
|
}
|
|
|
|
|
with (
|
|
|
|
|
patch.dict(
|
|
|
|
|
os.environ,
|
|
|
|
|
clean_env,
|
|
|
|
|
clear=True,
|
|
|
|
|
),
|
|
|
|
|
patch.dict(
|
|
|
|
|
"sys.modules",
|
|
|
|
|
{
|
|
|
|
|
"proxy_server": mock_proxy_module,
|
|
|
|
|
# Prevent real import of proxy_server inside Click's
|
|
|
|
|
# isolation context (heavy side effects cause stream
|
|
|
|
|
# lifecycle issues with Click 8.2+)
|
|
|
|
|
"litellm.proxy.proxy_server": mock_proxy_module,
|
|
|
|
|
},
|
|
|
|
|
),
|
|
|
|
|
patch(
|
|
|
|
|
"litellm.proxy.proxy_cli.ProxyInitializationHelpers._get_default_unvicorn_init_args"
|
|
|
|
|
) as mock_get_args,
|
|
|
|
|
):
|
2025-05-19 23:00:41 +08:00
|
|
|
mock_get_args.return_value = {
|
2025-05-27 05:41:42 +08:00
|
|
|
"app": "litellm.proxy.proxy_server:app",
|
|
|
|
|
"host": "localhost",
|
|
|
|
|
"port": 8000,
|
2025-05-19 23:00:41 +08:00
|
|
|
}
|
2025-05-27 05:41:42 +08:00
|
|
|
|
2026-01-24 04:10:08 +08:00
|
|
|
# --- skip startup ---
|
2025-05-19 23:00:41 +08:00
|
|
|
result = runner.invoke(run_server, ["--local", "--skip_server_startup"])
|
2025-05-27 05:41:42 +08:00
|
|
|
|
2026-04-18 04:02:59 +08:00
|
|
|
assert (
|
|
|
|
|
result.exit_code == 0
|
|
|
|
|
), f"exit_code={result.exit_code}, output={result.output}"
|
2026-01-24 04:10:08 +08:00
|
|
|
assert "Skipping server startup" in result.output
|
2025-05-19 23:00:41 +08:00
|
|
|
mock_uvicorn_run.assert_not_called()
|
2025-05-27 05:41:42 +08:00
|
|
|
|
2026-01-24 04:10:08 +08:00
|
|
|
# --- normal startup ---
|
2025-05-19 23:00:41 +08:00
|
|
|
mock_uvicorn_run.reset_mock()
|
2025-05-27 05:41:42 +08:00
|
|
|
|
2025-05-19 23:00:41 +08:00
|
|
|
result = runner.invoke(run_server, ["--local"])
|
2025-05-27 05:41:42 +08:00
|
|
|
|
2026-04-18 04:02:59 +08:00
|
|
|
assert (
|
|
|
|
|
result.exit_code == 0
|
|
|
|
|
), f"exit_code={result.exit_code}, output={result.output}"
|
2025-05-19 23:00:41 +08:00
|
|
|
mock_uvicorn_run.assert_called_once()
|
2025-06-11 04:30:19 +08:00
|
|
|
|
2026-05-12 05:49:38 +08:00
|
|
|
@pytest.mark.parametrize(
|
|
|
|
|
"timeout_config,expected_timeout",
|
|
|
|
|
[
|
|
|
|
|
({"database_connection_timeout": 30}, 30),
|
|
|
|
|
({"database_connection_pool_timeout": 45}, 45),
|
|
|
|
|
(
|
|
|
|
|
{
|
|
|
|
|
"database_connection_timeout": 30,
|
|
|
|
|
"database_connection_pool_timeout": 45,
|
|
|
|
|
},
|
|
|
|
|
30,
|
|
|
|
|
),
|
|
|
|
|
],
|
|
|
|
|
)
|
|
|
|
|
@patch("subprocess.run")
|
|
|
|
|
@patch("atexit.register")
|
|
|
|
|
@patch("litellm.proxy.db.prisma_client.PrismaManager.setup_database")
|
|
|
|
|
@patch(
|
|
|
|
|
"litellm.proxy.db.prisma_client.should_update_prisma_schema", return_value=False
|
|
|
|
|
)
|
|
|
|
|
def test_db_timeout_settings_are_forwarded_to_pool_timeout(
|
|
|
|
|
self,
|
|
|
|
|
mock_should_update,
|
|
|
|
|
mock_setup_db,
|
|
|
|
|
mock_atexit_register,
|
|
|
|
|
mock_subprocess_run,
|
|
|
|
|
timeout_config,
|
|
|
|
|
expected_timeout,
|
|
|
|
|
):
|
|
|
|
|
from click.testing import CliRunner
|
|
|
|
|
|
|
|
|
|
from litellm.proxy.proxy_cli import run_server
|
|
|
|
|
|
|
|
|
|
runner = CliRunner()
|
|
|
|
|
mock_subprocess_run.return_value = MagicMock(returncode=0)
|
|
|
|
|
|
|
|
|
|
mock_proxy_module = MagicMock(
|
|
|
|
|
app=MagicMock(),
|
|
|
|
|
ProxyConfig=MagicMock(),
|
|
|
|
|
KeyManagementSettings=MagicMock(),
|
|
|
|
|
save_worker_config=MagicMock(),
|
|
|
|
|
)
|
|
|
|
|
mock_proxy_module.ProxyConfig.return_value.get_config = AsyncMock(
|
|
|
|
|
return_value={
|
|
|
|
|
"general_settings": {
|
|
|
|
|
"database_url": "postgresql://test:test@localhost:5432/test",
|
|
|
|
|
"database_connection_pool_limit": 5,
|
|
|
|
|
**timeout_config,
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
clean_env = {
|
|
|
|
|
k: v
|
|
|
|
|
for k, v in os.environ.items()
|
|
|
|
|
if k not in ("DATABASE_URL", "DIRECT_URL")
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
with (
|
|
|
|
|
patch.dict(os.environ, clean_env, clear=True),
|
|
|
|
|
patch.dict(
|
|
|
|
|
"sys.modules",
|
|
|
|
|
{
|
|
|
|
|
"proxy_server": mock_proxy_module,
|
|
|
|
|
"litellm.proxy.proxy_server": mock_proxy_module,
|
|
|
|
|
},
|
|
|
|
|
),
|
|
|
|
|
patch(
|
|
|
|
|
"litellm.proxy.proxy_cli.ProxyInitializationHelpers._get_default_unvicorn_init_args"
|
|
|
|
|
) as mock_get_args,
|
|
|
|
|
patch(
|
|
|
|
|
"litellm.proxy.proxy_cli.append_query_params",
|
|
|
|
|
side_effect=lambda url, params: (
|
|
|
|
|
f"{url}?connection_limit={params['connection_limit']}&pool_timeout={params['pool_timeout']}"
|
|
|
|
|
),
|
|
|
|
|
) as mock_append_query_params,
|
|
|
|
|
):
|
|
|
|
|
mock_get_args.return_value = {
|
|
|
|
|
"app": "litellm.proxy.proxy_server:app",
|
|
|
|
|
"host": "localhost",
|
|
|
|
|
"port": 8000,
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
result = runner.invoke(
|
|
|
|
|
run_server,
|
|
|
|
|
["--local", "--config", "test-config.yaml", "--skip_server_startup"],
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
assert (
|
|
|
|
|
result.exit_code == 0
|
|
|
|
|
), f"exit_code={result.exit_code}, output={result.output}"
|
|
|
|
|
mock_append_query_params.assert_called()
|
|
|
|
|
appended_params = mock_append_query_params.call_args.args[1]
|
|
|
|
|
assert appended_params["connection_limit"] == 5
|
|
|
|
|
assert appended_params["pool_timeout"] == expected_timeout
|
|
|
|
|
|
2026-05-21 08:19:24 +08:00
|
|
|
def test_build_db_connection_url_params_defaults(self):
|
|
|
|
|
from litellm.proxy.proxy_cli import _build_db_connection_url_params
|
|
|
|
|
|
|
|
|
|
params = _build_db_connection_url_params(connection_limit=10, pool_timeout=60)
|
|
|
|
|
assert params == {"connection_limit": 10, "pool_timeout": 60}
|
|
|
|
|
|
|
|
|
|
def test_build_db_connection_url_params_omits_none_timeouts(self):
|
|
|
|
|
from litellm.proxy.proxy_cli import _build_db_connection_url_params
|
|
|
|
|
|
|
|
|
|
params = _build_db_connection_url_params(
|
|
|
|
|
connection_limit=10,
|
|
|
|
|
pool_timeout=60,
|
|
|
|
|
connect_timeout=None,
|
|
|
|
|
socket_timeout=None,
|
|
|
|
|
)
|
|
|
|
|
assert "connect_timeout" not in params
|
|
|
|
|
assert "socket_timeout" not in params
|
|
|
|
|
|
|
|
|
|
def test_build_db_connection_url_params_includes_optional_timeouts(self):
|
|
|
|
|
from litellm.proxy.proxy_cli import _build_db_connection_url_params
|
|
|
|
|
|
|
|
|
|
params = _build_db_connection_url_params(
|
|
|
|
|
connection_limit=10,
|
|
|
|
|
pool_timeout=60,
|
|
|
|
|
connect_timeout=15,
|
|
|
|
|
socket_timeout=120,
|
|
|
|
|
)
|
|
|
|
|
assert params["connect_timeout"] == 15
|
|
|
|
|
assert params["socket_timeout"] == 120
|
|
|
|
|
|
|
|
|
|
def test_build_db_connection_url_params_extras_override_defaults(self):
|
|
|
|
|
from litellm.proxy.proxy_cli import _build_db_connection_url_params
|
|
|
|
|
|
|
|
|
|
params = _build_db_connection_url_params(
|
|
|
|
|
connection_limit=10,
|
|
|
|
|
pool_timeout=60,
|
|
|
|
|
extra_params={
|
|
|
|
|
"pgbouncer": "true",
|
|
|
|
|
"statement_cache_size": 0,
|
|
|
|
|
"pool_timeout": 5,
|
|
|
|
|
},
|
|
|
|
|
)
|
|
|
|
|
assert params["pgbouncer"] == "true"
|
|
|
|
|
assert params["statement_cache_size"] == 0
|
|
|
|
|
assert params["pool_timeout"] == 5
|
|
|
|
|
|
|
|
|
|
@patch("subprocess.run")
|
|
|
|
|
@patch("atexit.register")
|
|
|
|
|
@patch("litellm.proxy.db.prisma_client.PrismaManager.setup_database")
|
|
|
|
|
@patch(
|
|
|
|
|
"litellm.proxy.db.prisma_client.should_update_prisma_schema", return_value=False
|
|
|
|
|
)
|
|
|
|
|
def test_db_connection_extra_params_forwarded_to_url(
|
|
|
|
|
self,
|
|
|
|
|
mock_should_update,
|
|
|
|
|
mock_setup_db,
|
|
|
|
|
mock_atexit_register,
|
|
|
|
|
mock_subprocess_run,
|
|
|
|
|
):
|
|
|
|
|
from click.testing import CliRunner
|
|
|
|
|
|
|
|
|
|
from litellm.proxy.proxy_cli import run_server
|
|
|
|
|
|
|
|
|
|
runner = CliRunner()
|
|
|
|
|
mock_subprocess_run.return_value = MagicMock(returncode=0)
|
|
|
|
|
|
|
|
|
|
mock_proxy_module = MagicMock(
|
|
|
|
|
app=MagicMock(),
|
|
|
|
|
ProxyConfig=MagicMock(),
|
|
|
|
|
KeyManagementSettings=MagicMock(),
|
|
|
|
|
save_worker_config=MagicMock(),
|
|
|
|
|
)
|
|
|
|
|
mock_proxy_module.ProxyConfig.return_value.get_config = AsyncMock(
|
|
|
|
|
return_value={
|
|
|
|
|
"general_settings": {
|
|
|
|
|
"database_url": "postgresql://test:test@localhost:5432/test",
|
|
|
|
|
"database_connect_timeout": 15,
|
|
|
|
|
"database_socket_timeout": 120,
|
|
|
|
|
"database_extra_connection_params": {
|
|
|
|
|
"pgbouncer": "true",
|
|
|
|
|
"statement_cache_size": 0,
|
|
|
|
|
},
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
clean_env = {
|
|
|
|
|
k: v
|
|
|
|
|
for k, v in os.environ.items()
|
|
|
|
|
if k not in ("DATABASE_URL", "DIRECT_URL")
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
with (
|
|
|
|
|
patch.dict(os.environ, clean_env, clear=True),
|
|
|
|
|
patch.dict(
|
|
|
|
|
"sys.modules",
|
|
|
|
|
{
|
|
|
|
|
"proxy_server": mock_proxy_module,
|
|
|
|
|
"litellm.proxy.proxy_server": mock_proxy_module,
|
|
|
|
|
},
|
|
|
|
|
),
|
|
|
|
|
patch(
|
|
|
|
|
"litellm.proxy.proxy_cli.ProxyInitializationHelpers._get_default_unvicorn_init_args"
|
|
|
|
|
) as mock_get_args,
|
|
|
|
|
patch(
|
|
|
|
|
"litellm.proxy.proxy_cli.append_query_params",
|
|
|
|
|
side_effect=lambda url, params: str(url),
|
|
|
|
|
) as mock_append_query_params,
|
|
|
|
|
):
|
|
|
|
|
mock_get_args.return_value = {
|
|
|
|
|
"app": "litellm.proxy.proxy_server:app",
|
|
|
|
|
"host": "localhost",
|
|
|
|
|
"port": 8000,
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
result = runner.invoke(
|
|
|
|
|
run_server,
|
|
|
|
|
["--local", "--config", "test-config.yaml", "--skip_server_startup"],
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
assert (
|
|
|
|
|
result.exit_code == 0
|
|
|
|
|
), f"exit_code={result.exit_code}, output={result.output}"
|
|
|
|
|
mock_append_query_params.assert_called()
|
|
|
|
|
appended_params = mock_append_query_params.call_args.args[1]
|
|
|
|
|
assert appended_params["connect_timeout"] == 15
|
|
|
|
|
assert appended_params["socket_timeout"] == 120
|
|
|
|
|
assert appended_params["pgbouncer"] == "true"
|
|
|
|
|
assert appended_params["statement_cache_size"] == 0
|
|
|
|
|
|
2026-06-11 07:06:32 +08:00
|
|
|
def test_build_db_connection_url_params_disable_prepared_statements(self):
|
|
|
|
|
from litellm.proxy.proxy_cli import _build_db_connection_url_params
|
|
|
|
|
|
|
|
|
|
params = _build_db_connection_url_params(
|
|
|
|
|
connection_limit=10,
|
|
|
|
|
pool_timeout=60,
|
|
|
|
|
disable_prepared_statements=True,
|
|
|
|
|
)
|
|
|
|
|
assert params["pgbouncer"] == "true"
|
|
|
|
|
|
|
|
|
|
def test_build_db_connection_url_params_no_pgbouncer_by_default(self):
|
|
|
|
|
from litellm.proxy.proxy_cli import _build_db_connection_url_params
|
|
|
|
|
|
|
|
|
|
params = _build_db_connection_url_params(
|
|
|
|
|
connection_limit=10,
|
|
|
|
|
pool_timeout=60,
|
|
|
|
|
)
|
|
|
|
|
assert "pgbouncer" not in params
|
|
|
|
|
|
|
|
|
|
def test_build_db_connection_url_params_extra_pgbouncer_overrides_flag(self):
|
|
|
|
|
from litellm.proxy.proxy_cli import _build_db_connection_url_params
|
|
|
|
|
|
|
|
|
|
params = _build_db_connection_url_params(
|
|
|
|
|
connection_limit=10,
|
|
|
|
|
pool_timeout=60,
|
|
|
|
|
disable_prepared_statements=True,
|
|
|
|
|
extra_params={"pgbouncer": "false"},
|
|
|
|
|
)
|
|
|
|
|
assert params["pgbouncer"] == "false"
|
|
|
|
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
|
|
|
"config_value, expect_pgbouncer",
|
|
|
|
|
[
|
|
|
|
|
(True, True),
|
|
|
|
|
(False, False),
|
|
|
|
|
("true", True),
|
|
|
|
|
("false", False),
|
|
|
|
|
("not-a-bool", False),
|
|
|
|
|
],
|
|
|
|
|
)
|
|
|
|
|
@patch("subprocess.run")
|
|
|
|
|
@patch("atexit.register")
|
|
|
|
|
@patch("litellm.proxy.db.prisma_client.PrismaManager.setup_database")
|
|
|
|
|
@patch(
|
|
|
|
|
"litellm.proxy.db.prisma_client.should_update_prisma_schema", return_value=False
|
|
|
|
|
)
|
|
|
|
|
def test_disable_prepared_statements_forwarded_to_url(
|
|
|
|
|
self,
|
|
|
|
|
mock_should_update,
|
|
|
|
|
mock_setup_db,
|
|
|
|
|
mock_atexit_register,
|
|
|
|
|
mock_subprocess_run,
|
|
|
|
|
config_value,
|
|
|
|
|
expect_pgbouncer,
|
|
|
|
|
):
|
|
|
|
|
from click.testing import CliRunner
|
|
|
|
|
|
|
|
|
|
from litellm.proxy.proxy_cli import run_server
|
|
|
|
|
|
|
|
|
|
runner = CliRunner()
|
|
|
|
|
mock_subprocess_run.return_value = MagicMock(returncode=0)
|
|
|
|
|
|
|
|
|
|
mock_proxy_module = MagicMock(
|
|
|
|
|
app=MagicMock(),
|
|
|
|
|
ProxyConfig=MagicMock(),
|
|
|
|
|
KeyManagementSettings=MagicMock(),
|
|
|
|
|
save_worker_config=MagicMock(),
|
|
|
|
|
)
|
|
|
|
|
mock_proxy_module.ProxyConfig.return_value.get_config = AsyncMock(
|
|
|
|
|
return_value={
|
|
|
|
|
"general_settings": {
|
|
|
|
|
"database_url": "postgresql://test:test@localhost:5432/test",
|
|
|
|
|
"database_disable_prepared_statements": config_value,
|
|
|
|
|
}
|
|
|
|
|
}
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
clean_env = {
|
|
|
|
|
k: v
|
|
|
|
|
for k, v in os.environ.items()
|
|
|
|
|
if k not in ("DATABASE_URL", "DIRECT_URL")
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
with (
|
|
|
|
|
patch.dict(os.environ, clean_env, clear=True),
|
|
|
|
|
patch.dict(
|
|
|
|
|
"sys.modules",
|
|
|
|
|
{
|
|
|
|
|
"proxy_server": mock_proxy_module,
|
|
|
|
|
"litellm.proxy.proxy_server": mock_proxy_module,
|
|
|
|
|
},
|
|
|
|
|
),
|
|
|
|
|
patch(
|
|
|
|
|
"litellm.proxy.proxy_cli.ProxyInitializationHelpers._get_default_unvicorn_init_args"
|
|
|
|
|
) as mock_get_args,
|
|
|
|
|
patch(
|
|
|
|
|
"litellm.proxy.proxy_cli.append_query_params",
|
|
|
|
|
side_effect=lambda url, params: str(url),
|
|
|
|
|
) as mock_append_query_params,
|
|
|
|
|
):
|
|
|
|
|
mock_get_args.return_value = {
|
|
|
|
|
"app": "litellm.proxy.proxy_server:app",
|
|
|
|
|
"host": "localhost",
|
|
|
|
|
"port": 8000,
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
result = runner.invoke(
|
|
|
|
|
run_server,
|
|
|
|
|
["--local", "--config", "test-config.yaml", "--skip_server_startup"],
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
assert (
|
|
|
|
|
result.exit_code == 0
|
|
|
|
|
), f"exit_code={result.exit_code}, output={result.output}"
|
|
|
|
|
mock_append_query_params.assert_called()
|
|
|
|
|
appended_params = mock_append_query_params.call_args.args[1]
|
|
|
|
|
if expect_pgbouncer:
|
|
|
|
|
assert appended_params["pgbouncer"] == "true"
|
|
|
|
|
else:
|
|
|
|
|
assert "pgbouncer" not in appended_params
|
|
|
|
|
|
2026-03-19 18:27:03 +08:00
|
|
|
@patch("uvicorn.run")
|
|
|
|
|
@patch("atexit.register")
|
|
|
|
|
@patch("litellm.proxy.db.prisma_client.PrismaManager.setup_database")
|
2026-04-18 04:02:59 +08:00
|
|
|
@patch(
|
|
|
|
|
"litellm.proxy.db.prisma_client.should_update_prisma_schema", return_value=False
|
|
|
|
|
)
|
2026-03-19 18:27:03 +08:00
|
|
|
def test_proxy_default_api_version_uses_azure_default(
|
|
|
|
|
self, mock_should_update, mock_setup_db, mock_atexit_register, mock_uvicorn_run
|
|
|
|
|
):
|
|
|
|
|
"""Proxy default api_version should match litellm.AZURE_DEFAULT_API_VERSION for consistency."""
|
|
|
|
|
from click.testing import CliRunner
|
|
|
|
|
|
|
|
|
|
import litellm
|
|
|
|
|
from litellm.proxy.proxy_cli import run_server
|
|
|
|
|
|
|
|
|
|
runner = CliRunner()
|
|
|
|
|
mock_proxy_module = MagicMock(
|
|
|
|
|
app=MagicMock(),
|
|
|
|
|
ProxyConfig=MagicMock(),
|
|
|
|
|
KeyManagementSettings=MagicMock(),
|
|
|
|
|
save_worker_config=MagicMock(),
|
|
|
|
|
)
|
2026-04-18 04:02:59 +08:00
|
|
|
clean_env = {
|
|
|
|
|
k: v
|
|
|
|
|
for k, v in os.environ.items()
|
|
|
|
|
if k not in ("DATABASE_URL", "DIRECT_URL")
|
|
|
|
|
}
|
|
|
|
|
with (
|
|
|
|
|
patch.dict(os.environ, clean_env, clear=True),
|
|
|
|
|
patch.dict(
|
|
|
|
|
"sys.modules",
|
|
|
|
|
{
|
|
|
|
|
"proxy_server": mock_proxy_module,
|
|
|
|
|
"litellm.proxy.proxy_server": mock_proxy_module,
|
|
|
|
|
},
|
|
|
|
|
),
|
|
|
|
|
patch(
|
|
|
|
|
"litellm.proxy.proxy_cli.ProxyInitializationHelpers._get_default_unvicorn_init_args"
|
|
|
|
|
) as mock_get_args,
|
|
|
|
|
):
|
2026-03-19 18:27:03 +08:00
|
|
|
mock_get_args.return_value = {
|
|
|
|
|
"app": "litellm.proxy.proxy_server:app",
|
|
|
|
|
"host": "localhost",
|
|
|
|
|
"port": 8000,
|
|
|
|
|
}
|
|
|
|
|
result = runner.invoke(run_server, ["--local", "--skip_server_startup"])
|
2026-04-18 04:02:59 +08:00
|
|
|
assert (
|
|
|
|
|
result.exit_code == 0
|
|
|
|
|
), f"exit_code={result.exit_code}, output={result.output}"
|
2026-03-19 18:27:03 +08:00
|
|
|
mock_proxy_module.save_worker_config.assert_called_once()
|
|
|
|
|
call_kwargs = mock_proxy_module.save_worker_config.call_args[1]
|
|
|
|
|
assert call_kwargs["api_version"] == litellm.AZURE_DEFAULT_API_VERSION
|
|
|
|
|
|
2025-06-11 04:30:19 +08:00
|
|
|
@patch("uvicorn.run")
|
|
|
|
|
@patch("builtins.print")
|
2026-05-16 07:51:45 +08:00
|
|
|
@patch("litellm.proxy.db.prisma_client.PrismaManager.setup_database")
|
|
|
|
|
@patch(
|
|
|
|
|
"litellm.proxy.db.prisma_client.should_update_prisma_schema", return_value=False
|
|
|
|
|
)
|
|
|
|
|
def test_keepalive_timeout_flag(
|
|
|
|
|
self, mock_should_update, mock_setup_db, mock_print, mock_uvicorn_run
|
|
|
|
|
):
|
2025-06-11 04:30:19 +08:00
|
|
|
"""Test that the keepalive_timeout flag is properly passed to uvicorn"""
|
|
|
|
|
from click.testing import CliRunner
|
|
|
|
|
|
|
|
|
|
from litellm.proxy.proxy_cli import run_server
|
|
|
|
|
|
|
|
|
|
runner = CliRunner()
|
|
|
|
|
|
|
|
|
|
mock_app = MagicMock()
|
|
|
|
|
mock_proxy_config = MagicMock()
|
|
|
|
|
mock_key_mgmt = MagicMock()
|
|
|
|
|
mock_save_worker_config = MagicMock()
|
|
|
|
|
|
2026-05-16 07:51:45 +08:00
|
|
|
# Strip DATABASE_URL/DIRECT_URL so run_server doesn't enter the prisma
|
|
|
|
|
# DB-setup block (un-timeout'd `subprocess.run(["prisma"])` +
|
|
|
|
|
# migrate-deploy retry loop) — same isolation every other run_server
|
|
|
|
|
# test in this file uses.
|
|
|
|
|
clean_env = {
|
|
|
|
|
k: v
|
|
|
|
|
for k, v in os.environ.items()
|
|
|
|
|
if k not in ("DATABASE_URL", "DIRECT_URL")
|
|
|
|
|
}
|
|
|
|
|
|
2026-04-18 04:02:59 +08:00
|
|
|
with (
|
2026-05-16 07:51:45 +08:00
|
|
|
patch.dict(os.environ, clean_env, clear=True),
|
2026-04-18 04:02:59 +08:00
|
|
|
patch.dict(
|
|
|
|
|
"sys.modules",
|
|
|
|
|
{
|
|
|
|
|
"proxy_server": MagicMock(
|
|
|
|
|
app=mock_app,
|
|
|
|
|
ProxyConfig=mock_proxy_config,
|
|
|
|
|
KeyManagementSettings=mock_key_mgmt,
|
|
|
|
|
save_worker_config=mock_save_worker_config,
|
|
|
|
|
)
|
|
|
|
|
},
|
|
|
|
|
),
|
|
|
|
|
patch(
|
|
|
|
|
"litellm.proxy.proxy_cli.ProxyInitializationHelpers._get_default_unvicorn_init_args"
|
|
|
|
|
) as mock_get_args,
|
|
|
|
|
patch(
|
|
|
|
|
"litellm.proxy.proxy_cli.ProxyInitializationHelpers._is_port_in_use",
|
|
|
|
|
return_value=False,
|
|
|
|
|
),
|
build: migrate packaging, CI, and Docker from Poetry to uv (#25007)
* build: migrate packaging metadata to uv
* ci: move automation and local tooling to uv
* docker: migrate image builds and runtime setup to uv
* docs: update install and deployment guidance for uv
* chore: align auxiliary scripts and tests with uv
* test: harden test_litellm isolation
* fix: keep release and health check images self-contained
* build: pin uv tooling and health check deps
* test: isolate bedrock image request formatting from suite state
* test: cover sandbox executor requirements flow
* ci: fix circleci no-op command steps
* ci: fix circleci publish workflow parsing
* fix: stabilize remaining uv migration CI checks
* ci: increase matrix test timeout headroom
* fix: restore published docker and license coverage
* fix: restore proxy runtime build parity
* fix: restore proxy extras parity and venv migrations
* ci: persist uv path across circleci steps
* fix: keep psycopg binary in default test env
* docker: preserve prisma cache across stages
* test: run local proxy checks through uv python
* build: restore runtime deps moved into ci
* build: refresh uv lock after upstream merge
* fix: restore module import in test_check_migration after merge
The conflict resolution imported only the function but the test body
references check_migration as a module throughout.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: revert dependency promotions, remove nodejs-wheel-binaries, fix Docker layer caching
- Move google-generativeai, Pillow, tenacity back to ci group (they are
lazily imported and bloat the base SDK install needlessly)
- Remove nodejs-wheel-binaries from extra_proxy and proxy-dev (redundant
in Docker where system Node.js is already installed via apk)
- Remove all nodejs-wheel node replacement and venv npm patching blocks
from Dockerfiles since the wheel is no longer installed
- Add --no-default-groups to CodSpeed benchmark workflow so the benchmark
environment matches the old minimal pip install footprint
- Apply standard uv two-phase Docker pattern: copy metadata first, install
deps (cached layer), then copy source and install project
- Replace CircleCI enterprise no-op with proper uv sync command
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* chore: regenerate uv.lock after removing nodejs-wheel-binaries
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(ci): use cache/restore instead of cache to prevent cache poisoning
The old workflow used actions/cache/restore (read-only). The uv migration
changed it to actions/cache (read-write), which zizmor flags as a cache
poisoning risk. Restore the safer read-only variant.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(ci): disable setup-uv built-in cache to silence cache-poisoning alert
The setup-uv action enables caching by default, which zizmor flags as a
cache poisoning risk. Disable it since we already use a read-only
cache/restore step.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(ci): disable setup-uv cache in publish workflow
Silences zizmor cache-poisoning alert. Publishing workflow runs
infrequently on protected branches so caching adds no real benefit.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(test): remove duplicate verbose_logger mock in test_check_migration
The logger was patched twice — first via mocker.patch() then via
mocker.patch.object(autospec=True). The second call fails because
autospec cannot inspect an already-mocked attribute. Remove the
redundant first patch.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(ci): free disk space before Docker build in test-server-root-path
The Dockerfile.non_root build ran out of disk on the CI runner. Remove
Android SDK, .NET, Boost, and GHC toolchains (~12GB) to free space.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-10 02:46:23 +08:00
|
|
|
):
|
2025-06-11 04:30:19 +08:00
|
|
|
mock_get_args.return_value = {
|
|
|
|
|
"app": "litellm.proxy.proxy_server:app",
|
|
|
|
|
"host": "localhost",
|
|
|
|
|
"port": 8000,
|
|
|
|
|
"timeout_keep_alive": 30,
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
result = runner.invoke(run_server, ["--local", "--keepalive_timeout", "30"])
|
|
|
|
|
|
|
|
|
|
assert result.exit_code == 0
|
|
|
|
|
mock_get_args.assert_called_once_with(
|
|
|
|
|
host="0.0.0.0",
|
|
|
|
|
port=4000,
|
|
|
|
|
log_config=None,
|
|
|
|
|
keepalive_timeout=30,
|
2026-04-28 02:06:56 +08:00
|
|
|
timeout_worker_healthcheck=None,
|
2025-06-11 04:30:19 +08:00
|
|
|
)
|
|
|
|
|
mock_uvicorn_run.assert_called_once()
|
2025-08-13 13:03:39 +08:00
|
|
|
|
2025-06-11 04:30:19 +08:00
|
|
|
# Check that the uvicorn.run was called with the timeout_keep_alive parameter
|
|
|
|
|
call_args = mock_uvicorn_run.call_args
|
|
|
|
|
assert call_args[1]["timeout_keep_alive"] == 30
|
2025-07-17 11:35:09 +08:00
|
|
|
|
2026-04-28 02:06:56 +08:00
|
|
|
@patch("uvicorn.run")
|
|
|
|
|
@patch("builtins.print")
|
2026-05-16 07:51:45 +08:00
|
|
|
@patch("litellm.proxy.db.prisma_client.PrismaManager.setup_database")
|
|
|
|
|
@patch(
|
|
|
|
|
"litellm.proxy.db.prisma_client.should_update_prisma_schema", return_value=False
|
|
|
|
|
)
|
|
|
|
|
def test_timeout_worker_healthcheck_flag(
|
|
|
|
|
self, mock_should_update, mock_setup_db, mock_print, mock_uvicorn_run
|
|
|
|
|
):
|
2026-04-28 02:06:56 +08:00
|
|
|
"""Test that the --timeout_worker_healthcheck flag is threaded through to the uvicorn init helper."""
|
|
|
|
|
from click.testing import CliRunner
|
|
|
|
|
|
|
|
|
|
from litellm.proxy.proxy_cli import run_server
|
|
|
|
|
|
|
|
|
|
runner = CliRunner()
|
|
|
|
|
|
|
|
|
|
mock_app = MagicMock()
|
|
|
|
|
mock_proxy_config = MagicMock()
|
|
|
|
|
mock_key_mgmt = MagicMock()
|
|
|
|
|
mock_save_worker_config = MagicMock()
|
|
|
|
|
|
2026-05-16 07:51:45 +08:00
|
|
|
# Strip DATABASE_URL/DIRECT_URL so run_server doesn't enter the prisma
|
|
|
|
|
# DB-setup block (un-timeout'd `subprocess.run(["prisma"])` +
|
|
|
|
|
# migrate-deploy retry loop) — same isolation every other run_server
|
|
|
|
|
# test in this file uses.
|
|
|
|
|
clean_env = {
|
|
|
|
|
k: v
|
|
|
|
|
for k, v in os.environ.items()
|
|
|
|
|
if k not in ("DATABASE_URL", "DIRECT_URL")
|
|
|
|
|
}
|
|
|
|
|
|
2026-04-28 02:06:56 +08:00
|
|
|
with (
|
2026-05-16 07:51:45 +08:00
|
|
|
patch.dict(os.environ, clean_env, clear=True),
|
2026-04-28 02:06:56 +08:00
|
|
|
patch.dict(
|
|
|
|
|
"sys.modules",
|
|
|
|
|
{
|
|
|
|
|
"proxy_server": MagicMock(
|
|
|
|
|
app=mock_app,
|
|
|
|
|
ProxyConfig=mock_proxy_config,
|
|
|
|
|
KeyManagementSettings=mock_key_mgmt,
|
|
|
|
|
save_worker_config=mock_save_worker_config,
|
|
|
|
|
)
|
|
|
|
|
},
|
|
|
|
|
),
|
|
|
|
|
patch(
|
|
|
|
|
"litellm.proxy.proxy_cli.ProxyInitializationHelpers._get_default_unvicorn_init_args"
|
|
|
|
|
) as mock_get_args,
|
|
|
|
|
patch(
|
|
|
|
|
"litellm.proxy.proxy_cli.ProxyInitializationHelpers._is_port_in_use",
|
|
|
|
|
return_value=False,
|
|
|
|
|
),
|
|
|
|
|
):
|
|
|
|
|
mock_get_args.return_value = {
|
|
|
|
|
"app": "litellm.proxy.proxy_server:app",
|
|
|
|
|
"host": "localhost",
|
|
|
|
|
"port": 8000,
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
result = runner.invoke(
|
|
|
|
|
run_server, ["--local", "--timeout_worker_healthcheck", "15"]
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
assert result.exit_code == 0
|
|
|
|
|
mock_get_args.assert_called_once_with(
|
|
|
|
|
host="0.0.0.0",
|
|
|
|
|
port=4000,
|
|
|
|
|
log_config=None,
|
|
|
|
|
keepalive_timeout=None,
|
|
|
|
|
timeout_worker_healthcheck=15,
|
|
|
|
|
)
|
|
|
|
|
|
2025-09-29 00:13:40 +08:00
|
|
|
@patch("uvicorn.run")
|
|
|
|
|
@patch("builtins.print")
|
2026-02-26 15:12:27 +08:00
|
|
|
@patch("litellm.proxy.db.prisma_client.PrismaManager.setup_database")
|
2026-04-18 04:02:59 +08:00
|
|
|
def test_max_requests_before_restart_flag(
|
|
|
|
|
self, mock_setup_db, mock_print, mock_uvicorn_run
|
|
|
|
|
):
|
2025-09-29 00:13:40 +08:00
|
|
|
"""Test that the max_requests_before_restart flag is passed to uvicorn as limit_max_requests"""
|
|
|
|
|
from click.testing import CliRunner
|
|
|
|
|
|
|
|
|
|
from litellm.proxy.proxy_cli import run_server
|
|
|
|
|
|
|
|
|
|
runner = CliRunner()
|
|
|
|
|
|
|
|
|
|
mock_app = MagicMock()
|
|
|
|
|
mock_proxy_config = MagicMock()
|
|
|
|
|
mock_key_mgmt = MagicMock()
|
|
|
|
|
mock_save_worker_config = MagicMock()
|
|
|
|
|
|
2026-04-18 04:02:59 +08:00
|
|
|
clean_env = {
|
|
|
|
|
k: v
|
|
|
|
|
for k, v in os.environ.items()
|
|
|
|
|
if k not in ("DATABASE_URL", "DIRECT_URL")
|
|
|
|
|
}
|
|
|
|
|
with (
|
|
|
|
|
patch.dict(
|
|
|
|
|
os.environ,
|
|
|
|
|
clean_env,
|
|
|
|
|
clear=True,
|
|
|
|
|
),
|
|
|
|
|
patch.dict(
|
|
|
|
|
"sys.modules",
|
|
|
|
|
{
|
|
|
|
|
"proxy_server": MagicMock(
|
|
|
|
|
app=mock_app,
|
|
|
|
|
ProxyConfig=mock_proxy_config,
|
|
|
|
|
KeyManagementSettings=mock_key_mgmt,
|
|
|
|
|
save_worker_config=mock_save_worker_config,
|
|
|
|
|
)
|
|
|
|
|
},
|
|
|
|
|
),
|
|
|
|
|
patch(
|
|
|
|
|
"litellm.proxy.proxy_cli.ProxyInitializationHelpers._get_default_unvicorn_init_args"
|
|
|
|
|
) as mock_get_args,
|
|
|
|
|
):
|
2025-09-29 00:13:40 +08:00
|
|
|
mock_get_args.return_value = {
|
|
|
|
|
"app": "litellm.proxy.proxy_server:app",
|
|
|
|
|
"host": "localhost",
|
|
|
|
|
"port": 8000,
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
result = runner.invoke(
|
|
|
|
|
run_server, ["--local", "--max_requests_before_restart", "123"]
|
|
|
|
|
)
|
|
|
|
|
|
2026-04-18 04:02:59 +08:00
|
|
|
assert (
|
|
|
|
|
result.exit_code == 0
|
|
|
|
|
), f"exit_code={result.exit_code}, output={result.output}"
|
2025-09-29 00:13:40 +08:00
|
|
|
mock_uvicorn_run.assert_called_once()
|
|
|
|
|
|
|
|
|
|
# Check that uvicorn.run was called with limit_max_requests parameter
|
|
|
|
|
call_args = mock_uvicorn_run.call_args
|
|
|
|
|
assert call_args[1]["limit_max_requests"] == 123
|
|
|
|
|
|
2025-08-01 04:52:56 +08:00
|
|
|
@patch.dict(os.environ, {}, clear=True)
|
|
|
|
|
def test_construct_database_url_from_env_vars(self):
|
|
|
|
|
"""Test the construct_database_url_from_env_vars function with various scenarios"""
|
|
|
|
|
from litellm.proxy.utils import construct_database_url_from_env_vars
|
|
|
|
|
|
|
|
|
|
# Test with all required variables present
|
|
|
|
|
test_env = {
|
|
|
|
|
"DATABASE_HOST": "localhost:5432",
|
|
|
|
|
"DATABASE_USERNAME": "testuser",
|
|
|
|
|
"DATABASE_PASSWORD": "testpass",
|
|
|
|
|
"DATABASE_NAME": "testdb",
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
with patch.dict(os.environ, test_env):
|
|
|
|
|
result = construct_database_url_from_env_vars()
|
|
|
|
|
expected_url = "postgresql://testuser:testpass@localhost:5432/testdb"
|
|
|
|
|
assert result == expected_url
|
|
|
|
|
|
|
|
|
|
# Test with special characters that need URL encoding
|
|
|
|
|
test_env_special = {
|
|
|
|
|
"DATABASE_HOST": "localhost:5432",
|
|
|
|
|
"DATABASE_USERNAME": "user@with+special",
|
2025-12-23 09:03:53 +08:00
|
|
|
"DATABASE_PASSWORD": "test-password-special-chars",
|
2025-08-01 04:52:56 +08:00
|
|
|
"DATABASE_NAME": "db_name/test",
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
with patch.dict(os.environ, test_env_special):
|
|
|
|
|
result = construct_database_url_from_env_vars()
|
2025-12-23 09:03:53 +08:00
|
|
|
expected_url = "postgresql://user%40with%2Bspecial:test-password-special-chars@localhost:5432/db_name%2Ftest"
|
2025-08-01 04:52:56 +08:00
|
|
|
assert result == expected_url
|
|
|
|
|
|
|
|
|
|
# Test without password (should still work)
|
|
|
|
|
test_env_no_password = {
|
|
|
|
|
"DATABASE_HOST": "localhost:5432",
|
|
|
|
|
"DATABASE_USERNAME": "testuser",
|
|
|
|
|
"DATABASE_NAME": "testdb",
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
with patch.dict(os.environ, test_env_no_password):
|
|
|
|
|
result = construct_database_url_from_env_vars()
|
|
|
|
|
expected_url = "postgresql://testuser@localhost:5432/testdb"
|
|
|
|
|
assert result == expected_url
|
|
|
|
|
|
|
|
|
|
# Test with missing required variables (should return None)
|
|
|
|
|
test_env_missing = {
|
|
|
|
|
"DATABASE_HOST": "localhost:5432",
|
|
|
|
|
"DATABASE_USERNAME": "testuser",
|
|
|
|
|
# Missing DATABASE_NAME
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
with patch.dict(os.environ, test_env_missing):
|
|
|
|
|
result = construct_database_url_from_env_vars()
|
|
|
|
|
assert result is None
|
|
|
|
|
|
|
|
|
|
# Test with empty environment (should return None)
|
|
|
|
|
with patch.dict(os.environ, {}, clear=True):
|
|
|
|
|
result = construct_database_url_from_env_vars()
|
|
|
|
|
assert result is None
|
|
|
|
|
|
|
|
|
|
@patch("uvicorn.run")
|
|
|
|
|
@patch("builtins.print")
|
|
|
|
|
def test_run_server_no_config_passed(self, mock_print, mock_uvicorn_run):
|
|
|
|
|
"""Test that run_server properly handles the case when no config is passed"""
|
2025-08-13 13:03:39 +08:00
|
|
|
import asyncio
|
|
|
|
|
|
2025-08-01 04:52:56 +08:00
|
|
|
from click.testing import CliRunner
|
2025-08-13 13:03:39 +08:00
|
|
|
|
2025-08-01 04:52:56 +08:00
|
|
|
from litellm.proxy.proxy_cli import run_server
|
|
|
|
|
|
|
|
|
|
runner = CliRunner()
|
|
|
|
|
|
|
|
|
|
mock_app = MagicMock()
|
|
|
|
|
mock_proxy_config = MagicMock()
|
|
|
|
|
mock_key_mgmt = MagicMock()
|
|
|
|
|
mock_save_worker_config = MagicMock()
|
|
|
|
|
|
|
|
|
|
# Mock the ProxyConfig.get_config method to return a proper async config
|
|
|
|
|
async def mock_get_config(config_file_path=None):
|
2025-08-13 13:03:39 +08:00
|
|
|
return {"general_settings": {}, "litellm_settings": {}}
|
|
|
|
|
|
2025-08-01 04:52:56 +08:00
|
|
|
mock_proxy_config_instance = MagicMock()
|
|
|
|
|
mock_proxy_config_instance.get_config = mock_get_config
|
|
|
|
|
mock_proxy_config.return_value = mock_proxy_config_instance
|
|
|
|
|
|
2026-02-17 06:45:05 +08:00
|
|
|
mock_proxy_server_module = MagicMock(app=mock_app)
|
|
|
|
|
|
|
|
|
|
# Only remove DATABASE_URL and DIRECT_URL to prevent the database setup
|
|
|
|
|
# code path from running. Do NOT use clear=True as it removes PATH, HOME,
|
|
|
|
|
# etc., which causes imports inside run_server to break in CI (the real
|
|
|
|
|
# litellm.proxy.proxy_server import at line 820 of proxy_cli.py has heavy
|
|
|
|
|
# side effects that fail without a proper environment).
|
|
|
|
|
env_overrides = {
|
|
|
|
|
"DATABASE_URL": "",
|
|
|
|
|
"DIRECT_URL": "",
|
|
|
|
|
"IAM_TOKEN_DB_AUTH": "",
|
|
|
|
|
"USE_AWS_KMS": "",
|
|
|
|
|
}
|
|
|
|
|
with patch.dict(os.environ, env_overrides):
|
|
|
|
|
# Remove DATABASE_URL entirely so the DB setup block is skipped
|
|
|
|
|
os.environ.pop("DATABASE_URL", None)
|
|
|
|
|
os.environ.pop("DIRECT_URL", None)
|
|
|
|
|
|
2026-04-18 04:02:59 +08:00
|
|
|
with (
|
|
|
|
|
patch.dict(
|
|
|
|
|
"sys.modules",
|
|
|
|
|
{
|
|
|
|
|
"proxy_server": MagicMock(
|
|
|
|
|
app=mock_app,
|
|
|
|
|
ProxyConfig=mock_proxy_config,
|
|
|
|
|
KeyManagementSettings=mock_key_mgmt,
|
|
|
|
|
save_worker_config=mock_save_worker_config,
|
|
|
|
|
),
|
|
|
|
|
# Also mock litellm.proxy.proxy_server to prevent the real
|
|
|
|
|
# import at line 820 of proxy_cli.py which has heavy side
|
|
|
|
|
# effects (FastAPI app init, logging setup, etc.)
|
|
|
|
|
"litellm.proxy.proxy_server": mock_proxy_server_module,
|
|
|
|
|
},
|
|
|
|
|
),
|
|
|
|
|
patch(
|
|
|
|
|
"litellm.proxy.proxy_cli.ProxyInitializationHelpers._get_default_unvicorn_init_args"
|
|
|
|
|
) as mock_get_args,
|
|
|
|
|
):
|
2025-08-01 04:52:56 +08:00
|
|
|
mock_get_args.return_value = {
|
|
|
|
|
"app": "litellm.proxy.proxy_server:app",
|
|
|
|
|
"host": "localhost",
|
|
|
|
|
"port": 8000,
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
# Test with no config parameter (config=None)
|
|
|
|
|
result = runner.invoke(run_server, ["--local"])
|
|
|
|
|
|
2026-02-17 06:45:05 +08:00
|
|
|
assert result.exit_code == 0, (
|
|
|
|
|
f"run_server failed with exit_code={result.exit_code}, "
|
|
|
|
|
f"output={result.output}, exception={result.exception}"
|
|
|
|
|
)
|
2025-08-13 13:03:39 +08:00
|
|
|
|
2025-08-01 04:52:56 +08:00
|
|
|
# Verify that uvicorn.run was called
|
|
|
|
|
mock_uvicorn_run.assert_called_once()
|
|
|
|
|
|
|
|
|
|
# Reset mocks for second test
|
|
|
|
|
mock_uvicorn_run.reset_mock()
|
|
|
|
|
|
|
|
|
|
# Test with explicit --config None (should behave the same)
|
|
|
|
|
result = runner.invoke(run_server, ["--local", "--config", "None"])
|
|
|
|
|
|
2026-02-17 06:45:05 +08:00
|
|
|
assert result.exit_code == 0, (
|
|
|
|
|
f"run_server failed with exit_code={result.exit_code}, "
|
|
|
|
|
f"output={result.output}, exception={result.exception}"
|
|
|
|
|
)
|
2025-08-13 13:03:39 +08:00
|
|
|
|
2025-08-01 04:52:56 +08:00
|
|
|
# Verify that uvicorn.run was called again
|
|
|
|
|
mock_uvicorn_run.assert_called_once()
|
|
|
|
|
|
2025-07-17 11:35:09 +08:00
|
|
|
|
2026-05-08 07:04:56 +08:00
|
|
|
class TestRunServerDbSetup:
|
|
|
|
|
"""Tests for run_server's prisma setup_database behavior."""
|
2025-08-13 13:03:39 +08:00
|
|
|
|
|
|
|
|
@patch("subprocess.run")
|
2026-02-22 07:23:55 +08:00
|
|
|
@patch("atexit.register")
|
2025-08-13 13:03:39 +08:00
|
|
|
@patch("litellm.proxy.db.prisma_client.PrismaManager.setup_database")
|
|
|
|
|
@patch("litellm.proxy.db.check_migration.check_prisma_schema_diff")
|
|
|
|
|
@patch("litellm.proxy.db.prisma_client.should_update_prisma_schema")
|
|
|
|
|
def test_use_prisma_db_push_flag_behavior(
|
|
|
|
|
self,
|
|
|
|
|
mock_should_update_schema,
|
|
|
|
|
mock_check_schema_diff,
|
|
|
|
|
mock_setup_database,
|
2026-02-22 06:28:02 +08:00
|
|
|
mock_atexit_register,
|
2025-08-13 13:03:39 +08:00
|
|
|
mock_subprocess_run,
|
|
|
|
|
):
|
|
|
|
|
"""Test that use_prisma_db_push flag correctly controls PrismaManager.setup_database use_migrate parameter"""
|
|
|
|
|
from litellm.proxy.proxy_cli import run_server
|
|
|
|
|
|
|
|
|
|
# Mock subprocess.run to simulate prisma being available
|
|
|
|
|
mock_subprocess_run.return_value = MagicMock(returncode=0)
|
|
|
|
|
|
|
|
|
|
# Mock should_update_prisma_schema to return True (so setup_database gets called)
|
|
|
|
|
mock_should_update_schema.return_value = True
|
|
|
|
|
|
2026-02-20 06:16:07 +08:00
|
|
|
mock_proxy_module = MagicMock(
|
|
|
|
|
app=MagicMock(),
|
|
|
|
|
ProxyConfig=MagicMock(),
|
|
|
|
|
KeyManagementSettings=MagicMock(),
|
|
|
|
|
save_worker_config=MagicMock(),
|
|
|
|
|
)
|
2025-08-13 13:03:39 +08:00
|
|
|
|
2026-02-22 06:11:48 +08:00
|
|
|
clean_env = {
|
|
|
|
|
k: v
|
|
|
|
|
for k, v in os.environ.items()
|
|
|
|
|
if k not in ("DATABASE_URL", "DIRECT_URL")
|
|
|
|
|
}
|
|
|
|
|
clean_env["DATABASE_URL"] = "postgresql://test:test@localhost:5432/test"
|
|
|
|
|
|
2026-04-18 04:02:59 +08:00
|
|
|
with (
|
|
|
|
|
patch.dict(os.environ, clean_env, clear=True),
|
|
|
|
|
patch.dict(
|
|
|
|
|
"sys.modules",
|
|
|
|
|
{
|
|
|
|
|
"proxy_server": mock_proxy_module,
|
|
|
|
|
"litellm.proxy.proxy_server": mock_proxy_module,
|
|
|
|
|
},
|
|
|
|
|
),
|
|
|
|
|
patch(
|
|
|
|
|
"litellm.proxy.proxy_cli.ProxyInitializationHelpers._get_default_unvicorn_init_args"
|
|
|
|
|
) as mock_get_args,
|
|
|
|
|
):
|
2025-08-13 13:03:39 +08:00
|
|
|
mock_get_args.return_value = {
|
|
|
|
|
"app": "litellm.proxy.proxy_server:app",
|
|
|
|
|
"host": "localhost",
|
|
|
|
|
"port": 8000,
|
|
|
|
|
}
|
|
|
|
|
|
2026-02-22 07:23:55 +08:00
|
|
|
# Use standalone_mode=False to bypass Click's CliRunner stream
|
|
|
|
|
# isolation which causes flaky "I/O operation on closed file"
|
|
|
|
|
# errors in CI environments (Click 8.3.x stream lifecycle issue).
|
|
|
|
|
|
2025-08-13 13:03:39 +08:00
|
|
|
# Test 1: Without --use_prisma_db_push flag (default behavior)
|
|
|
|
|
# use_prisma_db_push should be False (default), so use_migrate should be True
|
2026-04-18 04:02:59 +08:00
|
|
|
run_server.main(["--local", "--skip_server_startup"], standalone_mode=False)
|
2026-04-22 05:45:29 +08:00
|
|
|
mock_setup_database.assert_called_with(
|
|
|
|
|
use_migrate=True, use_v2_resolver=False
|
|
|
|
|
)
|
2025-08-13 13:03:39 +08:00
|
|
|
|
|
|
|
|
# Reset mocks
|
|
|
|
|
mock_setup_database.reset_mock()
|
|
|
|
|
mock_should_update_schema.reset_mock()
|
|
|
|
|
mock_should_update_schema.return_value = True
|
|
|
|
|
|
|
|
|
|
# Test 2: With --use_prisma_db_push flag set
|
|
|
|
|
# use_prisma_db_push should be True, so use_migrate should be False
|
2026-02-22 07:23:55 +08:00
|
|
|
run_server.main(
|
|
|
|
|
["--local", "--skip_server_startup", "--use_prisma_db_push"],
|
|
|
|
|
standalone_mode=False,
|
2025-08-13 13:03:39 +08:00
|
|
|
)
|
2026-04-22 05:45:29 +08:00
|
|
|
mock_setup_database.assert_called_with(
|
|
|
|
|
use_migrate=False, use_v2_resolver=False
|
|
|
|
|
)
|
2026-03-07 10:09:57 +08:00
|
|
|
|
2026-03-10 20:45:11 +08:00
|
|
|
@patch("subprocess.run")
|
|
|
|
|
@patch("atexit.register")
|
|
|
|
|
@patch("litellm.proxy.db.prisma_client.PrismaManager.setup_database")
|
|
|
|
|
@patch("litellm.proxy.db.check_migration.check_prisma_schema_diff")
|
|
|
|
|
@patch("litellm.proxy.db.prisma_client.should_update_prisma_schema")
|
|
|
|
|
def test_startup_fails_when_db_setup_fails(
|
|
|
|
|
self,
|
|
|
|
|
mock_should_update_schema,
|
|
|
|
|
mock_check_schema_diff,
|
|
|
|
|
mock_setup_database,
|
|
|
|
|
mock_atexit_register,
|
|
|
|
|
mock_subprocess_run,
|
|
|
|
|
):
|
[Infra] Merging RC Branch with Main (#23786)
* fix(test): add missing mocks for test_streamable_http_mcp_handler_mock
The test was missing mocks for extract_mcp_auth_context and set_auth_context,
causing the handler to fail silently in the except block instead of reaching
session_manager.handle_request. This mirrors the fix already applied to the
sibling test_sse_mcp_handler_mock.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(ci): route OpenAI models through chat completions in pass-through tests
The test_anthropic_messages_openai_model_streaming_cost_injection test fails
because the OpenAI Responses API returns 400 for requests routed through the
Anthropic Messages endpoint. Setting LITELLM_USE_CHAT_COMPLETIONS_URL_FOR_ANTHROPIC_MESSAGES=true
routes OpenAI models through the stable chat completions path instead.
Cost injection still works since it happens at the proxy level.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(ci): fix assemblyai custom auth and router wildcard test flakiness
1. custom_auth_basic.py: Add user_role='proxy_admin' so the custom auth
user can access management endpoints like /key/generate. The test
test_assemblyai_transcribe_with_non_admin_key was hidden behind an
earlier -x failure and was never reached before.
2. test_router_utils.py: Add flaky(retries=3) and increase sleep from 1s
to 2s for test_router_get_model_group_usage_wildcard_routes. The async
callback needs time to write usage to cache, and 1s is insufficient on
slower CI hardware.
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* ci: retrigger CI pipeline
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix(mypy): use LitellmUserRoles enum instead of raw string in custom_auth_basic
Fixes mypy error: Argument 'user_role' has incompatible type 'str'; expected 'LitellmUserRoles | None'
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
* fix: don't close HTTP/SDK clients on LLMClientCache eviction (#22926)
* fix: don't close HTTP/SDK clients on LLMClientCache eviction
Removing the _remove_key override that eagerly called aclose()/close()
on evicted clients. Evicted clients may still be held by in-flight
streaming requests; closing them causes:
RuntimeError: Cannot send a request, as the client has been closed.
This is a regression from commit fb72979432. Clients that are no longer
referenced will be garbage-collected naturally. Explicit shutdown cleanup
happens via close_litellm_async_clients().
Fixes production crashes after the 1-hour cache TTL expires.
* test: update LLMClientCache unit tests for no-close-on-eviction behavior
Flip the assertions: evicted clients must NOT be closed. Replace
test_remove_key_closes_async_client → test_remove_key_does_not_close_async_client
and equivalents for sync/eviction paths.
Add test_remove_key_removes_plain_values for non-client cache entries.
Remove test_background_tasks_cleaned_up_after_completion (no more _background_tasks).
Remove test_remove_key_no_event_loop variant that depended on old behavior.
* test: add e2e tests for OpenAI SDK client surviving cache eviction
Add two new e2e tests using real AsyncOpenAI clients:
- test_evicted_openai_sdk_client_stays_usable: verifies size-based eviction
doesn't close the client
- test_ttl_expired_openai_sdk_client_stays_usable: verifies TTL expiry
eviction doesn't close the client
Both tests sleep after eviction so any create_task()-based close would
have time to run, making the regression detectable.
Also expand the module docstring to explain why the sleep is required.
* docs(AGENTS.md): add rule — never close HTTP/SDK clients on cache eviction
* docs(CLAUDE.md): add HTTP client cache safety guideline
* [Fix] Install bsdmainutils for column command in security scans
The security_scans.sh script uses `column` to format vulnerability
output, but the package wasn't installed in the CI environment.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix: handle string callback values in prometheus multiproc setup
When callbacks are configured as a plain string (e.g., `callbacks: "my_callback"`)
instead of a list, the proxy crashes on startup with:
TypeError: can only concatenate str (not "list") to str
Normalize each callback setting to a list before concatenating.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* bump: version 1.82.2 → 1.82.3
* fix(test): update test_startup_fails_when_db_setup_fails for opt-in enforcement
The --enforce_prisma_migration_check flag is now required to trigger
sys.exit(1) on DB migration failure, after #23675 flipped the default
behavior to warn-and-continue.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
* fix(cost_calculator): use model name for per-request custom pricing when router_model_id has no pricing
When custom pricing is passed as per-request kwargs (input_cost_per_token/output_cost_per_token),
completion() registers pricing under the model name, but _select_model_name_for_cost_calc was
selecting the router deployment hash (which has no pricing data), causing response_cost to be 0.0.
Now checks whether the router_model_id entry actually has pricing before preferring it.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
---------
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Ishaan Jaff <ishaan-jaff@users.noreply.github.com>
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-17 06:32:20 +08:00
|
|
|
"""Test that proxy exits with code 1 when PrismaManager.setup_database returns False and --enforce_prisma_migration_check is set"""
|
2026-03-10 20:45:11 +08:00
|
|
|
from litellm.proxy.proxy_cli import run_server
|
|
|
|
|
|
|
|
|
|
mock_subprocess_run.return_value = MagicMock(returncode=0)
|
|
|
|
|
mock_should_update_schema.return_value = True
|
|
|
|
|
mock_setup_database.return_value = False
|
|
|
|
|
|
|
|
|
|
mock_proxy_module = MagicMock(
|
|
|
|
|
app=MagicMock(),
|
|
|
|
|
ProxyConfig=MagicMock(),
|
|
|
|
|
KeyManagementSettings=MagicMock(),
|
|
|
|
|
save_worker_config=MagicMock(),
|
|
|
|
|
)
|
|
|
|
|
|
|
|
|
|
clean_env = {
|
|
|
|
|
k: v
|
|
|
|
|
for k, v in os.environ.items()
|
|
|
|
|
if k not in ("DATABASE_URL", "DIRECT_URL")
|
|
|
|
|
}
|
|
|
|
|
clean_env["DATABASE_URL"] = "postgresql://test:test@localhost:5432/test"
|
|
|
|
|
|
2026-04-18 04:02:59 +08:00
|
|
|
with (
|
|
|
|
|
patch.dict(os.environ, clean_env, clear=True),
|
|
|
|
|
patch.dict(
|
|
|
|
|
"sys.modules",
|
|
|
|
|
{
|
|
|
|
|
"proxy_server": mock_proxy_module,
|
|
|
|
|
"litellm.proxy.proxy_server": mock_proxy_module,
|
|
|
|
|
},
|
|
|
|
|
),
|
|
|
|
|
patch(
|
|
|
|
|
"litellm.proxy.proxy_cli.ProxyInitializationHelpers._get_default_unvicorn_init_args"
|
|
|
|
|
) as mock_get_args,
|
|
|
|
|
):
|
2026-03-10 20:45:11 +08:00
|
|
|
mock_get_args.return_value = {
|
|
|
|
|
"app": "litellm.proxy.proxy_server:app",
|
|
|
|
|
"host": "localhost",
|
|
|
|
|
"port": 8000,
|
|
|
|
|
}
|
|
|
|
|
|
|
|
|
|
with pytest.raises(SystemExit) as exc_info:
|
|
|
|
|
run_server.main(
|
2026-04-18 04:02:59 +08:00
|
|
|
[
|
|
|
|
|
"--local",
|
|
|
|
|
"--skip_server_startup",
|
|
|
|
|
"--enforce_prisma_migration_check",
|
|
|
|
|
],
|
|
|
|
|
standalone_mode=False,
|
2026-03-10 20:45:11 +08:00
|
|
|
)
|
|
|
|
|
assert exc_info.value.code == 1
|
2026-04-22 05:45:29 +08:00
|
|
|
mock_setup_database.assert_called_once_with(
|
|
|
|
|
use_migrate=True, use_v2_resolver=False
|
|
|
|
|
)
|
2026-03-10 20:45:11 +08:00
|
|
|
|
2026-03-07 10:09:57 +08:00
|
|
|
|
|
|
|
|
# --- Module-level helpers for worker startup hook tests ---
|
|
|
|
|
|
|
|
|
|
_dummy_hook_called = False
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _dummy_hook():
|
|
|
|
|
"""A simple sync hook used by test_should_run_worker_startup_hooks."""
|
|
|
|
|
global _dummy_hook_called
|
|
|
|
|
_dummy_hook_called = True
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
_dummy_async_hook_called = False
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
async def _dummy_async_hook():
|
|
|
|
|
"""A simple async hook used by test_should_run_async_worker_startup_hook."""
|
|
|
|
|
global _dummy_async_hook_called
|
|
|
|
|
_dummy_async_hook_called = True
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
def _failing_hook():
|
|
|
|
|
"""A hook that always raises, used by test_should_raise_on_failing_hook."""
|
|
|
|
|
raise RuntimeError("Hook failed on purpose")
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
class TestWorkerStartupHooks:
|
|
|
|
|
"""Tests for the LITELLM_WORKER_STARTUP_HOOKS mechanism in proxy_startup_event."""
|
|
|
|
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
|
|
|
async def test_should_run_worker_startup_hooks(self):
|
|
|
|
|
"""Sync worker startup hook is called during proxy_startup_event."""
|
|
|
|
|
global _dummy_hook_called
|
|
|
|
|
_dummy_hook_called = False
|
|
|
|
|
|
|
|
|
|
from litellm.proxy.proxy_server import proxy_startup_event
|
|
|
|
|
|
|
|
|
|
env_overrides = {
|
|
|
|
|
"LITELLM_WORKER_STARTUP_HOOKS": "tests.test_litellm.proxy.test_proxy_cli:_dummy_hook",
|
|
|
|
|
}
|
|
|
|
|
# Remove DATABASE_URL to avoid real DB setup
|
|
|
|
|
clean_env = {
|
2026-04-18 04:02:59 +08:00
|
|
|
k: v
|
|
|
|
|
for k, v in os.environ.items()
|
|
|
|
|
if k not in ("DATABASE_URL", "DIRECT_URL")
|
2026-03-07 10:09:57 +08:00
|
|
|
}
|
|
|
|
|
clean_env.update(env_overrides)
|
|
|
|
|
|
|
|
|
|
with patch.dict(os.environ, clean_env, clear=True):
|
|
|
|
|
try:
|
|
|
|
|
async with proxy_startup_event(app=None) as _:
|
|
|
|
|
pass
|
|
|
|
|
except Exception:
|
|
|
|
|
pass # We expect errors after the hook (no DB, etc.)
|
|
|
|
|
|
|
|
|
|
assert _dummy_hook_called is True, "Sync startup hook was not called"
|
|
|
|
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
|
|
|
async def test_should_run_async_worker_startup_hook(self):
|
|
|
|
|
"""Async worker startup hook is awaited during proxy_startup_event."""
|
|
|
|
|
global _dummy_async_hook_called
|
|
|
|
|
_dummy_async_hook_called = False
|
|
|
|
|
|
|
|
|
|
from litellm.proxy.proxy_server import proxy_startup_event
|
|
|
|
|
|
|
|
|
|
env_overrides = {
|
|
|
|
|
"LITELLM_WORKER_STARTUP_HOOKS": "tests.test_litellm.proxy.test_proxy_cli:_dummy_async_hook",
|
|
|
|
|
}
|
|
|
|
|
clean_env = {
|
2026-04-18 04:02:59 +08:00
|
|
|
k: v
|
|
|
|
|
for k, v in os.environ.items()
|
|
|
|
|
if k not in ("DATABASE_URL", "DIRECT_URL")
|
2026-03-07 10:09:57 +08:00
|
|
|
}
|
|
|
|
|
clean_env.update(env_overrides)
|
|
|
|
|
|
|
|
|
|
with patch.dict(os.environ, clean_env, clear=True):
|
|
|
|
|
try:
|
|
|
|
|
async with proxy_startup_event(app=None) as _:
|
|
|
|
|
pass
|
|
|
|
|
except Exception:
|
|
|
|
|
pass
|
|
|
|
|
|
|
|
|
|
assert _dummy_async_hook_called is True, "Async startup hook was not called"
|
|
|
|
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
|
|
|
async def test_should_raise_on_failing_worker_startup_hook(self):
|
|
|
|
|
"""A failing worker startup hook propagates the error."""
|
|
|
|
|
from litellm.proxy.proxy_server import proxy_startup_event
|
|
|
|
|
|
|
|
|
|
env_overrides = {
|
|
|
|
|
"LITELLM_WORKER_STARTUP_HOOKS": "tests.test_litellm.proxy.test_proxy_cli:_failing_hook",
|
|
|
|
|
}
|
|
|
|
|
clean_env = {
|
2026-04-18 04:02:59 +08:00
|
|
|
k: v
|
|
|
|
|
for k, v in os.environ.items()
|
|
|
|
|
if k not in ("DATABASE_URL", "DIRECT_URL")
|
2026-03-07 10:09:57 +08:00
|
|
|
}
|
|
|
|
|
clean_env.update(env_overrides)
|
|
|
|
|
|
|
|
|
|
with patch.dict(os.environ, clean_env, clear=True):
|
|
|
|
|
with pytest.raises(RuntimeError, match="Hook failed on purpose"):
|
|
|
|
|
async with proxy_startup_event(app=None) as _:
|
|
|
|
|
pass
|
|
|
|
|
|
|
|
|
|
def test_should_skip_when_no_hooks_set(self):
|
|
|
|
|
"""When LITELLM_WORKER_STARTUP_HOOKS is not set, no hooks are executed."""
|
|
|
|
|
global _dummy_hook_called
|
|
|
|
|
_dummy_hook_called = False
|
|
|
|
|
|
|
|
|
|
with patch.dict(os.environ, {}, clear=False):
|
|
|
|
|
os.environ.pop("LITELLM_WORKER_STARTUP_HOOKS", None)
|
|
|
|
|
# The hook block should be skipped entirely when env var is absent
|
|
|
|
|
assert "LITELLM_WORKER_STARTUP_HOOKS" not in os.environ
|
|
|
|
|
# Verify that an empty env var value also results in no hook execution
|
|
|
|
|
assert os.environ.get("LITELLM_WORKER_STARTUP_HOOKS", "") == ""
|
|
|
|
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
|
|
|
async def test_should_run_multiple_hooks(self):
|
|
|
|
|
"""Multiple comma-separated hooks are all called."""
|
|
|
|
|
global _dummy_hook_called, _dummy_async_hook_called
|
|
|
|
|
_dummy_hook_called = False
|
|
|
|
|
_dummy_async_hook_called = False
|
|
|
|
|
|
|
|
|
|
from litellm.proxy.proxy_server import proxy_startup_event
|
|
|
|
|
|
|
|
|
|
hooks = (
|
|
|
|
|
"tests.test_litellm.proxy.test_proxy_cli:_dummy_hook,"
|
|
|
|
|
"tests.test_litellm.proxy.test_proxy_cli:_dummy_async_hook"
|
|
|
|
|
)
|
|
|
|
|
env_overrides = {
|
|
|
|
|
"LITELLM_WORKER_STARTUP_HOOKS": hooks,
|
|
|
|
|
}
|
|
|
|
|
clean_env = {
|
2026-04-18 04:02:59 +08:00
|
|
|
k: v
|
|
|
|
|
for k, v in os.environ.items()
|
|
|
|
|
if k not in ("DATABASE_URL", "DIRECT_URL")
|
2026-03-07 10:09:57 +08:00
|
|
|
}
|
|
|
|
|
clean_env.update(env_overrides)
|
|
|
|
|
|
|
|
|
|
with patch.dict(os.environ, clean_env, clear=True):
|
|
|
|
|
try:
|
|
|
|
|
async with proxy_startup_event(app=None) as _:
|
|
|
|
|
pass
|
|
|
|
|
except Exception:
|
|
|
|
|
pass
|
|
|
|
|
|
|
|
|
|
assert _dummy_hook_called is True, "First hook was not called"
|
|
|
|
|
assert _dummy_async_hook_called is True, "Second hook was not called"
|