Commit Graph

23429 Commits

Author SHA1 Message Date
Ishaan Jaff
93badc72dc ui fix formatNumberWithCommas 2025-07-19 14:42:41 -07:00
Ishaan Jaff
76d461dcae fix add model 2025-07-19 14:32:41 -07:00
Jugal D. Bhatt
b443817a56
[Key Access] Litellm disabled callbacks for UI (#12769)
* add disabled callbacks to ui

* added body

* update edit settings

* add tests
2025-07-19 14:32:05 -07:00
Krish Dholakia
ee066481f8
UI - Support 'batch' model health checks + make 'team-only' model concept clearer (#12770)
* fix(add_model_modes.tsx): add 'batch' mode to ui

* fix(main.py): support health checks on batches + support litellm_credentials on batches

* fix(add_model_tab.tsx): clarify what 'team' on add model means
2025-07-19 14:30:38 -07:00
Jugal D. Bhatt
92c9e38eca
[JSON Logs] fix ciruclar ref error by adding safe dumps (#12764)
* fix ciruclar ref error by adding safe dumps

* fix ruff

* fix ruff

* Update spend_tracking_utils.py
2025-07-19 13:45:27 -07:00
Cole McIntosh
ceb4a143c9
fix: correct Groq model naming convention for moonshotai/kimi-k2-instruct (#12768)
- Changed groq/moonshotai-kimi-k2-instruct to groq/moonshotai/kimi-k2-instruct in model_prices_and_context_window.json
- Added groq/moonshotai/kimi-k2-instruct and groq/qwen-qwq-32b to the supported models table in Groq documentation
2025-07-19 13:36:40 -07:00
Ishaan Jaff
05af269425 Revert "ui fix linting"
This reverts commit 85184c7f82.
2025-07-19 12:42:45 -07:00
Ishaan Jaff
5c7e5d4324 Revert "Regenerate Key State Management and Authentication Issues (#12729)"
This reverts commit 663abbe275.
2025-07-19 12:42:20 -07:00
Krish Dholakia
e03bc3ec7e
feat(proxy_server.py): add model hub to the swagger (#12767)
user request
2025-07-19 12:33:57 -07:00
Ishaan Jaff
7e2546da2d docs vllm rerank 2025-07-19 12:11:50 -07:00
Ishaan Jaff
a305c4a54c docs vLLM Rerank 2025-07-19 12:11:19 -07:00
Ishaan Jaff
3ca3772ef0 docs Vector Stores 2025-07-19 11:55:58 -07:00
Ishaan Jaff
4b13e3e214
[Docs] 1.74.6.rc note (#12765)
* draft 1.74.6

* add correct models

* fix

* update moonshot pricing

* docs

* docs fix

* changes till HELM

* Helm Chart

* upto circular references

* docs Groq

* fix typo

* docs

* docs fix

* docs fix
2025-07-19 11:54:22 -07:00
Krish Dholakia
ab09d0621d
Litellm gemini grounding metadata stream (#12673)
* fix(prompt_templates/factory.py): handle anthropic cache control on individual tool results

Fixes issue where cache control on individual tool result was being ignored

* test(test_vertex_And_google_ai_studio_gemini.py): initial unit test covering translation for grounding metadata on streaming chunk

* fix(vertex_and_google_ai_studio.py): ensure grounding metadata is preserved on streaming

Closes https://github.com/BerriAI/litellm/issues/10237

* fix(core_helpers.py): include usage in expected openai keys
2025-07-19 11:52:12 -07:00
Krrish Dholakia
6d0e575f74 docs(docusaurus.config.js): route to new support onboarding form
gives user both slack + discord invites
2025-07-19 11:43:36 -07:00
Krish Dholakia
d72b3389a1
Bulk Edit Users on UI (#12763)
* feat(bulk_edit_user.tsx): initial working ui for editing users in bulk on the ui

easier to give access / assign to a default team

* feat(team_endpoints-+-bulk_edit_users.tsx): add bulk adding users to teams

make it easier to add existing users to a default team

* fix(bulk_edit_user.tsx): fix ui linting error

* fix: fix linting error
2025-07-19 11:04:23 -07:00
Ishaan Jaff
96f7eb6f78 bump litellm enterprise version 2025-07-19 10:12:33 -07:00
Ishaan Jaff
06574a72b5
[Feat] Backend - Add support for disabling callbacks in request body (#12762)
* allow using standard_callback_dynamic_params to disable callbacks

* fix is_callback_disabled_dynamically

* test_callback_disabled_via_request_body_multiple
2025-07-19 10:10:30 -07:00
tanjiro
657ca3b81a
Fix Y-axis labels overlap on Spend per Tag (#12754)
* fix spend per tag bar chart labels

* remove console.log
2025-07-19 09:40:40 -07:00
tanjiro
f8e5dc5034
copy button (#12760) 2025-07-19 09:39:14 -07:00
Ishaan Jaff
4915d15cca bump: version 1.74.5 → 1.74.6 2025-07-18 18:48:39 -07:00
Ishaan Jaff
d0af0db766 ui new build 2025-07-18 18:47:53 -07:00
Ishaan Jaff
c89fad06e0 fix vtx linting 2025-07-18 18:45:13 -07:00
Ishaan Jaff
85184c7f82 ui fix linting 2025-07-18 18:42:24 -07:00
Ishaan Jaff
7eca3efc92 Revert "ui fix linting"
This reverts commit 586918dddb.
2025-07-18 18:38:53 -07:00
Ishaan Jaff
02f987d63d Revert "fix linting ui"
This reverts commit 1ff766e738.
2025-07-18 18:38:44 -07:00
Ishaan Jaff
2880fc9837 add list of v0 models by provider 2025-07-18 18:37:48 -07:00
Ishaan Jaff
1ff766e738 fix linting ui 2025-07-18 18:36:44 -07:00
Ishaan Jaff
586918dddb ui fix linting 2025-07-18 18:35:32 -07:00
Ishaan Jaff
bf50544b97 fix code qa 2025-07-18 18:34:51 -07:00
Cole McIntosh
bf046c9d5d
feat: add v0 provider support (#12751)
* feat: add v0 provider support to LiteLLM

- Add v0 as a new OpenAI-compatible provider
- Support all three v0 models: v0-1.0-md, v0-1.5-md, v0-1.5-lg
- Configure correct token limits and pricing for each model
- Enable vision support for all v0 models (multimodal)
- Add provider detection for v0/ prefix and api.v0.dev endpoint
- Include comprehensive unit tests for the provider

The v0 provider uses the standard OpenAI-compatible implementation
and supports all standard features including streaming, function
calling, and system messages.

* fix: add v0 provider to ProviderConfigManager

Add V0ChatConfig to the get_provider_chat_config method to fix
test_supports_tool_choice test failure. The v0 provider needs to
be included in the provider config manager to return the correct
configuration for tool choice support detection.

* docs: add documentation for v0 provider

- Add comprehensive v0 provider documentation
- Cover all supported models and their capabilities
- Include examples for SDK usage, proxy configuration, and all features
- Document supported OpenAI parameters based on v0 API docs
- Add v0 to the providers sidebar navigation

* fix: correct v0 supported OpenAI parameters

Based on review feedback and v0 API documentation:
- v0 only supports: messages, model, stream, tools, tool_choice
- Remove unsupported parameters like temperature, max_tokens, etc.
- Update tests to verify correct parameter set
- Update documentation to reflect actual API capabilities
- Remove JSON mode example as response_format is not supported

Reference: https://v0.dev/docs/v0-model-api#request-body

* fix: remove supports_response_schema from v0 models

Remove the supports_response_schema property from all v0 models in the model configuration files as v0 does not support this feature.

Models updated:
- v0/v0-1.0-md
- v0/v0-1.5-md
- v0/v0-1.5-lg
2025-07-18 18:26:44 -07:00
Ishaan Jaff
81eb2fdd30
[Feat] UI Vector Stores - Allow adding Vertex RAG Engine, OpenAI, Azure (#12752)
* fix _pass_through_endpoint_without_required_model

* add get_litellm_managed_vector_store_from_registry

* undo router change

* fix for using router + vector search methods

* add simple helper for _update_request_data_with_litellm_managed_vector_store_registry

* add vector_stores routes

* test_router_avector_store_search_passes_correct_args

* [Feat] UI - Allow clicking into Vector Stores (#12741)

* Add View Vector Store

* add /info for vector store

* fix updated_at

* allow easily testing the KB on litellm

* fix

* rename test

* test_init_vector_store_api_endpoints

* add get_vertex_ai_project

* fixes to vertex transformation for RAG Engine

* fix vectorStoreProviderFields

* Add Vertex Rag engine

* add oai, azure

* fix validate_environment

* fix provider name

* fix tester

* working vertex vector store
2025-07-18 18:25:26 -07:00
Ishaan Jaff
5802a5bbe3
[Feat] LLM API Endpoint - Expose OpenAI Compatible /vector_stores/{vector_store_id}/search endpoint (#12749)
* fix _pass_through_endpoint_without_required_model

* add get_litellm_managed_vector_store_from_registry

* undo router change

* fix for using router + vector search methods

* add simple helper for _update_request_data_with_litellm_managed_vector_store_registry

* add vector_stores routes

* test_router_avector_store_search_passes_correct_args

* [Feat] UI - Allow clicking into Vector Stores (#12741)

* Add View Vector Store

* add /info for vector store

* fix updated_at

* allow easily testing the KB on litellm

* fix

* rename test

* test_init_vector_store_api_endpoints

* test_update_request_data_with_litellm_managed_vector_store_registry
2025-07-18 18:18:53 -07:00
Jugal D. Bhatt
8d35a00974
[LLM Translation] Added model name formats (#12745)
* Added model supports

* invert logic

* Added gpt 35 turbo check

* add test check

* fix ruff check
2025-07-18 17:08:35 -07:00
Jugal D. Bhatt
be60d12ff7
[LLM Translation - Redis] fix: redis caching for embedding response models (#12750)
* fix: redis caching for embedding responses

* add helper

* add mypy fixes

* lint fix

* review changes

* remove file

* fix ruff

* add if check

* add if check
2025-07-18 16:31:10 -07:00
Jugal D. Bhatt
c3c6255689
[LLM Translation] Change System prompts to assistant prompts as a workaround for GH Copilot (#12742)
* add changes for copilot

* Add test

* reverse flag settings

* add settings

* utils changes

* fix tests
2025-07-18 15:48:27 -07:00
dependabot[bot]
a1c06e9f23
build(deps): bump on-headers and compression in /docs/my-website (#12721)
Bumps [on-headers](https://github.com/jshttp/on-headers) and [compression](https://github.com/expressjs/compression). These dependencies needed to be updated together.

Updates `on-headers` from 1.0.2 to 1.1.0
- [Release notes](https://github.com/jshttp/on-headers/releases)
- [Changelog](https://github.com/jshttp/on-headers/blob/master/HISTORY.md)
- [Commits](https://github.com/jshttp/on-headers/compare/v1.0.2...v1.1.0)

Updates `compression` from 1.8.0 to 1.8.1
- [Release notes](https://github.com/expressjs/compression/releases)
- [Changelog](https://github.com/expressjs/compression/blob/master/HISTORY.md)
- [Commits](https://github.com/expressjs/compression/compare/1.8.0...v1.8.1)

---
updated-dependencies:
- dependency-name: on-headers
  dependency-version: 1.1.0
  dependency-type: indirect
- dependency-name: compression
  dependency-version: 1.8.1
  dependency-type: indirect
...

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
2025-07-18 15:20:01 -07:00
Cole McIntosh
d4f5180212
fix(lowest_latency.py): Handle ZeroDivisionError with zero completion tokens (#12734)
* fix(lowest_latency.py): Handle ZeroDivisionError with zero completion tokens (#12641)

Fixes ZeroDivisionError when LLM responses have zero completion tokens, which can
occur with Gemini models on very long contexts that only use tool calls.

Changes:
- Add check for completion_tokens > 0 before division in log_success_event
- Handle both timedelta and float types for response times (supporting both time.time() and datetime)
- Apply fix to both occurrences of the division operation in the file
- Add comprehensive tests for zero completion token scenarios

This ensures the lowest latency routing continues to work properly even when
models return responses with no completion tokens.

* test: Move lowest_latency zero tokens test to test_litellm for CI execution

* refactor: Use safe_divide helper to eliminate code duplication

- Added safe_divide utility function to litellm_core_utils.core_helpers
- Handles both timedelta and float types for numerator
- Prevents ZeroDivisionError and negative denominator issues
- Replaced duplicated division logic in lowest_latency.py
- Added comprehensive unit tests for safe_divide function

This improves code quality by reducing duplication and centralizing the
division safety logic in a reusable helper function.

* refactor: Rename to safe_divide_seconds for clarity

- Renamed safe_divide to safe_divide_seconds to better indicate it handles time durations
- Simplified implementation with single-line ternary for seconds conversion
- Updated all references and tests accordingly

The function name now clearly indicates it's specifically for dividing
time durations (in seconds) by a denominator.

* refactor: Simplify safe_divide_seconds to only accept float arguments

- Changed safe_divide_seconds to accept only float parameters for consistency
- Callers now handle timedelta to seconds conversion explicitly
- Removed timedelta-specific test cases

* fix: Remove unused timedelta import

* fix: Remove Union[timedelta, float] type annotations

Since safe_divide_seconds now only accepts floats, we handle the
timedelta conversion explicitly in the code. The type annotations
are no longer needed and can be simplified.

* Revert "fix: Remove Union[timedelta, float] type annotations"

This reverts commit 19c6c62154fceb9c431fea3e448753495dec35d2.

* fix: Clean up implementation

- Remove unnecessary type annotations
- Use safe_divide_seconds utility for zero-division protection
- Handle both timedelta and float types explicitly
- Maintain compatibility with both datetime and time.time() usage
2025-07-18 15:19:47 -07:00
Ishaan Jaff
99e2ea081d
[Feat] UI - Allow clicking into Vector Stores (#12741)
* Add View Vector Store

* add /info for vector store

* fix updated_at
2025-07-18 14:28:57 -07:00
Ryan Richard
474ac2dd6a
add project_id from auth metadata to credentials cache if a user does not specify a project_id (#12661)
add unit tests for new cached credentials with project_id

Co-authored-by: Ryan Richard <ryanirichard07@gmail.com>
2025-07-18 13:33:09 -07:00
Krrish Dholakia
075032349f build: update litellm-proxy-extras 2025-07-18 13:18:59 -07:00
Krrish Dholakia
31ca4da734 build(migration.sql): add new sql migration 2025-07-18 12:57:28 -07:00
Cole McIntosh
506dc80b15
feat: integrate Google Cloud Model Armor guardrails (#12492)
* feat: integrate Google Cloud Model Armor guardrails (LIT-298)

- Add ModelArmorGuardrail class that extends CustomGuardrail and VertexBase
- Support for both pre-call (sanitizeUserPrompt) and post-call (sanitizeModelResponse) sanitization
- Integrate with existing Vertex AI authentication using VertexBase
- Add configuration model for Model Armor in guardrail types
- Register Model Armor in guardrail initializers and registry
- Include comprehensive test suite for Model Armor functionality
- Support content masking for both requests and responses
- Handle streaming responses with content sanitization

This integration allows LiteLLM to use Google Cloud Model Armor API for
content moderation and sanitization, providing similar functionality to
Bedrock Guards but using Google Cloud's security infrastructure.

* fix: remove unused imports flagged by ruff linter

- Remove unused asyncio import at top level (moved to local import where needed)
- Remove unused TextCompletionResponse import

* fix: remove additional unused imports

- Remove unused TextChoices import from line 34
- Remove duplicate asyncio import from line 384
- Replace asyncio.iscoroutine() with hasattr check for __await__

* fix: remove final unused imports

- Remove unused 'import sys' from line 9
- Remove unused 'StreamingChoices' from imports

* fix: remove unused import from model_armor.py

- Remove unused 'import os' from the top of the file

* fix: remove commented-out header from model_armor.py

- Eliminate unnecessary comments at the top of the file to improve code clarity.

* feat(guardrails): Add Model Armor UI support

- Convert model_armor.py to a directory structure with __init__.py for dynamic discovery
- Add get_config_model() method to ModelArmorGuardrail class for UI integration
- Add ui_friendly_name() to ModelArmorConfigModel returning "Google Cloud Model Armor"
- Remove manual registration from guardrail_registry.py to use dynamic discovery
- Model Armor now appears in the UI guardrails dropdown with proper configuration fields

This enables users to configure Model Armor guardrails through the LiteLLM UI interface.

* fix(guardrails): Fix undefined name 'GuardrailConfigModel' in Model Armor

- Import TYPE_CHECKING and GuardrailConfigModel from base module
- Fixes F821 linting error for undefined name in type annotation
- Follows same pattern as other guardrail implementations

* fix(guardrails): Fix Model Armor type errors and config model inheritance

- Create ModelArmorGuardrailConfigModel that properly inherits from GuardrailConfigModel base class
- Move config model to litellm/types/proxy/guardrails/guardrail_hooks/model_armor.py following convention
- Update get_config_model() to return the properly typed config model
- Remove ModelArmorConfigModel from LitellmParams inheritance chain
- Add template_id field to BaseLitellmParams instead

This fixes the mypy type errors and follows the same pattern as other guardrail implementations.

* fix(guardrails): Add missing Model Armor fields to BaseLitellmParams

- Add location, credentials, api_endpoint, and fail_on_error fields
- Fixes mypy errors about missing attributes in LitellmParams
- All Model Armor configuration parameters are now properly defined

* fix: Apply PR review feedback for Model Armor guardrail

- Move test file to tests/test_litellm/ for GitHub Actions
- Extract only last consecutive user messages to avoid context limits
- Use get_content_from_model_response helper for response extraction
- Handle non-ModelResponse types (e.g., TTS) gracefully
- Maintain newline separation for multi-part content

* refactor: Simplify message content extraction in ModelArmorGuardrail

- Removed the custom _extract_content_from_messages method.
- Integrated get_last_user_message helper for improved content extraction.
- Updated test to reflect changes in content formatting.

* refactor: Remove unused import in model_armor.py

- Deleted the unused import of AllMessageValues to clean up the codebase.

* Add unit tests for Model Armor guardrail functionality

- Implement tests for pre-call and post-call hooks, including content sanitization and blocking behavior.
- Validate error handling for API responses and credential management.
- Test streaming responses and handling of list content in user messages.
- Ensure proper assertions for API interactions and response sanitization.

* Add comprehensive test coverage for Model Armor guardrail

- Add test for requests with no messages field
- Add test for empty message content handling
- Add test for system/assistant-only messages
- Add test for fail_on_error=False behavior
- Add test for custom API endpoint configuration
- Add test for dictionary credentials (non-file path)
- Add test for action=NONE response handling
- Add test for missing sanitized_text field fallback
- Add test for non-text response types (TTS/image)
- Add test for auth token refresh behavior

All tests ensure robust edge case handling and proper error management.

* feat: improve Model Armor handling of non-ModelResponse types

- Add debug logging when skipping non-text responses (TTS, images, etc.)
- Improve docstring to clarify behavior for non-text responses
- Add test coverage for non-ModelResponse handling
- Ensure guardrail gracefully skips processing for response types it cannot handle
2025-07-18 12:51:09 -07:00
Jugal D. Bhatt
208484ce65
[jais-30b-chat] added model to prices and context window (#12739)
* added model to prices and context window

* add comma
2025-07-18 11:56:33 -07:00
Jugal D. Bhatt
a46b9d376f
[Prometheus] Move Prometheus to enterprise folder (#12659)
* fix tools fetch for keys

* add promethues to enterprise

* remove old prom

* remove old prom

* fix tests

* safe imports

* add if

* fix enterprise test

* rename imports

* added label import

* added label import

* move tests to enterprise

* fix tests

* add log

* build: update versions

---------

Co-authored-by: Krrish Dholakia <krrishdholakia@gmail.com>
2025-07-18 11:54:47 -07:00
Krish Dholakia
5004b915d1
Guardrails AI - support llmOutput based guardrails as pre-call hooks (#12674)
* build: move build_and_test to use prisma migrate

* fix(guardrails_ai.py): default to guardrail accepting 'llmOutput' as the input param

enables same guardrail to work for pre call and post call

* fix(__init__.py): set default value

* fix(guardrails_ai.py): updates

* fix: fix linting error
2025-07-18 11:31:30 -07:00
Jugal D. Bhatt
a112ec5b02
Health check app on separate port (#12718)
* add separate health app

* add new docs

* refactor

* fix colons

* Update config_settings.md

* refactor

* docs

* add unit test

* added supervisord

* remove app

* add supervisor conf

* Add markdown

* add video to md

* remove test

* docs build failure

* add to all docker files, change prod.md and add tests

* change dockerfiles

* remove extra file

* remove extra file

* remove extra file

* change apt->apk

* remove rdb file

* add fixed file
2025-07-18 11:17:15 -07:00
Krish Dholakia
f6f3f151f1
Anthropic - add tool cache control support (#12668)
* fix(prompt_templates/factory.py): handle anthropic cache control on individual tool results

Fixes issue where cache control on individual tool result was being ignored

* test(test_vertex_And_google_ai_studio_gemini.py): initial unit test covering translation for grounding metadata on streaming chunk
2025-07-18 11:14:03 -07:00
Krish Dholakia
60c7537cc7
/streamGenerateContent - non-gemini model support (#12647)
* fix(google_genai/adapters/transformation.py): enable calling non-googlegenai models via streaming

Fixes https://github.com/BerriAI/litellm/issues/12562

* test(test_openai.py): add unit test asserting streaming works as expected
2025-07-18 10:56:29 -07:00
Jugal D. Bhatt
7c49197f29
Add Hosted VLLM rerank provider integration (#12738)
* Vllm rerank (#12737)

* Add Hosted VLLM rerank provider integration

This commit implements the Hosted VLLM rerank provider integration for LiteLLM. The integration includes:
Adding Hosted VLLM as a supported rerank provider in the main rerank function
Implementing the HostedVLLMRerank handler class for making API requests
Creating a transformation class to convert Hosted VLLM responses to LiteLLM's standardized format
The integration supports both synchronous and asynchronous rerank operations. API credentials can be provided directly or through environment variables (HOSTED_VLLM_API_KEY and HOSTED_VLLM_API_BASE).
Notable features:
Proper error handling for missing credentials
Standard response transformation
Support for common rerank parameters (top_n, return_documents, etc.)
Proper token usage tracking
This expands LiteLLM's rerank provider ecosystem to include Hosted VLLM alongside existing providers like Cohere, Together AI, Azure AI, and Bedrock.

* refactor(rerank): use base_llm_http_handler for hosted_vllm rerank

- Replace custom HostedVLLMRerank handler with base_llm_http_handler
- Implement proper HostedVLLMRerankConfig inheriting from BaseRerankConfig
- Follow Cohere-compatible implementation pattern
- Clean up unnecessary comments

* Fix lint errors in hosted_vllm rerank transformer: remove unused imports

* Fix linting errors in rerank transformation modules

* fix: resolve type errors in Hosted VLLM rerank module

---------

Co-authored-by: Philip D'Souza <philip.dsouza@macro4.com>
Co-authored-by: Philip D'Souza <philip.a.dsouza@gmail.com>

* added a few tests

---------

Co-authored-by: Philip D'Souza <philip.dsouza@macro4.com>
Co-authored-by: Philip D'Souza <philip.a.dsouza@gmail.com>
2025-07-18 10:55:50 -07:00