* fix(vertex_ai): support pluggable (executable) credential_source for WIF auth (#24700) The WIF credential dispatch in load_auth() only handled identity_pool and aws credential types. When credential_source.executable was present (used for Azure Managed Identity via Workload Identity Federation), it fell through to identity_pool.Credentials which rejected it with MalformedError. Add dispatch to google.auth.pluggable.Credentials for executable-type credential sources, following the same pattern as the existing identity_pool and aws helpers. Fixes authentication for Azure Container Apps → GCP Vertex AI via WIF with executable credential sources. * feat(logging): add component and logger fields to JSON logs for 3rd p… (#24447) * feat(logging): add component and logger fields to JSON logs for 3rd party filtering * Let user-supplied extra fields win over auto-generated component/logger, tighten test assertions * Feat - Add organization into the metrics metadata for org_id & org_alias (#24440) * Add org_id and org_alias label names to Prometheus metric definitions * Add user_api_key_org_alias to StandardLoggingUserAPIKeyMetadata * Populate user_api_key_org_alias in pre-call metadata * Pass org_id and org_alias into per-request Prometheus metric labels * Add test for org labels on per-request Prometheus metrics * chore: resolve test mockdata * Address review: populate org_alias from DB view, add feature flag, use .get() for org metadata * Add org labels to failure path and verify flag behavior in test * Fix test: build flag-off enum_values without org fields * Gate org labels behind feature flag in get_labels() instead of static metric lists * Scope org label injection to metrics that carry team context, remove orphaned budget label defs, add test teardown * Use explicit metric allowlist for org label injection instead of team heuristic * Fix duplicate org label guard, move _org_label_metrics to class constant * Reset custom_prometheus_metadata_labels after duplicate label assertion * fix: emit org labels by default, remove flag, fix missing org_alias in all metadata paths * fix: emit org labels by default, no opt-in flag required * fix: write org_alias to metadata unconditionally in proxy_server.py * fix: 429s from batch creation being converted to 500 (#24703) * add us gov models (#24660) * add us gov models * added max tokens * Litellm dev 04 02 2026 p1 (#25052) * fix: replace hardcoded url * fix: Anthropic web search cost not tracked for Chat Completions The ModelResponse branch in response_object_includes_web_search_call() only checked url_citation annotations and prompt_tokens_details, missing Anthropic's server_tool_use.web_search_requests field. This caused _handle_web_search_cost() to never fire for Anthropic Claude models. Also routes vertex_ai/claude-* models to the Anthropic cost calculator instead of the Gemini one, since Claude on Vertex uses the same server_tool_use billing structure as the direct Anthropic API. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> * fix(anthropic): pass logging_obj to client.post for litellm_overhead_time_ms (#24071) When LITELLM_DETAILED_TIMING=true, litellm_overhead_time_ms was null for Anthropic because the handler did not pass logging_obj to client.post(), so track_llm_api_timing could not set llm_api_duration_ms. Pass logging_obj=logging_obj at all four post() call sites (make_call, make_sync_call, acompletion, completion). Add test to ensure make_call passes logging_obj to client.post. Made-with: Cursor * sap - add additional parameters for grounding - additional parameter for grounding added for the sap provider * sap - fix models * (sap) add filtering, masking, translation SAP GEN AI Hub modules * (sap) add tests and docs for new SAP modules * (sap) add support of multiple modules config * (sap) code refactoring * (sap) rename file * test(): add safeguard tests * (sap) update tests * (sap) update docs, solve merge conflict in transformation.py * (sap) linter fix * (sap) Align embedding request transformation with current API * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) mock commit * (sap) run black formater * (sap) add literals to models, add negative tests, fix test for tool transformation * (sap) fix formating * (sap) fix models * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) commit for rerun bot review * (sap) minor improve * (sap) fix after bot review * (sap) lint fix * docs(sap): update documentation * fix(sap): change creds priority * fix(sap): change creds priority * fix(sap): fix sap creds unit test * fix(sap): linter fix * fix(sap): linter fix * linter fix * (sap) update logic of fetching creds, add additional tests * (sap) clean up code * (sap) fix after review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) add a possibility to put the service key by both variants * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) update test * (sap) update service key resolve function * (sap) run black formater * (sap) fix validate credentials, add negative tests for credential fetching * (sap) fix validate credentials, add negative tests for credential fetching * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) fix after bot review * (sap) lint fix * (sap) lint fix * feat: support service_tier in gemini * chore: add a service_tier field mapping from openai to gemini * fix: use x-gemini-service-tier header in response * docs: add service_tier to gemini docs * chore: add defaut/standard mapping, and some tests * chore: tidying up some case insensitivity * chore: remove unnecessary guard * fix: remove redundant test file * fix: handle 'auto' case-insensitively * fix: return service_tier on final steamed chunk * chore: black * feat: enable supports_service_tier to gemini models * Fix get_standard_logging_metadata tests * Fix test_get_model_info_bedrock_models * Fix test_get_model_info_bedrock_models * Fix remaining tests * Fix mypy issues * Fix tests * Fix merge conflicts * Fix code qa * Fix code qa * Fix code qa * Fix greptile review --------- Co-authored-by: michelligabriele <gabriele.michelli@icloud.com> Co-authored-by: Josh <36064836+J-Byron@users.noreply.github.com> Co-authored-by: mubashir1osmani <mubashir.osmani777@gmail.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: milan-berri <milan@berri.ai> Co-authored-by: Alperen Kömürcü <alperen.koemuercue@sap.com> Co-authored-by: Vasilisa Parshikova <vasilisa.parshikova@sap.com> Co-authored-by: Lin Xu <lin.xu03@sap.com> Co-authored-by: Mark McDonald <macd@google.com> Co-authored-by: Sameer Kankute <sameer@berri.ai>
196 lines
7.5 KiB
Python
196 lines
7.5 KiB
Python
import importlib
|
|
import os
|
|
from typing import TYPE_CHECKING, Dict, Optional, Type
|
|
|
|
from litellm._logging import verbose_logger
|
|
from litellm.types.utils import CallTypes
|
|
|
|
from . import *
|
|
|
|
if TYPE_CHECKING:
|
|
from litellm.llms.base_llm.guardrail_translation.base_translation import (
|
|
BaseTranslation,
|
|
)
|
|
from litellm.types.utils import ModelInfo, Usage
|
|
|
|
|
|
def get_cost_for_web_search_request(
|
|
custom_llm_provider: str, usage: "Usage", model_info: "ModelInfo"
|
|
) -> Optional[float]:
|
|
"""
|
|
Get the cost for a web search request for a given model.
|
|
|
|
Args:
|
|
custom_llm_provider: The custom LLM provider.
|
|
usage: The usage object.
|
|
model_info: The model info.
|
|
"""
|
|
if custom_llm_provider == "gemini":
|
|
from .gemini.cost_calculator import cost_per_web_search_request
|
|
|
|
return cost_per_web_search_request(usage=usage, model_info=model_info)
|
|
elif custom_llm_provider == "anthropic":
|
|
from .anthropic.cost_calculation import get_cost_for_anthropic_web_search
|
|
|
|
return get_cost_for_anthropic_web_search(model_info=model_info, usage=usage)
|
|
elif custom_llm_provider.startswith("vertex_ai"):
|
|
# Anthropic Claude models on Vertex AI populate server_tool_use.web_search_requests
|
|
# (same as the direct Anthropic API), not prompt_tokens_details.web_search_requests
|
|
# (which is the Gemini field). Route claude-* models to the Anthropic calculator.
|
|
model_key: str = model_info.get("key", "") if model_info else ""
|
|
if "claude" in model_key.lower():
|
|
from .anthropic.cost_calculation import get_cost_for_anthropic_web_search
|
|
|
|
verbose_logger.debug(
|
|
"vertex_ai/claude model detected — routing web search cost to Anthropic calculator"
|
|
)
|
|
return get_cost_for_anthropic_web_search(model_info=model_info, usage=usage)
|
|
|
|
from .vertex_ai.gemini.cost_calculator import (
|
|
cost_per_web_search_request as cost_per_web_search_request_vertex_ai,
|
|
)
|
|
|
|
return cost_per_web_search_request_vertex_ai(usage=usage, model_info=model_info)
|
|
elif custom_llm_provider == "perplexity":
|
|
# Perplexity handles search costs internally in its own cost calculator
|
|
# Return 0.0 to indicate costs are already accounted for
|
|
return 0.0
|
|
elif custom_llm_provider == "xai":
|
|
from .xai.cost_calculator import cost_per_web_search_request
|
|
|
|
return cost_per_web_search_request(usage=usage, model_info=model_info)
|
|
else:
|
|
return None
|
|
|
|
|
|
def discover_guardrail_translation_mappings() -> (
|
|
Dict[CallTypes, Type["BaseTranslation"]]
|
|
):
|
|
"""
|
|
Discover guardrail translation mappings by scanning the llms directory structure.
|
|
|
|
Scans for modules with guardrail_translation_mappings dictionaries and aggregates them.
|
|
|
|
Returns:
|
|
Dict[CallTypes, Type[BaseTranslation]]: A dictionary mapping call types to their translation handler classes
|
|
"""
|
|
discovered_mappings: Dict[CallTypes, Type["BaseTranslation"]] = {}
|
|
|
|
try:
|
|
# Get the path to the llms directory
|
|
current_dir = os.path.dirname(__file__)
|
|
llms_dir = current_dir
|
|
|
|
if not os.path.exists(llms_dir):
|
|
verbose_logger.debug("llms directory not found")
|
|
return discovered_mappings
|
|
|
|
# Recursively scan for guardrail_translation directories
|
|
for root, dirs, files in os.walk(llms_dir):
|
|
# Skip __pycache__ and base_llm directories
|
|
dirs[:] = [d for d in dirs if not d.startswith("__") and d != "base_llm"]
|
|
|
|
# Check if this is a guardrail_translation directory with __init__.py
|
|
if (
|
|
os.path.basename(root) == "guardrail_translation"
|
|
and "__init__.py" in files
|
|
):
|
|
# Build the module path relative to litellm
|
|
rel_path = os.path.relpath(root, os.path.dirname(llms_dir))
|
|
module_path = "litellm." + rel_path.replace(os.sep, ".")
|
|
|
|
try:
|
|
# Import the module
|
|
verbose_logger.debug(
|
|
f"Discovering guardrail translations in: {module_path}"
|
|
)
|
|
|
|
module = importlib.import_module(module_path)
|
|
|
|
# Check for guardrail_translation_mappings dictionary
|
|
if hasattr(module, "guardrail_translation_mappings"):
|
|
mappings = getattr(module, "guardrail_translation_mappings")
|
|
if isinstance(mappings, dict):
|
|
discovered_mappings.update(mappings)
|
|
verbose_logger.debug(
|
|
f"Found guardrail_translation_mappings in {module_path}: {list(mappings.keys())}"
|
|
)
|
|
|
|
except ImportError as e:
|
|
verbose_logger.error(f"Could not import {module_path}: {e}")
|
|
continue
|
|
except Exception as e:
|
|
verbose_logger.error(f"Error processing {module_path}: {e}")
|
|
continue
|
|
|
|
try:
|
|
from litellm.proxy._experimental.mcp_server.guardrail_translation import (
|
|
guardrail_translation_mappings as mcp_guardrail_translation_mappings,
|
|
)
|
|
|
|
discovered_mappings.update(mcp_guardrail_translation_mappings)
|
|
verbose_logger.debug(
|
|
"Loaded MCP guardrail translation mappings: %s",
|
|
list(mcp_guardrail_translation_mappings.keys()),
|
|
)
|
|
except ImportError:
|
|
verbose_logger.debug(
|
|
"MCP guardrail translation mappings not available; skipping"
|
|
)
|
|
|
|
verbose_logger.debug(
|
|
f"Discovered {len(discovered_mappings)} guardrail translation mappings: {list(discovered_mappings.keys())}"
|
|
)
|
|
|
|
except Exception as e:
|
|
verbose_logger.error(f"Error discovering guardrail translation mappings: {e}")
|
|
|
|
return discovered_mappings
|
|
|
|
|
|
# Cache the discovered mappings
|
|
endpoint_guardrail_translation_mappings: Optional[
|
|
Dict[CallTypes, Type["BaseTranslation"]]
|
|
] = None
|
|
|
|
|
|
def load_guardrail_translation_mappings():
|
|
global endpoint_guardrail_translation_mappings
|
|
if endpoint_guardrail_translation_mappings is None:
|
|
endpoint_guardrail_translation_mappings = (
|
|
discover_guardrail_translation_mappings()
|
|
)
|
|
return endpoint_guardrail_translation_mappings
|
|
|
|
|
|
def get_guardrail_translation_mapping(call_type: CallTypes) -> Type["BaseTranslation"]:
|
|
"""
|
|
Get the guardrail translation handler for a given call type.
|
|
|
|
Args:
|
|
call_type: The type of call (e.g., completion, acompletion, anthropic_messages)
|
|
|
|
Returns:
|
|
The translation handler class for the given call type
|
|
|
|
Raises:
|
|
ValueError: If no translation mapping exists for the given call type
|
|
"""
|
|
global endpoint_guardrail_translation_mappings
|
|
|
|
# Lazy load the mappings on first access
|
|
if endpoint_guardrail_translation_mappings is None:
|
|
endpoint_guardrail_translation_mappings = (
|
|
discover_guardrail_translation_mappings()
|
|
)
|
|
|
|
# Get the translation handler class for the call type
|
|
if call_type not in endpoint_guardrail_translation_mappings:
|
|
raise ValueError(
|
|
f"No guardrail translation mapping found for call_type: {call_type}. "
|
|
f"Available mappings: {list(endpoint_guardrail_translation_mappings.keys())}"
|
|
)
|
|
|
|
# Return the handler class directly
|
|
return endpoint_guardrail_translation_mappings[call_type]
|