Debnil Sur
d94af171db
fix(exception_mapping): handle exceptions without response parameter ( #18919 )
...
When extract_and_raise_litellm_exception tries to raise a LiteLLM exception
from an error string, it was always passing the response parameter. However,
some exceptions like APIConnectionError don't accept this parameter, causing
a TypeError.
This fix tries to raise the exception with the response parameter first,
and falls back to raising without it if a TypeError occurs.
This fixes the error:
TypeError: APIConnectionError.__init__() got an unexpected keyword argument 'response'
Which was occurring when Gemini returned UNEXPECTED_TOOL_CALL finish reason
and LiteLLM tried to convert the error to an APIConnectionError.
Fixes: cascading error when Gemini uses thinking feature (__thought__ tool calls)
2026-01-14 04:03:05 +05:30
nulone
478bdcb60b
fix(model_prices): sync DeepSeek chat/reasoner to V3.2 pricing ( #18884 )
2026-01-14 03:52:50 +05:30
Ryan Malloy
f76938af5e
fix(ollama): set finish_reason to tool_calls and remove broken capability check ( #18924 )
...
* Update CLAUDE.md with qwen3 tool_calls bug fix instructions (#18922 )
* fix(ollama): set finish_reason to "tool_calls" when tool_calls present
When qwen3 models return tool_calls through Ollama, the finish_reason
was incorrectly left as "stop" instead of being set to "tool_calls".
This caused clients to miss the tool_calls in the response.
Added _get_finish_reason helper method following OpenAI provider's
pattern, and fixed both streaming and non-streaming response paths.
Fixes: https://github.com/BerriAI/litellm/issues/18922
* fix(ollama): pass tools directly without model capability check
The previous code tried to check model capability via get_model_info()
which made network calls to localhost:11434. When Ollama is remote,
this fails and falls back to JSON format, breaking tool calling.
Ollama 0.4+ supports native tool calling - let Ollama handle
model capability detection instead of LiteLLM.
Fixes #18922
* fix(ollama): transform tool_calls response to OpenAI format
Ollama returns tool_calls with arguments as dict, but OpenAI format
requires arguments to be a JSON string. Also ensures 'type': 'function'
field is present.
Completes the fix for #18922
* fix(ollama): set finish_reason to "tool_calls" when tool_calls present
Fixes #18922
Two issues addressed:
1. Remove broken model capability check
- get_model_info() fails when Ollama runs on remote server
- Broken fallback triggered JSON prompt injection
- Now passes tools directly - Ollama 0.4+ handles detection
2. Set finish_reason correctly
- Was hardcoded to "stop" even with tool_calls present
- Clients use this to know how to process the response
- Now returns "tool_calls" when tool_calls are in response
Both streaming and non-streaming responses are fixed.
Tests:
- All 14 existing Ollama tests pass
- Added 3 focused tests for the fixes
2026-01-14 03:52:26 +05:30
Cesar Garcia
d03c5017ff
fix: correct context window sizes for GPT-5 model variants ( #18928 )
...
* fix: correct context window sizes for GPT-5 model variants
Updates max_input_tokens for GPT-5, GPT-5.1, and GPT-5.2 model variants
to match OpenAI's official specifications, resolving issue #18927 .
Changes:
- GPT-5.1 (base, codex variants): 272k → 400k tokens
- GPT-5.1-chat variants: 272k → 128k tokens (with max_output 16,384)
- GPT-5 (base): 272k → 400k tokens
- GPT-5-chat: 272k → 128k tokens (with max_output 16,384)
- GPT-5-codex: 272k → 400k tokens
- GPT-5-mini: 272k → 400k tokens
- GPT-5-nano: 272k → 400k tokens
- GPT-5-pro: 272k → 400k tokens (max_output 272k)
- GPT-5 dated versions (2025-08-07): 272k → 400k tokens
Affected providers: OpenAI, Azure (all regions), OpenRouter
Fixes #18927
* fix: correct Azure GPT-5 context window limits to match Azure docs
Azure OpenAI has different limits than OpenAI for GPT-5 models.
Changes:
- Azure GPT-5 models: max_input_tokens 400k → 272k (Azure limit)
- Azure GPT-5 Pro: max_output_tokens 272k → 128k (Azure limit)
- OpenAI GPT-5 models: remain at 400k (correct)
- OpenRouter models: remain at 400k (routes to OpenAI)
Azure docs specify 272k input + 128k output = 400k total context.
OpenAI allows full 400k input + 128k output.
* fix: correct Azure GPT-5 max_input_tokens to 272k
Azure has explicit input limit of 272k tokens (not 400k like OpenAI).
Context window 400k = 272k input + 128k output for Azure.
OpenAI allows flexible input up to 400k (context - output).
2026-01-14 03:49:47 +05:30
Dominic Feliton
1f3d75a67e
Add QueryClient to model hub
2026-01-13 14:19:38 -08:00
xiaofan
f8836cb2a7
Fix Swagger UI path with server_root_path in OpenAPI schema ( #18947 )
...
Adds 'servers' field to OpenAPI schema when server_root_path is set, ensuring correct Swagger UI execute path for reverse proxies and subpath deployments. Includes tests to verify correct server URL handling for various root path formats.
2026-01-14 03:48:43 +05:30
Mateusz Szewczyk
72dc65fbb4
chore: allow passing scope id for watsonx inferencing ( #18959 )
...
* chore: allow inference with space
* make lint and make format
2026-01-14 03:47:20 +05:30
berkeyalciin
c2e01735e4
Fix: change delete() to delete_many() for prompt deletion to handle non-unique prompt_id ( #18966 )
...
Co-authored-by: Berke Yalcin <berke.yalcin@beko.com>
2026-01-14 03:42:43 +05:30
Matthias Humt
9adc19deab
Normalize OpenAI SDK BaseModel choices/messages to avoid Pydantic serializer warnings ( #18972 )
...
* Normalize BaseModel choices + suppress serializer warnings
* Fix ModelResponse normalization and test deps
2026-01-14 03:40:11 +05:30
Robin
b7c5662273
Fix: update novita models prices ( #19005 )
...
* feat: ci
* feat: fix novita models prices
2026-01-14 03:30:19 +05:30
houdataali
cbb72045a3
fix(ui): use non-streaming method for endpoint v1/a2a/message/send in… ( #19025 )
...
* Add end to end integration tests for batches
* Add end to end integration tests for batches
* Add end to end integration tests for batches
* Fix linter errors: remove unused imports and variables
* Add end to end integration tests for batches
* Add end to end integration tests for batches
* Add end to end integration tests for batches
* Add end to end integration tests for batches
* chore: document temporary grype ignore for CVE-2019-1010022
* chore: add config option
* chore: add ALLOWED_CVES
* refetch after key create
* test: remove flaky azure oidc embedding test
* fixing build
* bump: version 1.80.15 → 1.80.16
* [Fix] MSFT SSO - allow setting custom MSFT Base URLs (#18977 )
* fix TestCustomMicrosoftSSO
* init CustomMicrosoftSSO
* use CustomMicrosoftSSO
* docs fix
* docs fix
* [Feat] UI Feedback Form - why LiteLLM (#18999 )
* init survey prompt
* init survey modal
* init Survey Modal
* POST feedback hook
* survey Modal
* add other
* in product survey fixes
* fix survey prompt
* fix survey
* fix build
* ui new build
* [Feat] MSFT SSO - allow overriding env var attribute names (#18998 )
* add MSFT SSO constants
* fix MSFT SSO env vars
* test_microsoft_sso_handler_openid_from_response_with_custom_attributes
* Add pricing of azure_ai/claude-opus-4-5
* test: temporarily disable flaky responses_id_security tests
* fix(ui): use non-streaming method for endpoint v1/a2a/message/send in A2A playground
'
---------
Co-authored-by: Ephrim Stanley <ephrim.stanley@point72.com>
Co-authored-by: Yuta Saito <uc4w6c@bma.biglobe.ne.jp>
Co-authored-by: yuneng-jiang <yuneng.jiang@gmail.com>
Co-authored-by: YutaSaito <36355491+uc4w6c@users.noreply.github.com>
Co-authored-by: Ishaan Jaff <ishaanjaffer0324@gmail.com>
Co-authored-by: Sameer Kankute <sameer@berri.ai>
2026-01-14 03:29:10 +05:30
Harshit Jain
181c626d83
fix: properly handle custom guardrails parameters ( #18978 )
2026-01-14 03:23:54 +05:30
Harshit Jain
ebeea46fc4
fix(langsmith.py): hoist thread grouping metadata (session_id, thread_id, conversation_id) ( #18982 )
2026-01-14 03:21:37 +05:30
Yuta Saito
22aad95bb1
fix: rest_endpoints own allowed_mcp_servers for MCP calls
2026-01-14 06:39:46 +09:00
Jack Temple
9e08c2207f
fix: enable JSON logging via configuration and add regression test
2026-01-13 09:38:19 -07:00
Sameer Kankute
93203cda7c
Merge pull request #19027 from BerriAI/litellm_add_0_budget_model_bypass
...
[Feat] Add support for 0 cost models
2026-01-13 18:05:36 +05:30
Sameer Kankute
e98c2e4425
Merge pull request #19012 from BerriAI/litellm_fix_model_deployment_routing
...
Fix: Model matching priority in configuration
2026-01-13 17:55:01 +05:30
Sameer Kankute
bb0ab38636
Merge pull request #19007 from BerriAI/litellm_fix_header_forwarding_passthrough
...
Fix: Header forwarding in bedrock passthrough
2026-01-13 17:50:14 +05:30
Sameer Kankute
54f6f55c98
Merge pull request #18873 from BerriAI/litellm_staging_01_09_2026
...
staging 01/09/2025
2026-01-13 17:31:44 +05:30
Sameer Kankute
f2cb861d6a
Merge pull request #18976 from BerriAI/litellm_staging_12_19_2025
...
Staging 12/19/2025 - implement failopen option default to True on grayswan guardrail (#18266 )
2026-01-13 17:00:36 +05:30
Sameer Kankute
de6330b6b6
Fix test_async_otel_callback[False]
2026-01-13 16:59:17 +05:30
Sameer Kankute
1932d03aed
Add docs on Zero-Cost Models
2026-01-13 16:44:02 +05:30
Sameer Kankute
762a3ef090
Add support for 0 cost models
2026-01-13 16:39:57 +05:30
Sameer Kankute
d656f01bc9
Merge pull request #19009 from Dima-Mediator/fix-image-tokens-spend-logging
...
Fix image tokens spend logging for /images/generations
2026-01-13 15:03:37 +05:30
Igal Boxerman
8cff86ff01
fix(guardrails): use clean error messages for blocked requests ( #19022 )
...
- Add `should_wrap_with_default_message` parameter to GuardrailRaisedException
- Update Generic Guardrail API to use clean error messages without wrapper
- When should_wrap_with_default_message=False, exception shows the original
blocked_reason directly (e.g., "pii detected") instead of verbose format
- Update test to verify GuardrailRaisedException is raised with clean message
2026-01-13 11:02:06 +02:00
Sameer Kankute
5349d8922c
Merge pull request #19003 from BerriAI/litellm_add_azure_ai_claude_opus
...
Add pricing of azure_ai/claude-opus-4-5
2026-01-13 13:52:47 +05:30
Yuta Saito
658bbcc2d5
fix: mcp rest auth check
2026-01-13 17:18:58 +09:00
YutaSaito
b3e126222f
Merge pull request #19013 from BerriAI/litellm_test_comment_out_flaky
...
[test] temporarily disable flaky responses_id_security tests
2026-01-13 15:52:57 +09:00
Yuta Saito
2c8ac2c3f1
test: temporarily disable flaky responses_id_security tests
2026-01-13 15:51:37 +09:00
Sameer Kankute
ecb3959c3c
Merge pull request #18208 from Chesars/fix/case-insensitive-model-cost-lookup
...
fix: case-insensitive model cost map lookup
2026-01-13 11:53:43 +05:30
Sameer Kankute
dfece51f8c
Fix: Model matching priority in configuration
2026-01-13 11:44:48 +05:30
yuneng-jiang
da902c5c55
Migrate User and Team filters to use reusable components
2026-01-12 21:04:16 -08:00
Sameer Kankute
005541075b
Fix: Header forwarding in bedrock passthrough
2026-01-13 09:45:14 +05:30
Dima-Mediator
7c61933bc5
Fix image tokens spend logging for /images/generations
2026-01-12 23:07:08 -05:00
Sameer Kankute
5a51b74658
Add pricing of azure_ai/claude-opus-4-5
2026-01-13 09:15:05 +05:30
Ishaan Jaff
a1bba8c99b
[Feat] MSFT SSO - allow overriding env var attribute names ( #18998 )
...
* add MSFT SSO constants
* fix MSFT SSO env vars
* test_microsoft_sso_handler_openid_from_response_with_custom_attributes
2026-01-12 18:56:35 -08:00
Ishaan Jaffer
0feedfdf3d
ui new build
2026-01-12 18:55:18 -08:00
Ishaan Jaffer
dd959790bb
fix build
2026-01-12 18:53:42 -08:00
Ishaan Jaff
c4e6ae4d9e
[Feat] UI Feedback Form - why LiteLLM ( #18999 )
...
* init survey prompt
* init survey modal
* init Survey Modal
* POST feedback hook
* survey Modal
* add other
* in product survey fixes
* fix survey prompt
* fix survey
2026-01-12 18:48:17 -08:00
Sameer Kankute
a727aa9980
Merge pull request #18340 from Point72/ephrimstanley/fix-batch
...
Fix batch deletion and retrieve
2026-01-13 08:13:43 +05:30
yuneng-jiang
9875ef9720
Remove Notification Check
2026-01-12 18:34:06 -08:00
yuneng-jiang
0dbeb56346
Simplify Key Generate Permission Error
2026-01-12 18:28:29 -08:00
YutaSaito
5bbd22070b
Merge pull request #18996 from BerriAI/litellm_release
...
bump: version 1.80.15 → 1.80.16
2026-01-13 11:27:35 +09:00
Ishaan Jaff
21d611554b
[Fix] MSFT SSO - allow setting custom MSFT Base URLs ( #18977 )
...
* fix TestCustomMicrosoftSSO
* init CustomMicrosoftSSO
* use CustomMicrosoftSSO
* docs fix
* docs fix
2026-01-12 18:26:53 -08:00
Yuta Saito
f149491498
bump: version 1.80.15 → 1.80.16
2026-01-13 11:21:57 +09:00
yuneng-jiang
81b1becd95
remove only
2026-01-12 18:06:41 -08:00
yuneng-jiang
caa26de581
Merge remote-tracking branch 'origin' into litellm_e2e_create_key_test
2026-01-12 17:57:52 -08:00
yuneng-jiang
a135bfc376
Add create key test
2026-01-12 17:57:22 -08:00
yuneng-jiang
c126cfd8db
Merge pull request #18994 from BerriAI/litellm_ui_key_refresh_fix
...
[Fix] UI - Refetch Keys after Key Create
2026-01-12 17:56:47 -08:00
YutaSaito
474107da0c
Merge pull request #18987 from BerriAI/litellm_fix_security_test
...
[fix] security test
2026-01-13 10:41:13 +09:00