Go to file
Idris Mokhtarzada 9b89280a90
Use underscores
Datadog does not play nice with special characters (as in "(seconds)").  Also just makes sense to standardize on either underscores or camelCase, but not mix-and-match.
2024-07-26 16:38:54 -04:00
.circleci support using */* 2024-07-25 18:48:56 -07:00
.devcontainer Add devcontainer. 2024-05-07 11:33:04 +00:00
.github build(ghcr_deploy.yml): fix discord release note push 2024-07-06 19:06:04 -07:00
ci_cd (fix) pre commit hook to sync backup context_window mapping 2024-02-05 15:03:04 -08:00
cookbook docs grafana dashboard litellm 2024-07-05 14:08:55 -07:00
deploy move index and helm pacakge location 2024-07-10 00:54:36 +08:00
docker Revert "build(Dockerfile): move prisma build to dockerfile" 2024-01-06 09:51:44 +05:30
docs/my-website docs(stream.md): add streaming token usage info to docs 2024-07-26 10:51:17 -07:00
enterprise fix lakera ai tests 2024-07-19 07:49:57 -07:00
litellm Use underscores 2024-07-26 16:38:54 -04:00
litellm-js build(deps): bump @hono/node-server in /litellm-js/spend-logs 2024-04-25 23:43:28 +00:00
tests test proxy all model 2024-07-25 18:54:30 -07:00
ui Revert "[Ui] add together AI, Mistral, PerplexityAI, OpenRouter models on Admin UI " 2024-07-20 19:04:22 -07:00
.dockerignore Fix .dockerignore 2024-04-12 11:06:24 +01:00
.env.example feat: added support for OPENAI_API_BASE 2023-08-28 14:57:34 +02:00
.flake8 chore: list all ignored flake8 rules explicit 2023-12-23 09:07:59 +01:00
.git-blame-ignore-revs Add my commit to .git-blame-ignore-revs 2024-05-12 10:21:10 -07:00
.gitattributes ignore ipynbs 2023-08-31 16:58:54 -07:00
.gitignore fix(team_endpoints.py): fix check 2024-07-16 22:05:48 -07:00
.pre-commit-config.yaml Added poetry-check to pre-commit 2024-07-02 13:02:26 -04:00
build_admin_ui.sh (fix) build command 2024-02-21 21:09:15 -08:00
check_file_length.py refactor(check_file_length.py): add local pre-commit check for file length 2024-06-15 09:18:53 -07:00
docker-compose.yml build(docker-compose.yml): add prometheus scraper to docker compose 2024-07-24 10:09:23 -07:00
Dockerfile build(dockerfile): remove --config proxy_server_config.yaml from docker run 2024-04-08 13:23:56 -07:00
Dockerfile.alpine (fix) alpine Docker image 2024-01-10 22:18:37 +05:30
Dockerfile.database Revert "Create litellm user to fix issue with prisma in k8s " 2024-06-25 18:19:24 -07:00
entrypoint.sh fix(prisma_migration.py): support decrypting variables in a python script 2024-06-28 16:31:37 -07:00
index.yaml update index.yaml 2024-07-10 01:00:10 +08:00
LICENSE refactor: creating enterprise folder 2024-02-15 12:54:13 -08:00
litellm-helm-0.2.0.tgz build(bump-helm-chart-app-version): bump helm chart app version to latest 2024-05-06 10:26:01 -07:00
litellm-helm-0.2.1.tgz update index.yaml 2024-07-10 01:00:10 +08:00
log.txt fix(vertex_httpx.py): add function calling support to httpx route 2024-06-12 21:11:00 -07:00
model_prices_and_context_window.json Merge branch 'main' into bedrock-llama3.1-405b 2024-07-25 19:29:10 -07:00
mypy.ini ci(mypy.ini): ignore missing imports 2024-04-04 10:19:13 -07:00
package-lock.json (fix) create key flow 2024-03-29 10:08:35 -07:00
package.json (fix) create key flow 2024-03-29 10:08:35 -07:00
poetry.lock Bump azure-identity from 1.16.0 to 1.16.1 2024-07-12 01:28:58 +00:00
prometheus.yml build(docker-compose.yml): add prometheus scraper to docker compose 2024-07-24 10:09:23 -07:00
proxy_server_config.yaml support using */* 2024-07-25 18:48:56 -07:00
pyproject.toml bump: version 1.42.2 → 1.42.3 2024-07-25 22:18:17 -07:00
README.md Update README.md 2024-07-25 20:12:32 -07:00
render.yaml build(render.yaml): fix health check route 2024-05-24 09:45:28 -07:00
requirements.txt build(requirements.txt): bump openai version 2024-07-10 11:52:29 -07:00
retry_push.sh build(Dockerfile): moves prisma logic to dockerfile 2024-01-06 14:59:10 +05:30
ruff.toml fix(utils.py): improved predibase exception mapping 2024-06-08 14:32:43 -07:00
schema.prisma fix DB accept null values for api_base, user, etc 2024-07-23 16:33:04 -07:00

🚅 LiteLLM

Deploy to Render Deploy on Railway

Call all LLM APIs using the OpenAI format [Bedrock, Huggingface, VertexAI, TogetherAI, Azure, OpenAI, Groq etc.]

OpenAI Proxy Server | Hosted Proxy (Preview) | Enterprise Tier

PyPI Version CircleCI Y Combinator W23 Whatsapp Discord

LiteLLM manages:

  • Translate inputs to provider's completion, embedding, and image_generation endpoints
  • Consistent output, text responses will always be available at ['choices'][0]['message']['content']
  • Retry/fallback logic across multiple deployments (e.g. Azure/OpenAI) - Router
  • Set Budgets & Rate limits per project, api key, model OpenAI Proxy Server

Jump to OpenAI Proxy Docs
Jump to Supported LLM Providers

🚨 Stable Release: Use docker images with the -stable tag. These have undergone 12 hour load tests, before being published.

Support for more providers. Missing a provider or LLM Platform, raise a feature request.

Usage (Docs)

Important

LiteLLM v1.0.0 now requires openai>=1.0.0. Migration guide here
LiteLLM v1.40.14+ now requires pydantic>=2.0.0. No changes required.

Open In Colab
pip install litellm
from litellm import completion
import os

## set ENV variables
os.environ["OPENAI_API_KEY"] = "your-openai-key"
os.environ["COHERE_API_KEY"] = "your-cohere-key"

messages = [{ "content": "Hello, how are you?","role": "user"}]

# openai call
response = completion(model="gpt-3.5-turbo", messages=messages)

# cohere call
response = completion(model="command-nightly", messages=messages)
print(response)

Call any model supported by a provider, with model=<provider_name>/<model_name>. There might be provider-specific details here, so refer to provider docs for more information

Async (Docs)

from litellm import acompletion
import asyncio

async def test_get_response():
    user_message = "Hello, how are you?"
    messages = [{"content": user_message, "role": "user"}]
    response = await acompletion(model="gpt-3.5-turbo", messages=messages)
    return response

response = asyncio.run(test_get_response())
print(response)

Streaming (Docs)

liteLLM supports streaming the model response back, pass stream=True to get a streaming iterator in response.
Streaming is supported for all models (Bedrock, Huggingface, TogetherAI, Azure, OpenAI, etc.)

from litellm import completion
response = completion(model="gpt-3.5-turbo", messages=messages, stream=True)
for part in response:
    print(part.choices[0].delta.content or "")

# claude 2
response = completion('claude-2', messages, stream=True)
for part in response:
    print(part.choices[0].delta.content or "")

Logging Observability (Docs)

LiteLLM exposes pre defined callbacks to send data to Lunary, Langfuse, DynamoDB, s3 Buckets, Helicone, Promptlayer, Traceloop, Athina, Slack

from litellm import completion

## set env variables for logging tools
os.environ["LUNARY_PUBLIC_KEY"] = "your-lunary-public-key"
os.environ["HELICONE_API_KEY"] = "your-helicone-auth-key"
os.environ["LANGFUSE_PUBLIC_KEY"] = ""
os.environ["LANGFUSE_SECRET_KEY"] = ""
os.environ["ATHINA_API_KEY"] = "your-athina-api-key"

os.environ["OPENAI_API_KEY"]

# set callbacks
litellm.success_callback = ["lunary", "langfuse", "athina", "helicone"] # log input/output to lunary, langfuse, supabase, athina, helicone etc

#openai call
response = completion(model="gpt-3.5-turbo", messages=[{"role": "user", "content": "Hi 👋 - i'm openai"}])

OpenAI Proxy - (Docs)

Track spend + Load Balance across multiple projects

Hosted Proxy (Preview)

The proxy provides:

  1. Hooks for auth
  2. Hooks for logging
  3. Cost tracking
  4. Rate Limiting

📖 Proxy Endpoints - Swagger Docs

Quick Start Proxy - CLI

pip install 'litellm[proxy]'

Step 1: Start litellm proxy

$ litellm --model huggingface/bigcode/starcoder

#INFO: Proxy running on http://0.0.0.0:4000

Step 2: Make ChatCompletions Request to Proxy

Important

💡 Use LiteLLM Proxy with Langchain (Python, JS), OpenAI SDK (Python, JS) Anthropic SDK, Mistral SDK, LlamaIndex, Instructor, Curl

import openai # openai v1.0.0+
client = openai.OpenAI(api_key="anything",base_url="http://0.0.0.0:4000") # set proxy to base_url
# request sent to model set on litellm proxy, `litellm --model`
response = client.chat.completions.create(model="gpt-3.5-turbo", messages = [
    {
        "role": "user",
        "content": "this is a test request, write a short poem"
    }
])

print(response)

Proxy Key Management (Docs)

Connect the proxy with a Postgres DB to create proxy keys

# Get the code
git clone https://github.com/BerriAI/litellm

# Go to folder
cd litellm

# Add the master key - you can change this after setup
echo 'LITELLM_MASTER_KEY="sk-1234"' > .env

# Add the litellm salt key - you cannot change this after adding a model
# It is used to encrypt / decrypt your LLM API Key credentials
# We recommned - https://1password.com/password-generator/ 
# password generator to get a random hash for litellm salt key
echo 'LITELLM_SALT_KEY="sk-1234"' > .env

source .env

# Start
docker-compose up

UI on /ui on your proxy server ui_3

Set budgets and rate limits across multiple projects POST /key/generate

Request

curl 'http://0.0.0.0:4000/key/generate' \
--header 'Authorization: Bearer sk-1234' \
--header 'Content-Type: application/json' \
--data-raw '{"models": ["gpt-3.5-turbo", "gpt-4", "claude-2"], "duration": "20m","metadata": {"user": "ishaan@berri.ai", "team": "core-infra"}}'

Expected Response

{
    "key": "sk-kdEXbIqZRwEeEiHwdg7sFA", # Bearer token
    "expires": "2023-11-19T01:38:25.838000+00:00" # datetime object
}

Supported Providers (Docs)

Provider Completion Streaming Async Completion Async Streaming Async Embedding Async Image Generation
openai
azure
aws - sagemaker
aws - bedrock
google - vertex_ai
google - palm
google AI Studio - gemini
mistral ai api
cloudflare AI Workers
cohere
anthropic
empower
huggingface
replicate
together_ai
openrouter
ai21
baseten
vllm
nlp_cloud
aleph alpha
petals
ollama
deepinfra
perplexity-ai
Groq AI
Deepseek
anyscale
IBM - watsonx.ai
voyage ai
xinference [Xorbits Inference]
FriendliAI

Read the Docs

Contributing

To contribute: Clone the repo locally -> Make a change -> Submit a PR with the change.

Here's how to modify the repo locally: Step 1: Clone the repo

git clone https://github.com/BerriAI/litellm.git

Step 2: Navigate into the project, and install dependencies:

cd litellm
poetry install -E extra_proxy -E proxy

Step 3: Test your change:

cd litellm/tests # pwd: Documents/litellm/litellm/tests
poetry run flake8
poetry run pytest .

Step 4: Submit a PR with your changes! 🚀

  • push your fork to your GitHub repo
  • submit a PR from there

Enterprise

For companies that need better security, user management and professional support

Talk to founders

This covers:

  • Features under the LiteLLM Commercial License:
  • Feature Prioritization
  • Custom Integrations
  • Professional Support - Dedicated discord + slack
  • Custom SLAs
  • Secure access with Single Sign-On

Support / talk with founders

Why did we build this

  • Need for simplicity: Our code started to get extremely complicated managing & translating calls between Azure, OpenAI and Cohere.

Contributors