Go to file
Eslam karim gaber 48dbdaa73e
Change quota project to the correct project being used for the call
if not set it will use the default project in the ADC to set that quota project which is usually different 
https://github.com/googleapis/python-aiplatform/issues/2557#issuecomment-1709284744
2024-01-28 19:55:01 +02:00
.circleci fix(utils.py): fix sagemaker async logging for sync streaming 2024-01-25 12:49:45 -08:00
.github Update ghcr_deploy.yml 2024-01-26 20:29:52 -08:00
cookbook (docs) misc/cookbook - OpenAI python timeout 2024-01-23 19:31:31 -08:00
dist fix: syncing changes 2024-01-12 11:41:40 +05:30
docker Revert "build(Dockerfile): move prisma build to dockerfile" 2024-01-06 09:51:44 +05:30
docs/my-website docs(virtual_keys.md): add key alias and key name to docs 2024-01-27 20:03:07 -08:00
litellm Change quota project to the correct project being used for the call 2024-01-28 19:55:01 +02:00
tests (test) /key/info 2024-01-26 19:26:55 -08:00
ui (ui) ui cleanup 2024-01-27 19:37:42 -08:00
.env.example feat: added support for OPENAI_API_BASE 2023-08-28 14:57:34 +02:00
.flake8 chore: list all ignored flake8 rules explicit 2023-12-23 09:07:59 +01:00
.gitattributes ignore ipynbs 2023-08-31 16:58:54 -07:00
.gitignore (chore) gitignore 2024-01-27 17:32:30 -08:00
.pre-commit-config.yaml fix(dynamo_db.py): if table create fails, tell user what the table + hash key needs to be 2024-01-11 23:01:28 +05:30
docker-compose.yml (ci/cd) docker compose up with ui 2024-01-25 17:13:19 -08:00
Dockerfile Revert "Merge branch 'main' into main" 2024-01-27 18:19:23 -08:00
Dockerfile.alpine (fix) alpine Docker image 2024-01-10 22:18:37 +05:30
Dockerfile.database build(dockerfile.database): clean up 2024-01-10 19:51:39 +05:30
entrypoint.sh (ci/cd) set litellm as entrypoint 2024-01-10 15:15:49 +05:30
LICENSE Initial commit 2023-07-26 17:09:52 -07:00
model_prices_and_context_window.json (feat) add gpt-4-0125-preview 2024-01-25 16:40:23 -08:00
mypy.ini fix(google_kms.py): support enums for key management system 2023-12-27 13:19:33 +05:30
poetry.lock (chore) bump poetry lock 2024-01-26 10:34:16 -08:00
proxy_server_config.yaml fix(utils.py): fix sagemaker async logging for sync streaming 2024-01-25 12:49:45 -08:00
pyproject.toml bump: version 1.20.0 → 1.20.1 2024-01-27 18:18:59 -08:00
README.md Update README.md 2024-01-11 09:00:33 +05:30
requirements.txt build(requirements.txt): add apscheduler to requirements 2024-01-23 17:56:11 -08:00
retry_push.sh build(Dockerfile): moves prisma logic to dockerfile 2024-01-06 14:59:10 +05:30
schema.prisma build(schema.prisma): update schema 2024-01-26 20:53:07 -08:00
template.yaml Use -function for naming. 2023-11-23 02:09:09 -05:00

🚅 LiteLLM

Call all LLM APIs using the OpenAI format [Bedrock, Huggingface, VertexAI, TogetherAI, Azure, OpenAI, etc.]

OpenAI Proxy Server

PyPI Version CircleCI Y Combinator W23 Whatsapp Discord

LiteLLM manages:

  • Translate inputs to provider's completion, embedding, and image_generation endpoints
  • Consistent output, text responses will always be available at ['choices'][0]['message']['content']
  • Retry/fallback logic across multiple deployments (e.g. Azure/OpenAI) - Router

Jump to OpenAI Proxy Docs
Jump to Supported LLM Providers

Usage (Docs)

Important

LiteLLM v1.0.0 now requires openai>=1.0.0. Migration guide here

Open In Colab
pip install litellm
from litellm import completion
import os

## set ENV variables 
os.environ["OPENAI_API_KEY"] = "your-openai-key" 
os.environ["COHERE_API_KEY"] = "your-cohere-key" 

messages = [{ "content": "Hello, how are you?","role": "user"}]

# openai call
response = completion(model="gpt-3.5-turbo", messages=messages)

# cohere call
response = completion(model="command-nightly", messages=messages)
print(response)

Async (Docs)

from litellm import acompletion
import asyncio

async def test_get_response():
    user_message = "Hello, how are you?"
    messages = [{"content": user_message, "role": "user"}]
    response = await acompletion(model="gpt-3.5-turbo", messages=messages)
    return response

response = asyncio.run(test_get_response())
print(response)

Streaming (Docs)

liteLLM supports streaming the model response back, pass stream=True to get a streaming iterator in response.
Streaming is supported for all models (Bedrock, Huggingface, TogetherAI, Azure, OpenAI, etc.)

from litellm import completion
response = completion(model="gpt-3.5-turbo", messages=messages, stream=True)
for part in response:
    print(part.choices[0].delta.content or "")

# claude 2
response = completion('claude-2', messages, stream=True)
for part in response:
    print(part.choices[0].delta.content or "")

Logging Observability (Docs)

LiteLLM exposes pre defined callbacks to send data to Langfuse, DynamoDB, s3 Buckets, LLMonitor, Helicone, Promptlayer, Traceloop, Slack

from litellm import completion

## set env variables for logging tools
os.environ["LANGFUSE_PUBLIC_KEY"] = ""
os.environ["LANGFUSE_SECRET_KEY"] = ""
os.environ["LLMONITOR_APP_ID"] = "your-llmonitor-app-id"

os.environ["OPENAI_API_KEY"]

# set callbacks
litellm.success_callback = ["langfuse", "llmonitor"] # log input/output to langfuse, llmonitor, supabase

#openai call
response = completion(model="gpt-3.5-turbo", messages=[{"role": "user", "content": "Hi 👋 - i'm openai"}])

OpenAI Proxy - (Docs)

Track spend across multiple projects/people

The proxy provides:

  1. Hooks for auth
  2. Hooks for logging
  3. Cost tracking
  4. Rate Limiting

📖 Proxy Endpoints - Swagger Docs

Quick Start Proxy - CLI

pip install 'litellm[proxy]'

Step 1: Start litellm proxy

$ litellm --model huggingface/bigcode/starcoder

#INFO: Proxy running on http://0.0.0.0:8000

Step 2: Make ChatCompletions Request to Proxy

import openai # openai v1.0.0+
client = openai.OpenAI(api_key="anything",base_url="http://0.0.0.0:8000") # set proxy to base_url
# request sent to model set on litellm proxy, `litellm --model`
response = client.chat.completions.create(model="gpt-3.5-turbo", messages = [
    {
        "role": "user",
        "content": "this is a test request, write a short poem"
    }
])

print(response)

Proxy Key Management (Docs)

Track Spend, Set budgets and create virtual keys for the proxy POST /key/generate

Request

curl 'http://0.0.0.0:8000/key/generate' \
--header 'Authorization: Bearer sk-1234' \
--header 'Content-Type: application/json' \
--data-raw '{"models": ["gpt-3.5-turbo", "gpt-4", "claude-2"], "duration": "20m","metadata": {"user": "ishaan@berri.ai", "team": "core-infra"}}'

Expected Response

{
    "key": "sk-kdEXbIqZRwEeEiHwdg7sFA", # Bearer token
    "expires": "2023-11-19T01:38:25.838000+00:00" # datetime object
}

[Beta] Proxy UI

A simple UI to add new models and let your users create keys.

Live here: https://dashboard.litellm.ai/

Code: https://github.com/BerriAI/litellm/tree/main/ui

Screenshot 2023-12-26 at 8 33 53 AM

Supported Providers (Docs)

Provider Completion Streaming Async Completion Async Streaming Async Embedding Async Image Generation
openai
azure
aws - sagemaker
aws - bedrock
google - vertex_ai [Gemini]
google - palm
google AI Studio - gemini
mistral ai api
cloudflare AI Workers
cohere
anthropic
huggingface
replicate
together_ai
openrouter
ai21
baseten
vllm
nlp_cloud
aleph alpha
petals
ollama
deepinfra
perplexity-ai
anyscale
voyage ai
xinference [Xorbits Inference]

Read the Docs

Contributing

To contribute: Clone the repo locally -> Make a change -> Submit a PR with the change.

Here's how to modify the repo locally: Step 1: Clone the repo

git clone https://github.com/BerriAI/litellm.git

Step 2: Navigate into the project, and install dependencies:

cd litellm
poetry install

Step 3: Test your change:

cd litellm/tests # pwd: Documents/litellm/litellm/tests
poetry run flake8
poetry run pytest .

Step 4: Submit a PR with your changes! 🚀

  • push your fork to your GitHub repo
  • submit a PR from there

Support / talk with founders

Why did we build this

  • Need for simplicity: Our code started to get extremely complicated managing & translating calls between Azure, OpenAI and Cohere.

Contributors