Skip to navigation

LiteLLM Provider Integration

LiteLLM Provider Integration

Use GovernanceAI as a drop-in LiteLLM-compatible proxy to add AI governance to any application using LiteLLM.

Overview

GovernanceAI provides a LiteLLM-compatible API endpoint that:

  • ✅ Accepts same requests as OpenAI/Claude/etc.
  • ✅ Applies guardrails and policies
  • ✅ Returns policy-compliant responses
  • ✅ Logs all activity for audit

No code changes needed - Just point your LiteLLM client to GovernanceAI.

Setup

Step 1: Get Proxy Endpoint

  • Integrations → LiteLLM
  • Copy your endpoint: https://litellm.governanceai.com/v1
  • Generate API key (or use existing)

Step 2: Configure LiteLLM

import litellm
# Point LiteLLM to GovernanceAI
litellm.api_base = "https://litellm.governanceai.com/v1"
litellm.api_key = "gak_prod_your_api_key"
# Use normally - guardrails applied automatically
response = litellm.completion(
model="openai/gpt-4",
messages=[{"role": "user", "content": "Hello"}]
)
print(response.choices[0].message.content)

Step 3: Supported Models

All models are supported by passing through to their provider:

# OpenAI models
litellm.completion(model="openai/gpt-4", ...)
litellm.completion(model="openai/gpt-3.5-turbo", ...)
# Claude models
litellm.completion(model="claude-3-opus", ...)
# Cohere models
litellm.completion(model="cohere/command", ...)
# And more...

Configuration

Model Routing

Route different models through different guardrails:

curl -X POST https://api.governanceai.com/v1/litellm/model-routing \
-H "Authorization: Bearer $API_KEY" \
-d '{
"routes": [
{
"model_pattern": "gpt-4",
"guardrail_policy": "strict"
},
{
"model_pattern": "gpt-3.5-turbo",
"guardrail_policy": "standard"
},
{
"model_pattern": "claude-*",
"guardrail_policy": "standard"
}
]
}'

Rate Limiting

Configure per-model rate limits:

curl -X POST https://api.governanceai.com/v1/litellm/rate-limits \
-H "Authorization: Bearer $API_KEY" \
-d '{
"limits": [
{
"model": "gpt-4",
"requests_per_minute": 60,
"tokens_per_minute": 300000
},
{
"model": "*",
"requests_per_minute": 1000,
"tokens_per_minute": 1000000
}
]
}'

Usage Example

Python Application

import litellm
import json
# Configure
litellm.api_base = "https://litellm.governanceai.com/v1"
litellm.api_key = "gak_prod_..."
# Make request with context (optional)
response = litellm.completion(
model="openai/gpt-4",
messages=[{
"role": "user",
"content": "What is the capital of France?"
}],
# GovernanceAI-specific context
metadata={
"org_id": "org_123",
"user_id": "user_456",
"workspace_id": "ws_789"
}
)
# Response includes GovernanceAI metadata
print(f"Content: {response.choices[0].message.content}")
print(f"Risk Score: {response.risk_score}") # GovernanceAI addition
print(f"Policy Violations: {response.policy_violations}") # GovernanceAI addition

LangChain Integration

from langchain.chat_models import ChatOpenAI
from langchain.schema import HumanMessage
# Configure to use GovernanceAI
chat = ChatOpenAI(
model_name="gpt-4",
openai_api_base="https://litellm.governanceai.com/v1",
openai_api_key="gak_prod_...",
temperature=0
)
# Use normally - all requests go through GovernanceAI
messages = [HumanMessage(content="Hello!")]
response = chat(messages)
print(response.content)

LlamaIndex Integration

from llama_index.llms import OpenAI
# Use GovernanceAI endpoint
llm = OpenAI(
model="gpt-4",
api_base="https://litellm.governanceai.com/v1",
api_key="gak_prod_..."
)
# All requests go through GovernanceAI guardrails
response = llm.complete("What is AI governance?")

Monitoring & Metrics

View Usage

# Get LiteLLM endpoint usage
curl -H "Authorization: Bearer $API_KEY" \
https://api.governanceai.com/v1/litellm/usage
# Returns:
{
"total_requests": 45230,
"total_tokens": 12453000,
"avg_latency_ms": 245,
"policy_violations": 123,
"blocked_requests": 45,
"transformed_responses": 78
}

Per-Model Metrics

curl -H "Authorization: Bearer $API_KEY" \
'https://api.governanceai.com/v1/litellm/usage/by-model'
# Returns metrics per model (gpt-4, gpt-3.5-turbo, etc.)

Error Handling

GovernanceAI returns standard OpenAI error codes:

try:
response = litellm.completion(
model="openai/gpt-4",
messages=[...]
)
except litellm.APIError as e:
# Handle API errors
print(f"Error: {e.http_status} - {e.message}")
# Common errors:
# 400 - Invalid request (malformed guardrail config)
# 401 - Authentication failed (invalid API key)
# 429 - Rate limit exceeded
# 500 - Server error (try again)

Performance

Latency Impact

GovernanceAI adds minimal latency:

  • Average overhead: 45-100ms
  • P95: 150ms
  • P99: 250ms

Varies based on:

  • Policy complexity
  • Model response size
  • Network latency to provider

Caching

Enable response caching:

curl -X POST https://api.governanceai.com/v1/litellm/caching \
-H "Authorization: Bearer $API_KEY" \
-d '{
"enabled": true,
"ttl_seconds": 3600,
"cache_identical_requests": true
}'

Best Practices

✅ Do:

  • Use org_id and user_id in metadata
  • Set appropriate rate limits
  • Monitor usage regularly
  • Test policies before production
  • Use different keys per environment

❌ Don’t:

  • Share API keys between environments
  • Disable logging for audit trails
  • Route sensitive data without PII guardrails
  • Forget to set up alerts

Troubleshooting

  • 401 Unauthorized - Check API key
  • Rate limit exceeded - Check configured limits
  • Slow responses - Check policy complexity
  • Connection refused - Verify endpoint URL

Next Steps