> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.governanceaicore.com/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.governanceaicore.com/_mcp/server.

# LiteLLM Provider Integration

> Use GovernanceAI as a LiteLLM-compatible LLM provider

# LiteLLM Provider Integration

Use GovernanceAI as a drop-in LiteLLM-compatible proxy to add AI governance to any application using LiteLLM.

## Overview

GovernanceAI provides a LiteLLM-compatible API endpoint that:

* ✅ Accepts same requests as OpenAI/Claude/etc.
* ✅ Applies guardrails and policies
* ✅ Returns policy-compliant responses
* ✅ Logs all activity for audit

**No code changes needed** - Just point your LiteLLM client to GovernanceAI.

## Setup

### Step 1: Get Proxy Endpoint

* **Integrations** → **LiteLLM**
* Copy your endpoint: `https://litellm.governanceai.com/v1`
* Generate API key (or use existing)

### Step 2: Configure LiteLLM

```python
import litellm

# Point LiteLLM to GovernanceAI
litellm.api_base = "https://litellm.governanceai.com/v1"
litellm.api_key = "gak_prod_your_api_key"

# Use normally - guardrails applied automatically
response = litellm.completion(
    model="openai/gpt-4",
    messages=[{"role": "user", "content": "Hello"}]
)

print(response.choices[0].message.content)
```

### Step 3: Supported Models

All models are supported by passing through to their provider:

```python
# OpenAI models
litellm.completion(model="openai/gpt-4", ...)
litellm.completion(model="openai/gpt-3.5-turbo", ...)

# Claude models
litellm.completion(model="claude-3-opus", ...)

# Cohere models
litellm.completion(model="cohere/command", ...)

# And more...
```

## Configuration

### Model Routing

Route different models through different guardrails:

```bash
curl -X POST https://api.governanceai.com/v1/litellm/model-routing \
  -H "Authorization: Bearer $API_KEY" \
  -d '{
    "routes": [
      {
        "model_pattern": "gpt-4",
        "guardrail_policy": "strict"
      },
      {
        "model_pattern": "gpt-3.5-turbo",
        "guardrail_policy": "standard"
      },
      {
        "model_pattern": "claude-*",
        "guardrail_policy": "standard"
      }
    ]
  }'
```

### Rate Limiting

Configure per-model rate limits:

```bash
curl -X POST https://api.governanceai.com/v1/litellm/rate-limits \
  -H "Authorization: Bearer $API_KEY" \
  -d '{
    "limits": [
      {
        "model": "gpt-4",
        "requests_per_minute": 60,
        "tokens_per_minute": 300000
      },
      {
        "model": "*",
        "requests_per_minute": 1000,
        "tokens_per_minute": 1000000
      }
    ]
  }'
```

## Usage Example

### Python Application

```python
import litellm
import json

# Configure
litellm.api_base = "https://litellm.governanceai.com/v1"
litellm.api_key = "gak_prod_..."

# Make request with context (optional)
response = litellm.completion(
    model="openai/gpt-4",
    messages=[{
        "role": "user",
        "content": "What is the capital of France?"
    }],
    # GovernanceAI-specific context
    metadata={
        "org_id": "org_123",
        "user_id": "user_456",
        "workspace_id": "ws_789"
    }
)

# Response includes GovernanceAI metadata
print(f"Content: {response.choices[0].message.content}")
print(f"Risk Score: {response.risk_score}")  # GovernanceAI addition
print(f"Policy Violations: {response.policy_violations}")  # GovernanceAI addition
```

### LangChain Integration

```python
from langchain.chat_models import ChatOpenAI
from langchain.schema import HumanMessage

# Configure to use GovernanceAI
chat = ChatOpenAI(
    model_name="gpt-4",
    openai_api_base="https://litellm.governanceai.com/v1",
    openai_api_key="gak_prod_...",
    temperature=0
)

# Use normally - all requests go through GovernanceAI
messages = [HumanMessage(content="Hello!")]
response = chat(messages)

print(response.content)
```

### LlamaIndex Integration

```python
from llama_index.llms import OpenAI

# Use GovernanceAI endpoint
llm = OpenAI(
    model="gpt-4",
    api_base="https://litellm.governanceai.com/v1",
    api_key="gak_prod_..."
)

# All requests go through GovernanceAI guardrails
response = llm.complete("What is AI governance?")
```

## Monitoring & Metrics

### View Usage

```bash
# Get LiteLLM endpoint usage
curl -H "Authorization: Bearer $API_KEY" \
  https://api.governanceai.com/v1/litellm/usage

# Returns:
{
  "total_requests": 45230,
  "total_tokens": 12453000,
  "avg_latency_ms": 245,
  "policy_violations": 123,
  "blocked_requests": 45,
  "transformed_responses": 78
}
```

### Per-Model Metrics

```bash
curl -H "Authorization: Bearer $API_KEY" \
  'https://api.governanceai.com/v1/litellm/usage/by-model'

# Returns metrics per model (gpt-4, gpt-3.5-turbo, etc.)
```

## Error Handling

GovernanceAI returns standard OpenAI error codes:

```python
try:
    response = litellm.completion(
        model="openai/gpt-4",
        messages=[...]
    )
except litellm.APIError as e:
    # Handle API errors
    print(f"Error: {e.http_status} - {e.message}")

# Common errors:
# 400 - Invalid request (malformed guardrail config)
# 401 - Authentication failed (invalid API key)
# 429 - Rate limit exceeded
# 500 - Server error (try again)
```

## Performance

### Latency Impact

GovernanceAI adds minimal latency:

* **Average overhead**: 45-100ms
* **P95**: 150ms
* **P99**: 250ms

Varies based on:

* Policy complexity
* Model response size
* Network latency to provider

### Caching

Enable response caching:

```bash
curl -X POST https://api.governanceai.com/v1/litellm/caching \
  -H "Authorization: Bearer $API_KEY" \
  -d '{
    "enabled": true,
    "ttl_seconds": 3600,
    "cache_identical_requests": true
  }'
```

## Best Practices

✅ **Do:**

* Use org\_id and user\_id in metadata
* Set appropriate rate limits
* Monitor usage regularly
* Test policies before production
* Use different keys per environment

❌ **Don't:**

* Share API keys between environments
* Disable logging for audit trails
* Route sensitive data without PII guardrails
* Forget to set up alerts

## Troubleshooting

* **401 Unauthorized** - Check API key
* **Rate limit exceeded** - Check configured limits
* **Slow responses** - Check policy complexity
* **Connection refused** - Verify endpoint URL

## Next Steps

* **[Setting Up Guardrails](/usage-guides/guardrails-setup)** - Configure policies
* **[Quick Start](../05-quick-start.mdx)** - First API call
* **[API Reference](/api)** - Full LiteLLM API docs