Analytics
Control
Privacy

AI Gateway for Complete Control

The intelligent gateway for AI infrastructure with budget controls, smart failover, and comprehensive analytics. Perfect for developers and enterprises scaling AI applications while maintaining privacy and cost efficiency.

We'll never share your email. No spam.

dashboard.forstream.ai
Analytics Dashboard

Complete AI Infrastructure Control

Manage every AI API call with intelligent routing, real-time cost monitoring, and enterprise-grade security. Route across OpenAI, Anthropic, and open-source models through a single API key with advanced policies and analytics.

AI GatewayCost ControlPrivacy FirstSmart Routing
Real-time Usage
Cost: $142.35Tokens: 2.4MModels: 8
GPT-4: 45%
Claude: 35%

This month

$4,287.12

Avg latency

234ms

Success rate

99.7%

OpenAI
$2,184.56
Anthropic
$1,542.87

AI Usage Analytics & Cost Tracking

Cost, usage, latency, and error breakdowns by model/provider/key. Export to your BI/observability stack.

Request Flow

Your App
Forstream Router
GPT-4
Claude
Llama
Auto-failover in 200ms

Smart Routing & Failover

Route around outages automatically. A/B test models without touching code. Policies and cost-based routing.

API Key Policies

Team Alpha
$2,450/$5,000

Monthly spend limit: $5,000

Customer Key
847/1000 req/min

Models: GPT-4, Claude-3.5 only

Expired Key
BLOCKED

Exceeded monthly limit

Granular Access Control

Per-key, per-team, per-model budgets and restrictions. Time-boxed keys, rate limits, and scoped access.

Supported Providers

Forstream API
GPT

OpenAI

CL

Claude

OSS

Open Source

Route to best available model

One API, Many Models

Route to OpenAI, Anthropic, and leading OSS models like Kimi K2, Qwen3, GPT-OSS. Drop-in OpenAI-compatible API.

Zero Prompt Retention

Request Processed
SECURE

Prompt: [NEVER STORED]

Response: [NEVER STORED]

Metadata Only
Tokens2,847
Model IDgpt-4
Timestamp14:32:15
Status200
GDPR Compliant

Privacy by Default

We never store prompts or completions. Logs contain only metadata — token counts, model IDs, timestamps, status codes.

OpenAI Compatible
Smart routing
Privacy-first
Usage analytics
Multi-provider
Rate limiting
Cost tracking

See it in action

Experience AI inference with control, insight, and privacy through interactive examples

Real-time Analytics

Usage Dashboard Screenshot
Cost Breakdown Charts Screenshot
Performance Metrics Screenshot

AI Models

Discover the latest and most powerful AI models from leading providers. From text generation to image creation, find the perfect model for your needs.

M

Llama 2 70B

Meta

Open Source

Open source model with strong performance on reasoning and code tasks, completely free to use

Key Capabilities

Text Generation
Code
Reasoning
Pricing per 1M tokens
Input
Free
Output
Free
M

Kimi K2

Moonshot AI

Open Source

Advanced open source model with excellent long-context understanding and multilingual capabilities

Key Capabilities

Text Generation
Long Context
Multilingual
Reasoning
Pricing per 1M tokens
Input
$0.20
Output
$0.80
A

Qwen3 235B

Alibaba

Open Source

Large-scale open source model with exceptional performance across diverse tasks and languages

Key Capabilities

Text Generation
Code
Multilingual
Analysis
Pricing per 1M tokens
Input
$22.00
Output
$22.00
A

Qwen3 Coder 480B

Alibaba

Open Source

Specialized coding model with massive parameters for complex programming tasks

Key Capabilities

Code Generation
Programming
Debug
Architecture
Pricing per 1M tokens
Input
$35.00
Output
$35.00
O

GPT-5

OpenAI Community

Open Source

Next-generation open source model with advanced reasoning and multimodal capabilities

Key Capabilities

Text Generation
Reasoning
Code
Multimodal
Pricing per 1M tokens
Input
Free
Output
Free
O

GPT-OSS 120B

Open Source Community

Open Source

Community-driven open source GPT model with strong performance and transparency

Key Capabilities

Text Generation
Code
Analysis
Creative Writing
Pricing per 1M tokens
Input
$0.10
Output
$0.50
D

DeepSeek V3

DeepSeek

Open Source

Advanced open source model optimized for reasoning, mathematics, and code generation

Key Capabilities

Reasoning
Mathematics
Code
Analysis
Pricing per 1M tokens
Input
$0.10
Output
$0.20
Z

GLM 4.5 Air

Zhipu AI

Open Source

Lightweight yet powerful open source model with excellent efficiency and performance

Key Capabilities

Text Generation
Code
Efficiency
Multilingual
Pricing per 1M tokens
Input
Free
Output
Free
D

DeepSeek R1

DeepSeek

Open Source

Latest reasoning-focused model with advanced chain-of-thought capabilities and superior performance

Key Capabilities

Reasoning
Chain-of-Thought
Problem Solving
Analysis
Pricing per 1M tokens
Input
$0.50
Output
$2.00

This Could Be You

Real scenarios from teams who need control, insight, and privacy for their AI infrastructure.

Control Your AI Costs

"One customer used $2,000 of AI credits in an hour"

As a SaaS founder, unpredictable AI costs can destroy your margins overnight. Set granular spending limits and never worry about runaway usage again.

Peace of mind + predictable costs

Per-Customer Budgets

Set spending caps for individual customers or entire customer tiers

Rate Limiting

Prevent abuse with customizable rate limits per API key

Real-time Alerts

Get notified before costs spiral out of control

One API. Full control. Complete privacy.

Drop-in. OpenAI-compatible. Routing built-in.

Call our /v1/chat/completions endpoint just like OpenAI's — but with budgets, analytics, and smart routing on every request.

Same API, More Control

OpenAI-Compatible

Drop-in replacement for OpenAI's API. No code changes required.

Smart Routing

Automatically route to the best available provider based on your policies.

Built-in Analytics

Every request tracked with costs, latency, and error rates.

Drop In Compatibility
curl -X POST https://proxy.forstream.ai/v1/chat/completions \
  -H “Content-Type: application/json” \
  -H “Authorization: Bearer YOUR_API_KEY” \
  -d '{
    “model”: “gpt-4”,
    “messages”: [
      {“role”: “user”, “content”: “Hello world”}
    ],
    “max_tokens”: 100
  }”'
Terminal
# Policy Configuration
name: “saas-production”
version: “1.0”

# Budget Controls
budgets:
  - name: “per-customer”
    limit: “$50/month”
    scope: “api_key”
  - name: “team-limit”
    limit: “$500/month”
    scope: “organization”

# Smart Routing
routing:
  primary: “openai”
  fallback: [“anthropic”, “cohere”]
  health_check: true
  timeout: “30s”

# Privacy Settings
privacy:
  log_prompts: false
  log_completions: false
  metadata_only: true
  retention: “90d”
policy.yaml

Policy-Driven Control

Flexible Budgets

Set spending limits per customer, team, or entire organization.

Smart Failover

Define primary and backup providers with automatic health checks.

Privacy by Default

Zero prompt logging with configurable data retention policies.

GitOps ready

Frequently Asked Questions

Everything you need to know about AI infrastructure with control, insight, and privacy

While you could call OpenAI directly, you'd miss out on critical business controls. We provide granular budgets, automatic failover to backup providers, real-time cost analytics, and privacy-by-default logging. It's like having a complete AI ops layer that scales with your business without changing your code.

No, never. We use a metadata-only approach - we track token counts, model usage, response times, and costs, but your actual prompts and completions never touch our servers or logs. This means you get full visibility without compromising privacy.

We support all major providers including OpenAI, Anthropic, Cohere, together.ai, and leading open-source models. Our smart routing automatically switches between providers based on availability, cost, or performance - you just call our OpenAI-compatible endpoint.

You can set granular budgets per API key, per customer, per team, or across your entire organization. When limits are reached, requests are automatically blocked or routed to cheaper alternatives. Get real-time alerts before you hit your caps, so there are never any surprises.

It's much more than a proxy. We provide intelligent routing, cost optimization, automatic failover, detailed analytics, compliance-ready logging, and policy enforcement. Think of it as your AI infrastructure control plane that saves you money while reducing operational complexity.

You can be up and running in under 5 minutes. Simply swap your OpenAI API endpoint for ours, add your API key, and optionally configure policies. No code changes required - it's a true drop-in replacement with immediate benefits.

Still have questions? We're here to help.

Early Access

Route smarter, see everything

Early access to AI inference with control, insight, and privacy. Join engineering leaders building the future of AI infrastructure.

Zero vendor lock-in
Privacy by design
Analytics

Rolling invites for early users. We'll never share your email. No spam.