Control and optimize every AI call.
Hicap gives AI teams everything they need to run AI at scale: an AI gateway with governance, control, and full observability over every request, all in one platform.
Production AI needs better inference infrastructure.
One endpoint, two products
Hicap Inference. Hicap Control Plane.
Get full visibility into every token, eliminate AI blind spots, and enforce governance across your organization. All through an AI gateway built for control and enterprise reliability.
Meet the Hicap Platform
openai/GPT-5.6 Sol
GPT · complex reasoning & agentic coding
openai/GPT-5.6 Terra
GPT · high intelligence, lower cost
openai/GPT-5.6 Luna
GPT · speed & efficiency
openai/GPT-5.5
GPT · frontier model
anthropic/claude-opus-4
Claude · reasoning model
openai/o3
GPT · deep reasoning
google/gemini-2.5-pro
Gemini · frontier model
openai/GPT-4.1
GPT · flagship model
meta/llama-4-maverick
Llama · open weights
openai/GPT-5.6 Sol
GPT · complex reasoning & agentic coding
openai/GPT-5.6 Terra
GPT · high intelligence, lower cost
openai/GPT-5.6 Luna
GPT · speed & efficiency
openai/GPT-5.5
GPT · frontier model
anthropic/claude-opus-4
Claude · reasoning model
openai/o3
GPT · deep reasoning
google/gemini-2.5-pro
Gemini · frontier model
openai/GPT-4.1
GPT · flagship model
meta/llama-4-maverick
Llama · open weights
openai/GPT-5.5
GPT · frontier model
openai/GPT-5.6 Sol
GPT · complex reasoning & agentic coding
anthropic/claude-sonnet-4
Claude · balanced model
openai/GPT-4.1-mini
GPT · fast model
openai/GPT-5.6 Terra
GPT · high intelligence, lower cost
google/gemini-2.5-flash
Gemini · fast multimodal
openai/o4-mini
GPT · fast reasoning
openai/GPT-5.6 Luna
GPT · speed & efficiency
deepseek/deepseek-r1
DeepSeek · reasoning model
openai/GPT-5.5
GPT · frontier model
openai/GPT-5.6 Sol
GPT · complex reasoning & agentic coding
anthropic/claude-sonnet-4
Claude · balanced model
openai/GPT-4.1-mini
GPT · fast model
openai/GPT-5.6 Terra
GPT · high intelligence, lower cost
google/gemini-2.5-flash
Gemini · fast multimodal
openai/o4-mini
GPT · fast reasoning
openai/GPT-5.6 Luna
GPT · speed & efficiency
deepseek/deepseek-r1
DeepSeek · reasoning model
Models
100+ models. One endpoint.
- Every major frontier and open-weight model behind a single key
- Smart routing picks the best model for cost, latency, or quality
- Add new models the day they launch, no code changes needed
Observability and control for every token.
Everything your AI does: logged, tagged, budgeted, and access-controlled — in one place.
Request log · requested → served
● live
| Time | Requested | Served | Route | Latency | Status |
|---|---|---|---|---|---|
| 14:02:11 | gpt-5.5 | gpt-5.5 | direct | 320ms | 200 |
| 14:02:11 | claude-opus-4-8 | claude-opus-4-8 | direct | 608ms | 200 |
| 14:02:12 | claude-opus-4-8 | → gemini-3.1-pro | failover · 429 | 441ms | 200 |
| 14:02:12 | gemini-3.1-flash | gemini-3.1-flash | direct | 184ms | 200 |
| 14:02:13 | auto | → mistral-large-3 | cost-optimized | 296ms | 200 |
Every reroute is logged: the model you asked for, the model that answered, and why.
Know which model actually answered.
Every request writes one log line: the model you asked for, the model that answered, and what it cost. Enable deeper analytics and tracing to see usage over the last 7 days, per team, app, or key.
- One log line per request
- Requested vs served model
- Usage analytics over the last 7 days
Model Usage
Cost, volume & tokens by model across all apps
Top Models by Credits
Spend leaderboard
#1 eleven_multilingual_v2
29,161.00
#2 eleven_v3
27,708.00
#3 scribe_v1
3,046.00
#4 scribe_v2
2,658.00
#5 claude-opus-4.7
1,191.00
Every request tagged by feature, team, app, and model.
Add tags — feature, team, app — to each API request, and Hicap records them on every log line and rolls them into spend attribution automatically. Set them once in your client; there's no separate tracking pipeline to maintain.
- Tag requests by feature, team, and app
- Every request logged with team, app, model, feature
- Drill down without extra setup
Set spend limits per team, user, or key.
Give every team, app, or key its own budget. When a limit is reached, access can be restricted — no more guessing what's allowed.
- Per-team, per-user, per-key budgets
- Configurable limits and quotas
- Restrict access when limits are reached
Application
Tooling-Devactive
Created Jun 10, 2026 · basic · demo@hicap.aiPrimary Key
RevealConnection Analytics
9.1M
Total Tokens
4.0K
Requests
40,102.00
Credits Used
7.7M
Input Tokens
2.6M
Output Tokens
962.8K
Cached Tokens
Daily Token Usage
All Time
Top API Keys Leaderboard
Highest-spending API keys across all applications
#1 Tooling-Dev
sub-mq88spd4
$401 4.0K
#2 Tooling-Prod
sub-mq88soii
$289 1.0K
Every key and model, under one policy.
Give each team scoped keys and a list of approved models, and every request they make is checked against it. You always know which apps run, on which models, with whose keys.
- Connection-scoped, rotatable keys
- Model allow-lists per team
- Full audit trail, exportable
One line to every model. Infrastructure built for inference.
Point your OpenAI-compatible client at api.hicap.ai — no rewrites, no new SDK. Every request rides a distributed gateway: load-balanced across providers and regions, health-tracked, and rerouted mid-flight when a provider fails. Your app stays up.
Routing strategies
1curl https://api.hicap.ai/v1/chat/completions \2 -H "api-key: $HICAP_API_KEY" \3 -H "Content-Type: application/json" \4 -d '{5 "model": "gpt-5.5",6 "messages": [7 { "role": "user", "content": "Hello" }8 ]9 }'
Every token, cheaper.
Hicap pools bulk reserved capacity so you can lock in throughput at up to 25% below standard pay-as-you-go. Spikes overflow to on-demand at the provider list price with no Hicap markup. Whatever you reserve but don't use sells on the spot market, and the revenue returns to you.
Reserved capacity
MonitoringListed38k tokens on the spot market
$0.62/tokWorks with your stack.
One endpoint in front of every major provider and tool. Keep your stack exactly as it is — no framework lock-in.
Drop-in access
Integrations
Model providers
100+ compatibilities
Secure by default. Sovereign from day one.
Enterprise-grade controls, on by default — not bolted on later.
Private networking
Direct model inference that bypasses the public internet
No data retention
Prompts, responses, and outputs are never stored — your data stays in the request path.
Data residency
Tenant isolation and residency controls per deployment
Role-based access
Permissions scoped down to the connection and key
The one-line migration
Bring order to your AI stack in one line of code.
Change one base URL — every model, provider, and dollar runs through one place.

