Skip to main content

Control and optimize every AI call.

Hicap gives AI teams everything they need to run AI at scale: an AI gateway with governance, control, and full observability over every request, all in one platform.

Production AI needs better inference infrastructure.

One endpoint, two products

Hicap Inference. Hicap Control Plane.

Get full visibility into every token, eliminate AI blind spots, and enforce governance across your organization. All through an AI gateway built for control and enterprise reliability.

Meet the Hicap Platform

api.hicap.ai/v1

openai/GPT-5.6 Sol

GPT · complex reasoning & agentic coding

openai/GPT-5.6 Terra

GPT · high intelligence, lower cost

openai/GPT-5.6 Luna

GPT · speed & efficiency

openai/GPT-5.5

GPT · frontier model

anthropic/claude-opus-4

Claude · reasoning model

openai/o3

GPT · deep reasoning

google/gemini-2.5-pro

Gemini · frontier model

openai/GPT-4.1

GPT · flagship model

meta/llama-4-maverick

Llama · open weights

openai/GPT-5.6 Sol

GPT · complex reasoning & agentic coding

openai/GPT-5.6 Terra

GPT · high intelligence, lower cost

openai/GPT-5.6 Luna

GPT · speed & efficiency

openai/GPT-5.5

GPT · frontier model

anthropic/claude-opus-4

Claude · reasoning model

openai/o3

GPT · deep reasoning

google/gemini-2.5-pro

Gemini · frontier model

openai/GPT-4.1

GPT · flagship model

meta/llama-4-maverick

Llama · open weights

openai/GPT-5.5

GPT · frontier model

openai/GPT-5.6 Sol

GPT · complex reasoning & agentic coding

anthropic/claude-sonnet-4

Claude · balanced model

openai/GPT-4.1-mini

GPT · fast model

openai/GPT-5.6 Terra

GPT · high intelligence, lower cost

google/gemini-2.5-flash

Gemini · fast multimodal

openai/o4-mini

GPT · fast reasoning

openai/GPT-5.6 Luna

GPT · speed & efficiency

deepseek/deepseek-r1

DeepSeek · reasoning model

openai/GPT-5.5

GPT · frontier model

openai/GPT-5.6 Sol

GPT · complex reasoning & agentic coding

anthropic/claude-sonnet-4

Claude · balanced model

openai/GPT-4.1-mini

GPT · fast model

openai/GPT-5.6 Terra

GPT · high intelligence, lower cost

google/gemini-2.5-flash

Gemini · fast multimodal

openai/o4-mini

GPT · fast reasoning

openai/GPT-5.6 Luna

GPT · speed & efficiency

deepseek/deepseek-r1

DeepSeek · reasoning model

Models

100+ models. One endpoint.

  • Every major frontier and open-weight model behind a single key
  • Smart routing picks the best model for cost, latency, or quality
  • Add new models the day they launch, no code changes needed

Observability and control for every token.

Everything your AI does: logged, tagged, budgeted, and access-controlled — in one place.

Request log · requested → served

● live

TimeRequestedServedRouteLatencyStatus
14:02:11gpt-5.5gpt-5.5direct320ms200
14:02:11claude-opus-4-8claude-opus-4-8direct608ms200
14:02:12claude-opus-4-8gemini-3.1-profailover · 429441ms200
14:02:12gemini-3.1-flashgemini-3.1-flashdirect184ms200
14:02:13automistral-large-3cost-optimized296ms200

Every reroute is logged: the model you asked for, the model that answered, and why.

Know which model actually answered.

Every request writes one log line: the model you asked for, the model that answered, and what it cost. Enable deeper analytics and tracing to see usage over the last 7 days, per team, app, or key.

  • One log line per request
  • Requested vs served model
  • Usage analytics over the last 7 days

Model Usage

Cost, volume & tokens by model across all apps

30 DaysAll Time
gpt-5.3-codex
1.30M
gpt-5.3-chat
1.24M
gpt-5.2
1.12M
claude-opus-4.7
1.10M
claude-opus-4.8
1.05M
claude-sonnet-4.6
1.04M
gemini-3-pro-preview
1.02M
0350K700K1.05M1.4M

Top Models by Credits

Spend leaderboard

#1 eleven_multilingual_v2

29,161.00

#2 eleven_v3

27,708.00

#3 scribe_v1

3,046.00

#4 scribe_v2

2,658.00

#5 claude-opus-4.7

1,191.00

Every request tagged by feature, team, app, and model.

Add tags — feature, team, app — to each API request, and Hicap records them on every log line and rolls them into spend attribution automatically. Set them once in your client; there's no separate tracking pipeline to maintain.

  • Tag requests by feature, team, and app
  • Every request logged with team, app, model, feature
  • Drill down without extra setup

Set spend limits per team, user, or key.

Give every team, app, or key its own budget. When a limit is reached, access can be restricted — no more guessing what's allowed.

  • Per-team, per-user, per-key budgets
  • Configurable limits and quotas
  • Restrict access when limits are reached
platform.hicap.ai/keys

Application

All Applications

Tooling-Devactive

Created Jun 10, 2026 · basic · demo@hicap.ai

Primary Key

Reveal
••••••••••••••••••••••••••••••••••••

Connection Analytics

30 DaysAll Time

9.1M

Total Tokens

4.0K

Requests

40,102.00

Credits Used

7.7M

Input Tokens

2.6M

Output Tokens

962.8K

Cached Tokens

Daily Token Usage

All Time

Apr 13May 4May 27Jun 17Jul 8

Top API Keys Leaderboard

Highest-spending API keys across all applications

#1 Tooling-Dev

sub-mq88spd4

$401 4.0K

#2 Tooling-Prod

sub-mq88soii

$289 1.0K

Every key and model, under one policy.

Give each team scoped keys and a list of approved models, and every request they make is checked against it. You always know which apps run, on which models, with whose keys.

  • Connection-scoped, rotatable keys
  • Model allow-lists per team
  • Full audit trail, exportable

One line to every model. Infrastructure built for inference.

Point your OpenAI-compatible client at api.hicap.ai — no rewrites, no new SDK. Every request rides a distributed gateway: load-balanced across providers and regions, health-tracked, and rerouted mid-flight when a provider fails. Your app stays up.

Routing strategies

BalancedLatency optimizedCost optimizedData residentReserved capacity first
1curl https://api.hicap.ai/v1/chat/completions \
2 -H "api-key: $HICAP_API_KEY" \
3 -H "Content-Type: application/json" \
4 -d '{
5 "model": "gpt-5.5",
6 "messages": [
7 { "role": "user", "content": "Hello" }
8 ]
9 }'

Every token, cheaper.

Hicap pools bulk reserved capacity so you can lock in throughput at up to 25% below standard pay-as-you-go. Spikes overflow to on-demand at the provider list price with no Hicap markup. Whatever you reserve but don't use sells on the spot market, and the revenue returns to you.

Reserved capacity

Monitoring
100k tokens/mokey prod-inference
62% used internally38% idle → listed

Listed38k tokens on the spot market

$0.62/tok
Cleared: 50k tokens demand38% below list
Platform fee (5%)$1.18

Works with your stack.

One endpoint in front of every major provider and tool. Keep your stack exactly as it is — no framework lock-in.

Drop-in access

OpenAI-compatible APIPython SDKTypeScript SDKREST / cURLBYOK

Integrations

ClineOpenClawAiderContinueOpenHandsClaude Code RouterLiteLLMVS Codeand many more...

Model providers

OpenAIAnthropicGoogle GeminiDeepSeekKimiZhipuElevenLabsMiniMaxand many more...

100+ compatibilities

LiteLLMOpenClawOpenClaudeClaude Code RouterClineAiderContinueRoo CodeOpenHandsCursorOpenAI Agents SDKcurl
LangChainLlamaIndexLangGraphVercel AI SDKPydantic AICrewAIAutoGenn8nDifyLangflowOllamaHugging FaceMLflowPromptfoo

Secure by default. Sovereign from day one.

Enterprise-grade controls, on by default — not bolted on later.

SOC 2Type ICertifiedISO27001In progress

Private networking

Direct model inference that bypasses the public internet

No data retention

Prompts, responses, and outputs are never stored — your data stays in the request path.

Data residency

Tenant isolation and residency controls per deployment

Role-based access

Permissions scoped down to the connection and key

The one-line migration

Bring order to your AI stack in one line of code.

Change one base URL — every model, provider, and dollar runs through one place.