Davis Polk is margin-negative — 3 days running
AI cost $2,480 vs MRR $2,000 · –24% AI margin · Gauge is auto-enforcing downgrade · Recommended: upgrade to Growth plan ($4,800/mo)
Weil Gotshal AI cost spiked 34% this week
89% driven by legal_reasoning (held on GPT-4o) · margin still healthy at 69.7% but trending down · Consider usage cap on legal_reasoning for this tier
Caching active
1 of 6 customers
Latham & Watkins · 34% hit rate
Cache hits this month
84,200
requests served at zero inference cost
Saved from caching
$420/mo
across all customers this month
Cache opportunity
$600/mo
uncaptured · 2 customers flagged ↓
Customer margin matrix
Configure floors →
| Customer | MRR | AI cost / mo | AI Margin | Model | Insight | Actions |
|---|---|---|---|---|---|---|
Kirkland & Ellis Enterprise |
$24,000 | $5,820 | +75.8% | Llama 8B | ⬡ Cache opportunity 62% of clause_lookup requests are repeat queries · caching could save ~$420/mo · margin +7pp |
|
Latham & Watkins Enterprise |
$18,000 | $3,210 | +82.2% | Llama 8B | Healthy · cache already active · 34% hit rate | |
Sullivan & Cromwell Growth |
$9,600 | $2,940 | +69.4% | Mixed | Margin drifting ↓ ⬡ Caching contract_summary could recover ~$180/mo |
|
Weil Gotshal Growth |
$6,000 | $1,820 | +69.7% | Llama 8B | Cost spike +34% · legal_reasoning heavy | |
Davis Polk ⚠ Starter |
$2,000 | $2,480 | –24.0% | GPT-4o ↓ | Margin-negative · auto-enforcing | |
Skadden Arps Growth |
$4,800 | $890 | +81.5% | Llama 8B | Healthy · no action needed |
Spend by model this month
AI margin trend · last 60 days
Requests routed
284,120
to cheaper models today
Avg quality score
0.87
threshold: 0.82
Cost per 1k tokens
$0.22
blended · down from $2.50
SLA breaches
0
p50 latency: 340ms
Live routing decisions
● streaming
Inference provider matrix · Llama 3.1 8B
| Provider | Listed $/1M | Infra cost | Effective $/1M | p50 | |
|---|---|---|---|---|---|
|
Fireworks AI
managed inference
|
$0.20 | bundled | $0.20 | 310ms | ✓ active |
|
Baseten
managed inference
|
$0.28 | bundled | $0.28 | 380ms | standby |
|
RunPod (self-hosted)
$3,100/mo lease · 40% util
|
$0.14 | $3,100/mo | $0.31 ↑ | 680ms | high eff. cost |
|
OpenAI GPT-4o
frontier model
|
$2.50 | bundled | $2.50 | 290ms | 12× cost |
// effective cost = listed token price + (monthly lease ÷ utilization ÷ projected tokens) · recalculated every 60s
Per-customer routing table
Routed cheaper
Shadow testing
Held on frontier
| Task type | Customer | Status | Samples | Score | Floor | Routes to | Saving/req |
|---|
// re-evaluated every 500 requests per task/customer pair
GPU crossover engine · auto-calculated from your live data
Pulling from LiteLLM · RunPod · OpenAI billing…
···
calculating…
Fetching your token volume and lease rates…
Current API spend
—
per month · OpenAI
H100 lease cost
—
2× H100 SXM5 · RunPod
Month 3 savings
—
per month after crossover
Cost trajectory · 90 days
API cost
Self-host
from LiteLLM
Token volume
—
from RunPod
GPU lease
—
from usage trend
Growth rate
—
How this connects — Gauge installs as a GitHub App on your repository. Every time a PR is opened, Gauge runs the prompt diff against a sample of your real production traffic and posts the cost projection as a comment before the PR can merge. Setup takes about 10 minutes. Connect GitHub →
Blocked this month
3
high-cost PRs blocked
Cost avoided
$4,280
per month
Pending review
2
awaiting decision
PRs approved
18
cost within threshold
PR #341 · feat/rag-context-expansion
BLOCKED
G
⚠ GAUGE COST ALERT
•Context window: +350 tokens / request
•Anthropic bill: +$1,450 / month (▲ 18.4%)
•Tier-1 margin impact: –4.2 pp
•Agentic loop risk: elevated
✓ Approve
⊘ Block
Simulation based on 847,000 sampled requests (last 7d)
PR #338 · fix/prompt-compression
APPROVED
G
✓ GAUGE COST APPROVAL
•Context window: –120 tokens / request
•Anthropic bill: –$520 / month (▼ 6.6%)
•Tier-1 margin impact: +1.4 pp
•Quality delta: <0.01 — negligible
✓ Auto-approved
⊘ Block
Simulation based on 847,000 sampled requests (last 7d)
PR #335 · feat/gpt4o-upgrade-all
BLOCKED
G
⚠ GAUGE COST ALERT
•Model change: Llama 8B → GPT-4o (all routes)
•Monthly impact: +$31,200 / month (▲ 178%)
•Margin impact: –38 pp across all tiers
•Quality gain: +0.09 — not worth the cost
✓ Approve anyway
⊘ Keep blocked
Simulation based on 847,000 sampled requests (last 7d)
PR #333 · refactor/cache-layer
APPROVED
G
✓ GAUGE COST APPROVAL
•Cache hit rate: +22% improvement
•Token reduction: –18% of total requests
•Monthly savings: $2,840 / month
•Quality delta: none — cache only
✓ Auto-approved
⊘ Block
Simulation based on 847,000 sampled requests (last 7d)
Connected
0 / 3
connect to go live
Setup time
~30 min
estimated total
SDK required
None
webhooks only
Code changes
None
on your side
Required connections
View docs →
⚡
Gateway
LiteLLM · Bifrost · OpenRouter — one webhook covers every model you're running
Add Gauge's webhook URL to your gateway config. Every completion event starts flowing automatically.
Not connected
💳
Revenue / Billing Platform
Gauge reads MRR per customer from your billing system — whichever one you use. Connect one below.
Without a billing connection, Gauge can track AI cost but can't show per-customer margin.
Not connected
🧾
AI Provider Billing
OpenAI · Anthropic · Google · AWS Bedrock · Azure — import your billing history and keep it synced
Connect whichever providers you use. Gauge computes blended effective cost across all of them.
Not connected
⑂
GitHub / GitLab optional
Installs gauge-bot as a GitHub App — posts cost impact comments on every PR before it merges
Required for PR cost guardrails. Not needed for margin matrix or routing.
Optional
Infrastructure billing optional · for GPU crossover engine
RunPod · Lambda Labs · CoreWeave GPU lease costs · utilization data
AWS · GCP · Azure reserved capacity · spot pricing