Sample scenario — You're viewing Gauge as Acme Legal Tech, an AI-native legal SaaS company. The margin matrix shows Acme's own law firm clients — their MRR flows in via their billing platform and AI costs via LiteLLM. Gauge tracks AI margin per customer in real time, creates alerts when margins are drifting, and automatically blocks or downgrades requests when a customer's margin drops below their configured floor.
Margin Matrix
Per-customer gross margin · sample data
Last 30 days ▾
Export ↗
Demo data
Total MRR
$60,000
AI Cost / mo
$14,680
AI Unit Margin
75.5%
Saved vs GPT-4o
$32,500
Margin Breaches
1
active · auto-enforcing
Davis Polk is margin-negative — 3 days running
AI cost $2,480 vs MRR $2,000 · –24% AI margin · Gauge is auto-enforcing downgrade · Recommended: upgrade to Growth plan ($4,800/mo)
Weil Gotshal AI cost spiked 34% this week
89% driven by legal_reasoning (held on GPT-4o) · margin still healthy at 69.7% but trending down · Consider usage cap on legal_reasoning for this tier
Caching active
1 of 6 customers
Latham & Watkins · 34% hit rate
Cache hits this month
84,200
requests served at zero inference cost
Saved from caching
$420/mo
across all customers this month
Cache opportunity
$600/mo
uncaptured · 2 customers flagged ↓
Customer margin matrix Configure floors →
Customer MRR AI cost / mo AI Margin Model Insight Actions
Kirkland & Ellis
Enterprise
$24,000 $5,820 +75.8% Llama 8B ⬡ Cache opportunity
62% of clause_lookup requests are repeat queries · caching could save ~$420/mo · margin +7pp
Latham & Watkins
Enterprise
$18,000 $3,210 +82.2% Llama 8B Healthy · cache already active · 34% hit rate
Sullivan & Cromwell
Growth
$9,600 $2,940 +69.4% Mixed Margin drifting ↓
⬡ Caching contract_summary could recover ~$180/mo
Weil Gotshal
Growth
$6,000 $1,820 +69.7% Llama 8B Cost spike +34% · legal_reasoning heavy
Davis Polk ⚠
Starter
$2,000 $2,480 –24.0% GPT-4o ↓ Margin-negative · auto-enforcing
Skadden Arps
Growth
$4,800 $890 +81.5% Llama 8B Healthy · no action needed
Davis Polk breached –10% floor at 2:14am · downgrading remaining requests to Llama-3.1-8B via Fireworks
Auto-enforcing
Spend by model this month
AI margin trend · last 60 days
Requests routed
284,120
to cheaper models today
Avg quality score
0.87
threshold: 0.82
Cost per 1k tokens
$0.22
blended · down from $2.50
SLA breaches
0
p50 latency: 340ms
Live routing decisions ● streaming
Inference provider matrix · Llama 3.1 8B
ProviderListed $/1MInfra costEffective $/1Mp50
Fireworks AI
managed inference
$0.20 bundled $0.20 310ms ✓ active
Baseten
managed inference
$0.28 bundled $0.28 380ms standby
RunPod (self-hosted)
$3,100/mo lease · 40% util
$0.14 $3,100/mo $0.31 ↑ 680ms high eff. cost
OpenAI GPT-4o
frontier model
$2.50 bundled $2.50 290ms 12× cost
// effective cost = listed token price + (monthly lease ÷ utilization ÷ projected tokens) · recalculated every 60s
Per-customer routing table
Routed cheaper
Shadow testing
Held on frontier
Task type Customer Status Samples Score Floor Routes to Saving/req
// re-evaluated every 500 requests per task/customer pair
GPU crossover engine · auto-calculated from your live data
Pulling from LiteLLM · RunPod · OpenAI billing…
···
calculating…
Fetching your token volume and lease rates…
Current API spend
per month · OpenAI
H100 lease cost
2× H100 SXM5 · RunPod
Month 3 savings
per month after crossover
Cost trajectory · 90 days
API cost
Self-host
Cost vs self-host over 90 days.
from LiteLLM
Token volume
from RunPod
GPU lease
from usage trend
Growth rate
// recalculates every 24h as your data updates
Owned hardware scenario Two-stage crossover
vs leased GPU · vs managed inference
What Gauge calculates
Crossover 1 · lease beats API
Day 14
RunPod H100 lease beats Fireworks per-token cost
Crossover 2 · owned beats lease
Month 8
Owned cluster amortized capex beats RunPod lease
Owned cluster assumptions
Hardware 4× Mac Studio Ultra
Total capex $36,000
Amortization 36 months
Monthly cost $1,000/mo
Effective $/1M tokens $0.08/1M
Inference runtime exo · vLLM
// compute cost only · does not include ops overhead or procurement time
At Month 8, owned hardware becomes your lowest-cost tier. Gauge routes non-latency-sensitive tasks there first, spills to leased GPU when saturated, and reserves managed inference for burst and frontier-only tasks.
Compute resale offset Coming soon
via Vast.ai · RunPod marketplace
Off-peak window
2am – 8am
est. from your usage curve
Resaleable capacity
~$800/mo
at current Vast.ai spot rates
Adjusted crossover
Day 9
vs Day 14 without resale
Gauge will automatically list spare capacity when utilization drops below your threshold, and reclaim it when your traffic picks up. Your cluster becomes partially self-funding.
How this connects — Gauge installs as a GitHub App on your repository. Every time a PR is opened, Gauge runs the prompt diff against a sample of your real production traffic and posts the cost projection as a comment before the PR can merge. Setup takes about 10 minutes. Connect GitHub →
Blocked this month
3
high-cost PRs blocked
Cost avoided
$4,280
per month
Pending review
2
awaiting decision
PRs approved
18
cost within threshold
PR #341 · feat/rag-context-expansion BLOCKED
G
gauge-bot · just now
⚠ GAUGE COST ALERT
Context window: +350 tokens / request
Anthropic bill: +$1,450 / month (▲ 18.4%)
Tier-1 margin impact: –4.2 pp
Agentic loop risk: elevated
✓ Approve
⊘ Block
Simulation based on 847,000 sampled requests (last 7d)
PR #338 · fix/prompt-compression APPROVED
G
gauge-bot · 2h ago
✓ GAUGE COST APPROVAL
Context window: –120 tokens / request
Anthropic bill: –$520 / month (▼ 6.6%)
Tier-1 margin impact: +1.4 pp
Quality delta: <0.01 — negligible
✓ Auto-approved
⊘ Block
Simulation based on 847,000 sampled requests (last 7d)
PR #335 · feat/gpt4o-upgrade-all BLOCKED
G
gauge-bot · 1d ago
⚠ GAUGE COST ALERT
Model change: Llama 8B → GPT-4o (all routes)
Monthly impact: +$31,200 / month (▲ 178%)
Margin impact: –38 pp across all tiers
Quality gain: +0.09 — not worth the cost
✓ Approve anyway
⊘ Keep blocked
Simulation based on 847,000 sampled requests (last 7d)
PR #333 · refactor/cache-layer APPROVED
G
gauge-bot · 3d ago
✓ GAUGE COST APPROVAL
Cache hit rate: +22% improvement
Token reduction: –18% of total requests
Monthly savings: $2,840 / month
Quality delta: none — cache only
✓ Auto-approved
⊘ Block
Simulation based on 847,000 sampled requests (last 7d)
Connected
0 / 3
connect to go live
Setup time
~30 min
estimated total
SDK required
None
webhooks only
Code changes
None
on your side
Required connections View docs →
Gateway
LiteLLM · Bifrost · OpenRouter — one webhook covers every model you're running
Add Gauge's webhook URL to your gateway config. Every completion event starts flowing automatically.
Not connected
💳
Revenue / Billing Platform
Gauge reads MRR per customer from your billing system — whichever one you use. Connect one below.
Without a billing connection, Gauge can track AI cost but can't show per-customer margin.
Not connected
🧾
AI Provider Billing
OpenAI · Anthropic · Google · AWS Bedrock · Azure — import your billing history and keep it synced
Connect whichever providers you use. Gauge computes blended effective cost across all of them.
Not connected
GitHub / GitLab optional
Installs gauge-bot as a GitHub App — posts cost impact comments on every PR before it merges
Required for PR cost guardrails. Not needed for margin matrix or routing.
Optional
Infrastructure billing optional · for GPU crossover engine
RunPod · Lambda Labs · CoreWeave GPU lease costs · utilization data
AWS · GCP · Azure reserved capacity · spot pricing