Deploy
AI Inference TCO Calculator
Compare API vs self-host monthly cost across low/base/high traffic — with staff and utilization inputs.
Fill in your parameters below and click Calculate. All math runs locally in your browser — nothing is sent to a server or LLM.
Traffic & tokens
API pricing
Self-host
After changing any field, click Recalculate to update results.
How we calculate
Three traffic bands (low / base / high requests per day × 30). API cost = billable input tokens × $/1M + output tokens × $/1M, with partial prefix-cache credit. Self-host GPU ≈ (tokens in + out) ÷ tokens/sec ÷ utilization × GPU $/hr × hours/month. Staff and power are flat monthly lines.
Break-even is a rough crossover when API variable cost exceeds self-host variable + fixed staff — not procurement advice.
Pricing assumptions as of 2026-08-30. Not a quote — verify vendor list prices.