A planning reference, not a price list
What custom AI and infrastructure hosting actually costs to run.
We don't sell fixed packages, which means we also don't publish fixed prices. What we can do is show you the shape of the number: what compute, data, bandwidth, and management actually cost at current market rates, broken down by the variables that move your bill the most.
What actually moves the number
Eight variables account for nearly all of the spread.
Two companies running what looks like "the same" AI product can land 10x apart on monthly spend. These are the variables that explain why.
Concurrent users and API calls per month set your baseline compute and bandwidth floor before anything else is decided.
A 7B parameter model and a 70B+ model are different hosting problems entirely, and the GPU class you need follows directly from that choice.
Always-on dedicated capacity costs more per hour than autoscaled or serverless inference, but serverless costs more per request at sustained volume.
RAG and semantic search costs scale with the number of stored vectors and their dimensionality, not just with traffic.
Media-heavy products move far more data out of the cloud than text-only products, and egress pricing is tiered and provider-specific.
Single-region hosting and multi-region failover are not the same line item twice. Failover roughly doubles the infrastructure it protects.
Formal compliance programs and data residency requirements add dedicated infrastructure and audit overhead most teams underestimate.
Business-hours monitoring and 24/7 SRE-level management are priced very differently, and most outages happen outside business hours.
Compute reference
GPU cloud rates, current market range.
Inference and fine-tuning workloads are priced by the GPU-hour. Rates vary by provider tier, from budget neo-clouds to major hyperscalers, and by commitment level.
| GPU Class | Typical Use | Hourly Range | Full-Time Monthly (single GPU) |
|---|---|---|---|
| A100 (80GB) | Mid-scale training, fine-tuning, moderate inference | $1.25 to $2.50 | $900 to $1,800 |
| H100 (80GB) | Production inference, 30B to 70B class models | $1.50 to $7.00 | $1,100 to $5,000 |
| H200 (141GB) | Memory-bound inference, larger context windows | $0.50 to $4.00 | $400 to $2,900 |
| B200 (latest gen) | Frontier-scale training and high-throughput inference | $5.00 to $18.00 | $3,600 to $13,000 |
Multi-GPU clusters (4 to 8 GPUs, the realistic floor for serious training or high-availability inference) typically run $6,000 to $40,000 per month depending on GPU class, provider, and reservation term.
Data layer reference
Vector and RAG storage, by scale.
Retrieval-augmented systems price by stored vectors, query volume, and embedding dimensionality. Cost is rarely linear. It steps up hard at certain scale thresholds.
| Index Size | Typical Workload | Monthly Range |
|---|---|---|
| Under 1M vectors | Pilot, early product, small knowledge base | $0 to $75 |
| 1M to 10M vectors | Production RAG for a single product line | $65 to $400 |
| 10M to 100M vectors | Multi-tenant RAG, growing document corpus | $400 to $2,000 |
| 100M+ vectors | Enterprise-scale search, agent memory at scale | $2,000 to $10,000+ |
Self-hosting the vector layer can undercut managed platforms by 3x to 10x at high query volume, at the cost of taking on the operational overhead directly. That trade-off is one of the first things we model in an actual engagement.
Delivery reference
Bandwidth and egress, by provider type.
Egress, data leaving the cloud toward your users, is billed per gigabyte and tiered by volume. A small number of providers charge nothing for it at all.
| Provider Type | Egress Rate | 10TB/Month Estimate |
|---|---|---|
| Major hyperscaler (standard tier) | $0.07 to $0.12/GB | $700 to $1,200 |
| Major hyperscaler (high volume, negotiated) | $0.02 to $0.05/GB | $200 to $500 |
| CDN edge delivery (cache hit) | $0.004 to $0.02/GB | $40 to $200 |
| Zero-egress object storage | $0.00/GB | $0, storage billed separately |
Which of these actually applies depends on how much of your traffic is cache hits versus origin pulls, and whether your architecture can take advantage of zero-egress storage at all.
Illustrative spend tiers
Four shapes of the same question: what will this actually cost.
These are not packages we sell. They are composite ranges built from the reference tables above, illustrating how total monthly spend moves as traffic, redundancy, and compliance requirements scale up. Your actual number depends on your actual architecture.