How to size GPUs and AI inference costs
Context, concurrency, queueing and latency: measurement criteria for planning an AI product's infrastructure and costs.
Assess memory beyond model weights
The memory needed to load a model is only the first limit. Operations also need room for cache and concurrent work. Long contexts and more active requests change that budget.
Capacity per GPU depends on the product's model, configuration and input distribution. A user count without these conditions is not a useful sizing reference.
Prepare a representative workload
- Record input and output sizes, peak periods and expected volume.
- Vary concurrency and observe when the queue grows without recovering.
- Measure time to first token, total time and errors; inspect the distribution as well as the average.
- State whether latency includes queueing, network time and input preparation.
Calculate the cost per delivered result
Divide the period's cost by work completed within the quality and latency targets. Include idle time, storage, transfer and operational effort that are part of the contract.
Compare owned infrastructure, rented infrastructure and a managed API using the same workload. The lowest hourly price can still mean a higher cost per useful result.
Assess GPU sharing
Switching between LLMs, transcription and speech synthesis can allow shared hardware use. However, loading models takes time and can affect waiting requests. Measure that impact before adopting the architecture.
An infrastructure review can include workload tests, concurrency limits, a budget and criteria for expanding or separating infrastructure. Deliverables are defined in the proposal.
Consulting for your project
Infrastructure review, deployment and ongoing operations, with scope and pricing defined in the proposal.
Consulting
Infrastructure review, deployment and management for your company's AI systems.