Current estimate
Deployment analysis
ESTIMATED PROFILE
Required topology
—
Select a model
Capacity and throughput sizing
Usable VRAM consumed—
- Total demand
- —
- Configuration
- —
Weights—including metadata
KV cache—at peak reservation
Aggregate decode—roofline estimate
Per request—fair-share estimate
Memory allocation
Cluster total
Throughput limits
Roofline modelCompute—
Memory bandwidth—
Hardware candidates
Ranked by fit, topology, speed, and GPU countRankConfigurationVRAM reserveEst. decodeScore
Saved scenarios
Save the current estimate to compare models, quantization formats, or GPU options.
Calculation details
Method and limitations
Memory reserves the full configured context for each concurrent request. Decode throughput uses the lower of an active-parameter compute ceiling and a memory-bandwidth ceiling, then applies runtime, batching, and interconnect efficiency. Use this for planning, then benchmark the exact checkpoint and engine before procurement.