TP
TensorPlanLLM GPU deployment calculator
Local calculation Source
Current estimate

Deployment analysis

ESTIMATED PROFILE
Required topology

Select a model

Capacity and throughput sizing

Usable VRAM consumed
Total demand
Configuration
Weightsincluding metadata
KV cacheat peak reservation
Aggregate decoderoofline estimate
Per requestfair-share estimate

Memory allocation

Cluster total

Throughput limits

Roofline model
Compute
Memory bandwidth

Hardware candidates

Ranked by fit, topology, speed, and GPU count
RankConfigurationVRAM reserveEst. decodeScore

Saved scenarios

Save the current estimate to compare models, quantization formats, or GPU options.

Calculation details

Method and limitations

Memory reserves the full configured context for each concurrent request. Decode throughput uses the lower of an active-parameter compute ceiling and a memory-bandwidth ceiling, then applies runtime, batching, and interconnect efficiency. Use this for planning, then benchmark the exact checkpoint and engine before procurement.

Edit model configuration

Override the selected profile for this session.