🖥️ AI Datacenter - Multi-Vendor GPU Server Racks
❄️
Rack A1 - NVIDIA
NVIDIA DGX H100
8x H100 80GB GPUs
NVIDIA RTX 5090
4x RTX 5090 32GB
❄️
Rack B1 - AMD
AMD Instinct MI300X
8x MI300X 192GB
AMD Radeon Pro W7900
4x W7900 48GB
❄️
Rack C1 - Intel
Intel Gaudi 3
8x Gaudi 3 AI Accelerators
Intel Arc B580
4x Arc B580 16GB
❄️
Rack D1 - Specialized
Cerebras CS-3
Wafer-Scale Engine
Google TPU v5
TPU Pod Slice
32 GPUs
Total GPU Count
1.8 PetaFLOPS
Compute Power
18°C
Cooling Temperature
99.9%
Uptime SLA
GPU Performance Calculator for LLM Models
Estimate GPU requirements, inference performance, and costs for the latest open-weights Large Language Models you can self-host — including Kimi K3, Qwen3.8, DeepSeek-V4, GLM-5.2, Llama 4, Mistral, and gpt-oss
Performance Estimate
GPU Memory Required
-
VRAM needed per GPUGPUs Needed
-
Minimum GPU countInference Speed
-
Tokens per secondLatency (First Token)
-
Time to first tokenEst. Monthly Cost
-
Cloud GPU pricingPower Consumption
-
TDP per GPURecommendations
- Select model and GPU configuration to see recommendations
Important Notes
- Covers open-weights models only — Kimi K3, Qwen3.8, DeepSeek-V4, GLM-5.2, Llama 4, Mistral, and gpt-oss. Closed-weight APIs (GPT, Claude, Gemini, Grok) are excluded because they cannot be self-hosted
- Parameter counts for mixture-of-experts models are totals; active parameters per token are much lower, so real throughput is typically higher than estimated here
- Covers NVIDIA Blackwell (B300, B200), Hopper/Ampere (H100, A100, V100), workstation GPUs, GeForce RTX 50/40/30 series, and the DGX Spark desktop system
- Estimates are based on typical configurations and may vary based on implementation details
- Training requires significantly more memory than inference (typically 3-4x model size)
- Actual performance depends on model optimization, batch size, and sequence length
- MoE (Mixture of Experts) models like DeepSeek-V3 and Mixtral use fewer active parameters during inference
- Costs shown are approximate cloud pricing; on-premises costs will differ significantly
- Multi-GPU setups require high-speed interconnects (NVLink, InfiniBand) for optimal performance
Need Help Choosing the Right Infrastructure?
Our infrastructure experts can help you design the optimal GPU solution for your AI workloads.
Back to Infrastructure Page