01 The Silicon Trio: CPU vs GPU vs TPU vs Custom ASICs
Compute ParadigmsNo single processor can efficiently execute an entire LLM application stack. An enterprise AI system is a heterogeneous hierarchy where each chip excels at distinct computational patterns:
Complex Control Flow
Features few massive cores with deep instruction pipelines, branch prediction, and multi-megabyte L3 caches. Optimized for sequential code execution and low memory latency.
Massive Parallelism
Thousands of small SIMT (Single Instruction, Multiple Threads) cores paired with High Bandwidth Memory (HBM). Equipped with specialized Tensor Cores for mixed-precision matrix multiplication.
Direct Systolic Flow
Google's custom ASIC designed around 2D Matrix Multiply Units (MXUs). Data flows through a systolic array of multiply-accumulators without reading and writing back to registers at each step.
SRAM-Centric Architecture
Chips engineered strictly for inference without standard HBM latency. Groq's LPU holds weights entirely in ultra-fast on-chip SRAM (80 TB/s bandwidth), eliminating memory-bandwidth bottlenecks.
Processor Architecture & Workload Matrix
Silicon Tradeoffs| Architecture | Memory Subsystem | Bandwidth (Per Chip) | Compute Precision | Interconnect | Ideal SaaS Workload |
|---|---|---|---|---|---|
| Modern Server CPU | DDR5 Channels (up to 1.5 TB) | 300 – 600 GB/s | FP32, FP64, AVX-512, AMX (BF16/INT8) | PCIe Gen 5 / CXL 2.0 | API Gateway, Agent State Machine, RAG vector index |
| NVIDIA H100 / H200 SXM | 80GB HBM3 / 141GB HBM3e | 3.35 – 4.8 TB/s | FP64, TF32, FP16, BF16, FP8 (E4M3/E5M2) | NVLink 4 (900 GB/s) | Frontier model pretraining, high-throughput vLLM serving |
| NVIDIA B200 / GB200 | 192GB HBM3e (dual-die) | 8.0 TB/s | FP16, BF16, FP8, NVFP4 (Microscaling) | NVLink 5 (1,800 GB/s) | 405B+ LLM inference on single rack; reasoning models |
| Google TPU v5p | 95GB HBM2e | 2.76 TB/s | BF16, INT8, FP32 (accumulation) | Optical Circuit Switch (OCS) ICI | Multi-thousand-chip distributed model pretraining |
| Groq LPU (ASIC) | 230 MB on-chip SRAM | 80 TB/s (on-die) | FP16, INT8 | Direct Chip-to-Chip mesh | Voice agents, sub-second latency streaming generation |
02 The Memory Bandwidth Wall & Modern Hardware Trends
Physics of LLMsThe greatest misconception in AI engineering is that LLMs are limited by raw floating-point computing power (FLOPS). In reality, LLM generation is dominated by the Memory Bandwidth Wall. While computing power has grown exponentially, memory transfer speeds have lagged, forcing architectural bifurcations.
Parallel Matrix Multiplication
When a prompt arrives (e.g. 4,000 tokens of retrieved documents), all input tokens are processed simultaneously in a single large Matrix-Matrix multiplication (GEMM).
- High Arithmetic Intensity: Hundreds of FLOPS performed per byte read from memory.
- Fully saturates Tensor Cores: Modern GPUs operate at 50%–70% of peak theoretical compute capacity (MFU).
- Optimization Goal: FlashAttention-3, Chunked Prefill to avoid starving concurrent decode streams.
Sequential Memory Streaming
Autoregressive generation predicts exactly one token at a time. For each new token generated, the GPU must stream every single model weight and the historical Key-Value (KV) cache from HBM into on-chip registers.
- Near-Zero Arithmetic Intensity: Just ~1 FLOP performed per byte transferred.
- Tensor Cores Sit Idle: GPUs typically achieve only 1% to 5% compute utilization during single-user decode!
- Generation Speed Limit: Max Token/sec =
Memory Bandwidth (Bytes/sec) ÷ Model Size in VRAM (Bytes).
NVLink 5 & Optical Fabrics
NVIDIA’s GB200 NVL72 uses copper cartridge backplanes delivering 130 TB/s bisection bandwidth. The entire 72-GPU rack behaves as a single unified GPU with 13.8 TB of shared memory, allowing a 70B model to fit entirely in L2/L3 cache speeds.
NVFP4 & MX Formats
Transitioning weights and KV cache from FP16 to FP8 doubles memory bandwidth throughput. New NVFP4 (4-bit floating point with 16-element block scaling) cuts memory bandwidth requirements by 4x without losing reasoning accuracy.
Direct-to-Chip Liquid Cooling
With chips like Blackwell B200 dissipating 1,000W to 1,200W per package and racks consuming 120 kW+, traditional air conditioning is obsolete. Facilities now mandate closed-loop direct-to-chip water cooling and immersion tanks.
03 How It All Comes Together: Enterprise LLM / Agent SaaS Architecture
End-to-End BlueprintBuilding a production-ready AI Agent SaaS (supporting streaming completions, tool execution, multi-agent coordination, and custom fine-tuning) requires orchestrating CPUs, GPUs, high-speed networks, and specialized storage in a 5-tier architecture:
Heterogeneous Infrastructure Topology
Separation of Concerns: CPU Control Plane + GPU Compute PlaneIngress & Auth
API Gateway, JWT auth, rate limits, model routing (8B vs 70B), semantic cache (Redis).
Agent Workflow
LangGraph state machines, RAG retrieval (Qdrant), tool execution, sandboxed code execution.
Inference Engine
vLLM / TensorRT-LLM, continuous batching, PagedAttention, speculative decoding (H100/H200).
SFT & RLHF Pod
Ray / Slurm cluster, FSDP / DeepSpeed, DPO & RLHF fine-tuning, NVLink 5 + InfiniBand.
Storage & Checkpoints
GPUDirect Storage (GDS), FSx Lustre / Ceph, S3 model weight bucket, streaming weights.
Why Not Run Everything on GPUs?
GPU hours cost $2.50 to $4.50+ each. Having a GPU wait idly while an external web search API executes (taking 800ms) burns money. The CPU tier manages long-lived agent states, assembling prompts and calling GPU inference only when token generation is strictly required.
PD Disaggregation (Prefill vs Decode)
State-of-the-art SaaS setups separate physical nodes into Prefill Workers (compute-heavy, batching massive context prompts) and Decode Workers (memory-bandwidth optimized, generating tokens). The KV cache is transferred between them over 400 Gbps RDMA fabrics in under 5 milliseconds.
Speculative Decoding Pairs
A small 8B "draft" model runs on low-cost compute (or quantized FP4) predicting 4–6 tokens in advance. A massive 70B model validates all draft tokens in a single parallel forward pass. This doubles inference speed (2x–3x TPS) without changing mathematical outputs.
04 Production GPU Inference Engine: vLLM & PagedAttention
High-Throughput ServingRaw PyTorch inference suffers from severe memory fragmentation and static batch sizes. Modern production inference engines (vLLM, TensorRT-LLM, SGLang) solve this through PagedAttention (treating KV cache like virtual memory pages) and Continuous (Iteration-Level) Batching.
Production Multi-GPU vLLM Launch Configuration (Docker / Kubernetes)
Configures 4-way Tensor Parallelism, FP8 KV caching, speculative drafting, and 32k context handling.
Zero Memory Waste
Traditional LLM serving pre-allocated contiguous memory for maximum context (e.g. 8k), wasting 60%–80% of VRAM on short requests. PagedAttention allocates non-contiguous physical blocks on demand, boosting serving throughput by 2x–4x.
Iteration-Level Scheduling
Instead of waiting for an entire batch to finish generating, new requests enter the batch immediately after every single token generation step, keeping GPUs at 100% saturation continuously.
Instant Agent Prompts
AI agents repeatedly pass long system instructions, API schema definitions, and RAG documents. Automatic Prefix Caching stores prompt KV cache in VRAM, slashing Time-To-First-Token (TTFT) from 1,200ms to under 40ms.
05 Distributed Training Clusters: 3D Parallelism & Networking Fabric
Multi-Node ScalingWhen training or fine-tuning models that exceed single-GPU memory, engineering teams deploy 3D Parallelism across clusters connected by high-speed non-blocking fabrics:
Intra-Layer Sharding
Splits individual weight matrices (Linear & Attention projections) across multiple GPUs (e.g. 8 GPUs within one node).
Layer-by-Layer Stacking
Partitions transformer layers sequentially across nodes (e.g. Layers 1–20 on Node 1, Layers 21–40 on Node 2).
Zero Redundancy Sharding
Fully Sharded Data Parallel (FSDP) and DeepSpeed ZeRO-3 shard optimizer states, gradients, and model parameters across thousands of GPUs.
PyTorch Fully Sharded Data Parallel (FSDP) Code Setup
Production snippet showing auto-wrapping transformer blocks and gradient checkpointing.
06 VRAM Capacity Formula & AI SaaS Infrastructure Economics
Financial EngineeringMiscalculating GPU VRAM requirements leads either to out-of-memory (OOM) crashes in production or over-provisioning that bankrupts early-stage startups. Use these mathematical formulas for exact sizing:
1. Total VRAM Requirement Formula (Inference)
Params (Billion) × Precision (Bytes)- 70B @ 16-bit (BF16) = 140 GB
- 70B @ 8-bit (FP8) = 70 GB
- 70B @ 4-bit (AWQ/GPTQ) = 35 GB
2 × Layers × Heads × HeadDim × SeqLen × Batch × PrecisionBytes- Llama 3 70B with 128k context & Batch 8 = ~48 GB VRAM in FP16 (or 24 GB in FP8).
2. SaaS Infrastructure Decision Framework
Use OpenAI, Anthropic, or DeepSeek API endpoints. Zero infrastructure management, zero idle cost.
Rent on-demand or 1-year reserved 8x H100 nodes via specialized AI clouds (RunPod, Lambda, CoreWeave) at $2.20–$3.20/GPU-hr.
Purchase DGX / Supermicro clusters in colocation data centers. Amortized cost drops to ~$1.20/GPU-hr with custom liquid cooling.
07 Frequently Asked Questions (FAQ)
Why can't we run Tensor Parallelism across standard Ethernet or Cloud VPC networks? →
Tensor Parallelism executes an All-Reduce collective synchronization after every single transformer layer (over 80 times per forward pass). Standard 100G or 200G Ethernet has microsecond latencies that leave compute cores starving for data 90% of the time. TP requires NVLink (900 to 1,800 GB/s per GPU with sub-microsecond latency) or dense NVSwitch fabrics found strictly within a single physical server chassis.
What is the difference between H100 SXM5 and H100 PCIe? →
The H100 SXM5 connects via NVIDIA's proprietary mezzanine board with full NVLink 4 support (900 GB/s bidirectional inter-GPU bandwidth, 700W TDP, and maximum 3.35 TB/s HBM3 memory bandwidth). The H100 PCIe plugs into standard motherboard slots (350W TDP, reduced clock speeds, only 2 TB/s HBM2e bandwidth, and no direct board-level NVLink crossbar). SXM5 provides 30% to 50% higher throughput for multi-GPU training and distributed inference.
When should a startup use Google TPUs instead of NVIDIA GPUs? →
TPUs excel when: (1) You are building inside Google Cloud (GKE) and training models at massive scale using JAX or PyTorch/XLA; (2) You need predictable availability and pricing through Google TPU v5p/v6e pods connected via Optical Circuit Switching (OCS); (3) You are serving standardized models with high throughput. However, if your codebase relies on custom CUDA kernels, Triton kernels, FlashAttention-3, or NVIDIA-specific libraries (TensorRT-LLM), GPUs remain the industry standard.
How does Speculative Decoding achieve higher tokens-per-second without quality loss? →
Autoregressive decoding is bottlenecked by streaming weights from HBM for one token at a time. Speculative decoding runs a tiny draft model (e.g. 8B) to generate a speculative sequence of tokens quickly. The large target model (e.g. 70B) checks all proposed tokens simultaneously in a single parallel GEMM forward pass. Any accepted tokens are kept; the first rejected token triggers a correction. Because the large model's mathematical verification governs acceptance, output distribution is 100% mathematically identical to running the 70B model alone.
What is Prefill and Decode (PD) Disaggregation, and why is it trending? →
Traditionally, an inference worker performs both prefill (processing prompt) and decode (streaming tokens). However, when a massive 32k prompt arrives, it occupies the GPU for 500ms, stalling token generation for all other concurrent users and causing severe jitter in Inter-Token Latency (ITL). Disaggregation assigns dedicated GPU nodes to compute prefill, transfers the resulting KV cache across 400G RDMA networks in under 5ms, and lets dedicated decode nodes stream tokens uninterrupted.
How does FP8 and NVFP4 quantization affect LLM SaaS unit economics? →
Inference cost is directly proportional to how many GPUs are required to hold a model and stream its memory. Running a 70B model in FP16 requires at least two 80GB GPUs (160GB total). In FP8, it fits on a single 80GB GPU or 40GB partition, cutting infrastructure spend in half. In NVFP4 on Blackwell B200, a 70B model occupies ~35GB, allowing 4x greater concurrent user capacity per server with minimal loss in benchmark accuracy.