AI Silicon, High-Performance Compute & System Architecture

AI Hardware Architecture: GPU, TPU, CPU & Cluster Engineering

Modern AI engineering is fundamentally constrained by physics, thermodynamics, and memory bandwidth. Understand how CPUs, GPUs, TPUs, and custom NPUs fit together to power production-scale training and inference pipelines for Large Language Models (LLMs) and multi-agent SaaS platforms.

CPU vs GPU vs TPU vs NPU Memory Bandwidth Wall (HBM3e/4) Prefill vs Decode Bottlenecks NVLink 5 & InfiniBand Fabric PagedAttention & vLLM Serving LLM SaaS VRAM Economics

01 The Silicon Trio: CPU vs GPU vs TPU vs Custom ASICs

Compute Paradigms

No single processor can efficiently execute an entire LLM application stack. An enterprise AI system is a heterogeneous hierarchy where each chip excels at distinct computational patterns:

CPU (Host Orchestrator) Latency-Optimized

Complex Control Flow

Features few massive cores with deep instruction pipelines, branch prediction, and multi-megabyte L3 caches. Optimized for sequential code execution and low memory latency.

Key Hardware: AMD EPYC 9005 (Turin), Intel Xeon 6, ARM Grace / Graviton 4.
Role in SaaS: API gateway, tokenization, agent state loops, JSON parsing, tool calling, RAG vector indexing (HNSW).
GPU (Matrix Workhorse) Throughput-Optimized

Massive Parallelism

Thousands of small SIMT (Single Instruction, Multiple Threads) cores paired with High Bandwidth Memory (HBM). Equipped with specialized Tensor Cores for mixed-precision matrix multiplication.

Key Hardware: NVIDIA H100 / H200 / B200 (Blackwell), AMD Instinct MI300X.
Role in SaaS: General-purpose LLM pretraining, fine-tuning, and high-concurrency production inference (vLLM / TensorRT-LLM).
TPU (Systolic Matrix ASIC) ASIC Specialization

Direct Systolic Flow

Google's custom ASIC designed around 2D Matrix Multiply Units (MXUs). Data flows through a systolic array of multiply-accumulators without reading and writing back to registers at each step.

Key Hardware: Google TPU v5p (Training), TPU v6e Trillium (Cost-optimized inference & fine-tuning).
Role in SaaS: Cost-effective massive training pods and high-volume Gemini inference inside Google Cloud / GKE.
NPU / LPU (Custom Silicon) Ultra-Low Latency

SRAM-Centric Architecture

Chips engineered strictly for inference without standard HBM latency. Groq's LPU holds weights entirely in ultra-fast on-chip SRAM (80 TB/s bandwidth), eliminating memory-bandwidth bottlenecks.

Key Hardware: Groq LPU, AWS Inferentia2 / Trainium2, Cerebras CS-3 Wafer-Scale Engine.
Role in SaaS: Real-time voice agents, sub-50ms token generation, and predictable fixed-cost inference at scale.

Processor Architecture & Workload Matrix

Silicon Tradeoffs
Architecture Memory Subsystem Bandwidth (Per Chip) Compute Precision Interconnect Ideal SaaS Workload
Modern Server CPU DDR5 Channels (up to 1.5 TB) 300 – 600 GB/s FP32, FP64, AVX-512, AMX (BF16/INT8) PCIe Gen 5 / CXL 2.0 API Gateway, Agent State Machine, RAG vector index
NVIDIA H100 / H200 SXM 80GB HBM3 / 141GB HBM3e 3.35 – 4.8 TB/s FP64, TF32, FP16, BF16, FP8 (E4M3/E5M2) NVLink 4 (900 GB/s) Frontier model pretraining, high-throughput vLLM serving
NVIDIA B200 / GB200 192GB HBM3e (dual-die) 8.0 TB/s FP16, BF16, FP8, NVFP4 (Microscaling) NVLink 5 (1,800 GB/s) 405B+ LLM inference on single rack; reasoning models
Google TPU v5p 95GB HBM2e 2.76 TB/s BF16, INT8, FP32 (accumulation) Optical Circuit Switch (OCS) ICI Multi-thousand-chip distributed model pretraining
Groq LPU (ASIC) 230 MB on-chip SRAM 80 TB/s (on-die) FP16, INT8 Direct Chip-to-Chip mesh Voice agents, sub-second latency streaming generation

03 How It All Comes Together: Enterprise LLM / Agent SaaS Architecture

End-to-End Blueprint

Building a production-ready AI Agent SaaS (supporting streaming completions, tool execution, multi-agent coordination, and custom fine-tuning) requires orchestrating CPUs, GPUs, high-speed networks, and specialized storage in a 5-tier architecture:

Heterogeneous Infrastructure Topology

Separation of Concerns: CPU Control Plane + GPU Compute Plane
Tier 1 CPU Only

Ingress & Auth

API Gateway, JWT auth, rate limits, model routing (8B vs 70B), semantic cache (Redis).

Tier 2 CPU + RAM

Agent Workflow

LangGraph state machines, RAG retrieval (Qdrant), tool execution, sandboxed code execution.

Tier 3 GPU Fleet

Inference Engine

vLLM / TensorRT-LLM, continuous batching, PagedAttention, speculative decoding (H100/H200).

Tier 4 Training Pod

SFT & RLHF Pod

Ray / Slurm cluster, FSDP / DeepSpeed, DPO & RLHF fine-tuning, NVLink 5 + InfiniBand.

Tier 5 Storage Fabric

Storage & Checkpoints

GPUDirect Storage (GDS), FSx Lustre / Ceph, S3 model weight bucket, streaming weights.

Why Not Run Everything on GPUs?

GPU hours cost $2.50 to $4.50+ each. Having a GPU wait idly while an external web search API executes (taking 800ms) burns money. The CPU tier manages long-lived agent states, assembling prompts and calling GPU inference only when token generation is strictly required.

PD Disaggregation (Prefill vs Decode)

State-of-the-art SaaS setups separate physical nodes into Prefill Workers (compute-heavy, batching massive context prompts) and Decode Workers (memory-bandwidth optimized, generating tokens). The KV cache is transferred between them over 400 Gbps RDMA fabrics in under 5 milliseconds.

Speculative Decoding Pairs

A small 8B "draft" model runs on low-cost compute (or quantized FP4) predicting 4–6 tokens in advance. A massive 70B model validates all draft tokens in a single parallel forward pass. This doubles inference speed (2x–3x TPS) without changing mathematical outputs.

04 Production GPU Inference Engine: vLLM & PagedAttention

High-Throughput Serving

Raw PyTorch inference suffers from severe memory fragmentation and static batch sizes. Modern production inference engines (vLLM, TensorRT-LLM, SGLang) solve this through PagedAttention (treating KV cache like virtual memory pages) and Continuous (Iteration-Level) Batching.

Production Multi-GPU vLLM Launch Configuration (Docker / Kubernetes)

Configures 4-way Tensor Parallelism, FP8 KV caching, speculative drafting, and 32k context handling.

vllm-serve.sh
#!/usr/bin/env bash
# Production vLLM Launch Script for 4x NVIDIA H100 (80GB SXM5)
# Serves Llama-3.3-70B-Instruct with Tensor Parallelism & Speculative Decoding

python3 -m vllm.entrypoints.openai.api_server \
    --model meta-llama/Llama-3.3-70B-Instruct \
    --tensor-parallel-size 4 \
    --max-model-len 32768 \
    --gpu-memory-utilization 0.92 \
    --kv-cache-dtype fp8 \
    --enable-chunked-prefill \
    --max-num-batched-tokens 8192 \
    --speculative-model meta-llama/Llama-3.1-8B-Instruct \
    --num-speculative-tokens 5 \
    --enable-prefix-caching \
    --trust-remote-code \
    --port 8000 \
    --host 0.0.0.0

# ------------------------------------------------------------------------------
# WHY THESE FLAGS MATTER IN ENTERPRISE SAAS:
# 1. --tensor-parallel-size 4: Shards model weights across 4 GPUs via high-speed NVLink
# 2. --kv-cache-dtype fp8: Halves KV cache memory footprint, doubling max concurrency
# 3. --enable-chunked-prefill: Prevents long 32k prompt ingestion from starving active decodes
# 4. --speculative-model: 8B draft model proposes 5 tokens; 70B verifies in single pass (2.2x TPS)
# 5. --enable-prefix-caching: Reuses KV cache for identical system prompts & RAG context headers
# ------------------------------------------------------------------------------
PagedAttention

Zero Memory Waste

Traditional LLM serving pre-allocated contiguous memory for maximum context (e.g. 8k), wasting 60%–80% of VRAM on short requests. PagedAttention allocates non-contiguous physical blocks on demand, boosting serving throughput by 2x–4x.

Continuous Batching

Iteration-Level Scheduling

Instead of waiting for an entire batch to finish generating, new requests enter the batch immediately after every single token generation step, keeping GPUs at 100% saturation continuously.

Prefix Caching (APC)

Instant Agent Prompts

AI agents repeatedly pass long system instructions, API schema definitions, and RAG documents. Automatic Prefix Caching stores prompt KV cache in VRAM, slashing Time-To-First-Token (TTFT) from 1,200ms to under 40ms.

05 Distributed Training Clusters: 3D Parallelism & Networking Fabric

Multi-Node Scaling

When training or fine-tuning models that exceed single-GPU memory, engineering teams deploy 3D Parallelism across clusters connected by high-speed non-blocking fabrics:

Tensor Parallelism (TP) Intra-Node Only

Intra-Layer Sharding

Splits individual weight matrices (Linear & Attention projections) across multiple GPUs (e.g. 8 GPUs within one node).

Communication: All-Reduce per transformer layer.
Requirement: Ultra-fast NVLink (900–1800 GB/s). Never run TP across standard network cables!
Pipeline Parallelism (PP) Inter-Node Stages

Layer-by-Layer Stacking

Partitions transformer layers sequentially across nodes (e.g. Layers 1–20 on Node 1, Layers 21–40 on Node 2).

Communication: Point-to-point peer activation transfer.
Challenge: "Bubble" idle time mitigated by 1F1B (One Forward, One Backward) scheduling.
Data Parallelism (FSDP / ZeRO) Fleet-Scale Scaling

Zero Redundancy Sharding

Fully Sharded Data Parallel (FSDP) and DeepSpeed ZeRO-3 shard optimizer states, gradients, and model parameters across thousands of GPUs.

Communication: All-Gather during forward; Reduce-Scatter during backward pass.
Requirement: 400G InfiniBand or RoCE v2 with GPUDirect RDMA.

PyTorch Fully Sharded Data Parallel (FSDP) Code Setup

Production snippet showing auto-wrapping transformer blocks and gradient checkpointing.

train_fsdp.py
import torch
from torch.distributed.fsdp import (
    FullyShardedDataParallel as FSDP,
    MixedPrecision,
    ShardingStrategy,
    BackwardPrefetch,
)
from torch.distributed.fsdp.wrap import size_based_auto_wrap_policy
from transformers.models.llama.modeling_llama import LlamaDecoderLayer

# 1. Mixed Precision Policy: BF16 computation with FP32 reduction buffers
fpSixteen = MixedPrecision(
    param_dtype=torch.bfloat16,
    reduce_dtype=torch.float32,
    buffer_dtype=torch.bfloat16,
)

# 2. Wrap the model in FSDP (Shards weights, gradients & Adam optimizer)
model = FSDP(
    base_model,
    auto_wrap_policy=size_based_auto_wrap_policy,
    mixed_precision=fpSixteen,
    sharding_strategy=ShardingStrategy.FULL_SHARD, # ZeRO-3 equivalent
    backward_prefetch=BackwardPrefetch.BACKWARD_PRE,
    device_id=torch.cuda.current_device(),
    limit_all_gathers=True,
)

# 3. Enable Activation Checkpointing (Saves 60% VRAM by recomputing forward pass)
from torch.distributed.algorithms._checkpoint.checkpoint_wrapper import apply_activation_checkpointing
apply_activation_checkpointing(model, check_fn=lambda submodule: isinstance(submodule, LlamaDecoderLayer))

06 VRAM Capacity Formula & AI SaaS Infrastructure Economics

Financial Engineering

Miscalculating GPU VRAM requirements leads either to out-of-memory (OOM) crashes in production or over-provisioning that bankrupts early-stage startups. Use these mathematical formulas for exact sizing:

1. Total VRAM Requirement Formula (Inference)

Total_VRAM = VRAM_Weights + VRAM_KV_Cache + VRAM_Overhead
• VRAM_Weights: Params (Billion) × Precision (Bytes)
  - 70B @ 16-bit (BF16) = 140 GB
  - 70B @ 8-bit (FP8) = 70 GB
  - 70B @ 4-bit (AWQ/GPTQ) = 35 GB
• VRAM_KV_Cache: 2 × Layers × Heads × HeadDim × SeqLen × Batch × PrecisionBytes
  - Llama 3 70B with 128k context & Batch 8 = ~48 GB VRAM in FP16 (or 24 GB in FP8).
• VRAM_Overhead: Add 20% for CUDA context, activations, and fragmentation buffers.

2. SaaS Infrastructure Decision Framework

Tier A: < 10M Tokens / Day (Pay-Per-Token APIs)

Use OpenAI, Anthropic, or DeepSeek API endpoints. Zero infrastructure management, zero idle cost.

Tier B: 10M – 500M Tokens / Day (Cloud GPU Rentals)

Rent on-demand or 1-year reserved 8x H100 nodes via specialized AI clouds (RunPod, Lambda, CoreWeave) at $2.20–$3.20/GPU-hr.

Tier C: > 500M Tokens / Day (Dedicated Bare Metal / Colocation)

Purchase DGX / Supermicro clusters in colocation data centers. Amortized cost drops to ~$1.20/GPU-hr with custom liquid cooling.

07 Frequently Asked Questions (FAQ)

Why can't we run Tensor Parallelism across standard Ethernet or Cloud VPC networks? →

Tensor Parallelism executes an All-Reduce collective synchronization after every single transformer layer (over 80 times per forward pass). Standard 100G or 200G Ethernet has microsecond latencies that leave compute cores starving for data 90% of the time. TP requires NVLink (900 to 1,800 GB/s per GPU with sub-microsecond latency) or dense NVSwitch fabrics found strictly within a single physical server chassis.

What is the difference between H100 SXM5 and H100 PCIe? →

The H100 SXM5 connects via NVIDIA's proprietary mezzanine board with full NVLink 4 support (900 GB/s bidirectional inter-GPU bandwidth, 700W TDP, and maximum 3.35 TB/s HBM3 memory bandwidth). The H100 PCIe plugs into standard motherboard slots (350W TDP, reduced clock speeds, only 2 TB/s HBM2e bandwidth, and no direct board-level NVLink crossbar). SXM5 provides 30% to 50% higher throughput for multi-GPU training and distributed inference.

When should a startup use Google TPUs instead of NVIDIA GPUs? →

TPUs excel when: (1) You are building inside Google Cloud (GKE) and training models at massive scale using JAX or PyTorch/XLA; (2) You need predictable availability and pricing through Google TPU v5p/v6e pods connected via Optical Circuit Switching (OCS); (3) You are serving standardized models with high throughput. However, if your codebase relies on custom CUDA kernels, Triton kernels, FlashAttention-3, or NVIDIA-specific libraries (TensorRT-LLM), GPUs remain the industry standard.

How does Speculative Decoding achieve higher tokens-per-second without quality loss? →

Autoregressive decoding is bottlenecked by streaming weights from HBM for one token at a time. Speculative decoding runs a tiny draft model (e.g. 8B) to generate a speculative sequence of tokens quickly. The large target model (e.g. 70B) checks all proposed tokens simultaneously in a single parallel GEMM forward pass. Any accepted tokens are kept; the first rejected token triggers a correction. Because the large model's mathematical verification governs acceptance, output distribution is 100% mathematically identical to running the 70B model alone.

What is Prefill and Decode (PD) Disaggregation, and why is it trending? →

Traditionally, an inference worker performs both prefill (processing prompt) and decode (streaming tokens). However, when a massive 32k prompt arrives, it occupies the GPU for 500ms, stalling token generation for all other concurrent users and causing severe jitter in Inter-Token Latency (ITL). Disaggregation assigns dedicated GPU nodes to compute prefill, transfers the resulting KV cache across 400G RDMA networks in under 5ms, and lets dedicated decode nodes stream tokens uninterrupted.

How does FP8 and NVFP4 quantization affect LLM SaaS unit economics? →

Inference cost is directly proportional to how many GPUs are required to hold a model and stream its memory. Running a 70B model in FP16 requires at least two 80GB GPUs (160GB total). In FP8, it fits on a single 80GB GPU or 40GB partition, cutting infrastructure spend in half. In NVFP4 on Blackwell B200, a 70B model occupies ~35GB, allowing 4x greater concurrent user capacity per server with minimal loss in benchmark accuracy.