NVIDIA AI Infrastructure
This job-oriented course covers the infrastructure required to train, serve, monitor, secure, and operate AI workloads at scale. Learn NVIDIA GPU architecture, CUDA, data-center platforms, InfiniBand, NVLink, storage, NVIDIA AI Enterprise, NGC, containers, Kubernetes, Run:ai, training and inference infrastructure, generative AI, NIM, operations, and lifecycle management.
Curriculum Areas
- Module 1 - AI Infrastructure Fundamentals
- Module 2 - NVIDIA GPU Architecture
- Module 3 - NVIDIA Data Center GPUs
- Module 4 - CUDA & NVIDIA Software Stack
- Module 5 - AI Compute Infrastructure
- Module 6 - NVIDIA AI Networking
- Module 7 - NVLink & NVSwitch
- Module 8 - AI Storage Infrastructure
- Module 9 - NVIDIA AI Enterprise
- Module 10 - NVIDIA NGC
- Module 11 - AI Containers & Docker
- Module 12 - Kubernetes & NVIDIA GPUs
- Module 13 - NVIDIA Run:ai
- Module 14 - Virtualization & GPU Sharing
- Module 15 - AI Model Training Infrastructure
- Module 16 - AI Inference Infrastructure
- Module 17 - Generative AI Infrastructure
- Module 18 - NVIDIA NIM
- Module 19 - Monitoring & Troubleshooting
- Module 20 - Security & Operations
NVIDIA AI Infrastructure Curriculum
- What is AI Infrastructure?
- AI/ML workloads and infrastructure requirements
- CPU vs GPU architecture
- Training vs inference infrastructure
- AI data-center architecture
- Compute, networking and storage
- On-premises vs cloud AI infrastructure
- NVIDIA AI ecosystem overview
- NVIDIA GPU fundamentals
- GPU architecture evolution
- CUDA cores
- Tensor Cores
- GPU memory and HBM
- NVLink
- NVSwitch
- PCIe
- GPU partitioning concepts
- GPU performance metrics
- NVIDIA A100
- NVIDIA H100
- NVIDIA H200
- NVIDIA B200 / Blackwell architecture
- GPU generations and capabilities
- Training GPUs vs inference GPUs
- GPU selection
- GPU sizing
- Capacity planning
- CUDA fundamentals
- CUDA Toolkit
- CUDA libraries
- cuDNN
- NCCL
- TensorRT
- GPU-accelerated computing
- CUDA-aware applications
- NVIDIA software ecosystem
- GPU servers
- GPU nodes
- Multi-GPU systems
- NVIDIA DGX concepts
- NVIDIA HGX platforms
- GPU clusters
- CPU-GPU architecture
- Power and thermal requirements
- AI infrastructure sizing
- Why networking matters for AI
- InfiniBand fundamentals
- NVIDIA Quantum InfiniBand
- NVIDIA Spectrum Ethernet
- RDMA
- RoCE
- GPU-to-GPU communication
- Network topology
- High-performance AI clusters
- Network troubleshooting
- NVLink architecture
- GPU interconnects
- NVSwitch
- Multi-GPU communication
- Bandwidth optimization
- GPU scaling
- Distributed AI workloads
- AI storage requirements
- High-throughput storage
- Parallel file systems
- NVIDIA GPUDirect Storage
- Object storage
- NVMe storage
- Data pipelines
- Storage bottlenecks
- Data ingestion for AI training
- NVIDIA AI Enterprise overview
- Enterprise AI software stack
- Supported infrastructure
- Virtualization
- AI application deployment
- Enterprise security
- Software lifecycle
- Production AI environments
- NVIDIA NGC overview
- NGC containers
- Pre-trained models
- AI frameworks
- Container deployment
- Model repositories
- Security scanning
- Using NGC in enterprise environments
- Docker fundamentals
- NVIDIA Container Toolkit
- GPU-enabled containers
- CUDA containers
- Building AI containers
- GPU resource allocation
- Container troubleshooting
- Kubernetes fundamentals
- Kubernetes for AI workloads
- NVIDIA GPU Operator
- GPU scheduling
- Device plugins
- GPU resources
- Node configuration
- Multi-GPU workloads
- Kubernetes troubleshooting
- GPU resource management
- GPU scheduling
- Workload orchestration
- GPU sharing
- Fractional GPUs
- Job scheduling
- Cluster utilization
- Multi-tenant AI infrastructure
- GPU virtualization concepts
- NVIDIA vGPU
- MIG - Multi-Instance GPU
- GPU partitioning
- Workload isolation
- Resource allocation
- Enterprise virtualization
- AI training architecture
- Distributed training
- Data parallelism
- Model parallelism
- Pipeline parallelism
- Mixed precision
- GPU utilization
- NCCL communication
- Scaling AI training
- Training vs inference
- Inference architecture
- NVIDIA Triton Inference Server
- TensorRT
- TensorRT-LLM
- Model optimization
- Batch inference
- Real-time inference
- Inference performance tuning
- LLM infrastructure requirements
- GPU sizing for LLMs
- LLM training infrastructure
- LLM inference infrastructure
- RAG infrastructure
- Vector databases
- Model serving
- Token throughput
- Latency optimization
- NVIDIA NIM fundamentals
- NIM architecture
- Deploying prebuilt AI models
- Container-based model serving
- NIM + Kubernetes
- NIM + enterprise infrastructure
- LLM inference optimization
- GPU monitoring
- NVIDIA DCGM
- DCGM Exporter
- Prometheus
- Grafana
- GPU utilization
- GPU memory monitoring
- Temperature and power
- ECC errors
- Performance troubleshooting
- AI infrastructure security
- GPU workload isolation
- Identity and access management
- Container security
- Kubernetes security
- Network security
- Data protection
- Patch management
- Infrastructure lifecycle management