NVIDIA AI Infrastructure Training

Build and operate production AI infrastructure with NVIDIA GPUs, CUDA, AI networking, storage, Kubernetes, Run:ai, NIM, monitoring, security, and enterprise workloads.

NVIDIA AI Infrastructure

This job-oriented course covers the infrastructure required to train, serve, monitor, secure, and operate AI workloads at scale. Learn NVIDIA GPU architecture, CUDA, data-center platforms, InfiniBand, NVLink, storage, NVIDIA AI Enterprise, NGC, containers, Kubernetes, Run:ai, training and inference infrastructure, generative AI, NIM, operations, and lifecycle management.

Curriculum Areas

  1. Module 1 - AI Infrastructure Fundamentals
  2. Module 2 - NVIDIA GPU Architecture
  3. Module 3 - NVIDIA Data Center GPUs
  4. Module 4 - CUDA & NVIDIA Software Stack
  5. Module 5 - AI Compute Infrastructure
  6. Module 6 - NVIDIA AI Networking
  7. Module 7 - NVLink & NVSwitch
  8. Module 8 - AI Storage Infrastructure
  9. Module 9 - NVIDIA AI Enterprise
  10. Module 10 - NVIDIA NGC
  11. Module 11 - AI Containers & Docker
  12. Module 12 - Kubernetes & NVIDIA GPUs
  13. Module 13 - NVIDIA Run:ai
  14. Module 14 - Virtualization & GPU Sharing
  15. Module 15 - AI Model Training Infrastructure
  16. Module 16 - AI Inference Infrastructure
  17. Module 17 - Generative AI Infrastructure
  18. Module 18 - NVIDIA NIM
  19. Module 19 - Monitoring & Troubleshooting
  20. Module 20 - Security & Operations

NVIDIA AI Infrastructure Curriculum

  • What is AI Infrastructure?
  • AI/ML workloads and infrastructure requirements
  • CPU vs GPU architecture
  • Training vs inference infrastructure
  • AI data-center architecture
  • Compute, networking and storage
  • On-premises vs cloud AI infrastructure
  • NVIDIA AI ecosystem overview
  • NVIDIA GPU fundamentals
  • GPU architecture evolution
  • CUDA cores
  • Tensor Cores
  • GPU memory and HBM
  • NVLink
  • NVSwitch
  • PCIe
  • GPU partitioning concepts
  • GPU performance metrics
  • NVIDIA A100
  • NVIDIA H100
  • NVIDIA H200
  • NVIDIA B200 / Blackwell architecture
  • GPU generations and capabilities
  • Training GPUs vs inference GPUs
  • GPU selection
  • GPU sizing
  • Capacity planning
  • CUDA fundamentals
  • CUDA Toolkit
  • CUDA libraries
  • cuDNN
  • NCCL
  • TensorRT
  • GPU-accelerated computing
  • CUDA-aware applications
  • NVIDIA software ecosystem
  • GPU servers
  • GPU nodes
  • Multi-GPU systems
  • NVIDIA DGX concepts
  • NVIDIA HGX platforms
  • GPU clusters
  • CPU-GPU architecture
  • Power and thermal requirements
  • AI infrastructure sizing
  • Why networking matters for AI
  • InfiniBand fundamentals
  • NVIDIA Quantum InfiniBand
  • NVIDIA Spectrum Ethernet
  • RDMA
  • RoCE
  • GPU-to-GPU communication
  • Network topology
  • High-performance AI clusters
  • Network troubleshooting
  • NVLink architecture
  • GPU interconnects
  • NVSwitch
  • Multi-GPU communication
  • Bandwidth optimization
  • GPU scaling
  • Distributed AI workloads
  • AI storage requirements
  • High-throughput storage
  • Parallel file systems
  • NVIDIA GPUDirect Storage
  • Object storage
  • NVMe storage
  • Data pipelines
  • Storage bottlenecks
  • Data ingestion for AI training
  • NVIDIA AI Enterprise overview
  • Enterprise AI software stack
  • Supported infrastructure
  • Virtualization
  • AI application deployment
  • Enterprise security
  • Software lifecycle
  • Production AI environments
  • NVIDIA NGC overview
  • NGC containers
  • Pre-trained models
  • AI frameworks
  • Container deployment
  • Model repositories
  • Security scanning
  • Using NGC in enterprise environments
  • Docker fundamentals
  • NVIDIA Container Toolkit
  • GPU-enabled containers
  • CUDA containers
  • Building AI containers
  • GPU resource allocation
  • Container troubleshooting
  • Kubernetes fundamentals
  • Kubernetes for AI workloads
  • NVIDIA GPU Operator
  • GPU scheduling
  • Device plugins
  • GPU resources
  • Node configuration
  • Multi-GPU workloads
  • Kubernetes troubleshooting
  • GPU resource management
  • GPU scheduling
  • Workload orchestration
  • GPU sharing
  • Fractional GPUs
  • Job scheduling
  • Cluster utilization
  • Multi-tenant AI infrastructure
  • GPU virtualization concepts
  • NVIDIA vGPU
  • MIG - Multi-Instance GPU
  • GPU partitioning
  • Workload isolation
  • Resource allocation
  • Enterprise virtualization
  • AI training architecture
  • Distributed training
  • Data parallelism
  • Model parallelism
  • Pipeline parallelism
  • Mixed precision
  • GPU utilization
  • NCCL communication
  • Scaling AI training
  • Training vs inference
  • Inference architecture
  • NVIDIA Triton Inference Server
  • TensorRT
  • TensorRT-LLM
  • Model optimization
  • Batch inference
  • Real-time inference
  • Inference performance tuning
  • LLM infrastructure requirements
  • GPU sizing for LLMs
  • LLM training infrastructure
  • LLM inference infrastructure
  • RAG infrastructure
  • Vector databases
  • Model serving
  • Token throughput
  • Latency optimization
  • NVIDIA NIM fundamentals
  • NIM architecture
  • Deploying prebuilt AI models
  • Container-based model serving
  • NIM + Kubernetes
  • NIM + enterprise infrastructure
  • LLM inference optimization
  • GPU monitoring
  • NVIDIA DCGM
  • DCGM Exporter
  • Prometheus
  • Grafana
  • GPU utilization
  • GPU memory monitoring
  • Temperature and power
  • ECC errors
  • Performance troubleshooting
  • AI infrastructure security
  • GPU workload isolation
  • Identity and access management
  • Container security
  • Kubernetes security
  • Network security
  • Data protection
  • Patch management
  • Infrastructure lifecycle management
© 2026 All Rights Reserved by JOYATRES | Powered By Name Lelo