NVIDIA Certified Professional - AI Operations
Validates the ability to monitor, troubleshoot, and optimize AI infrastructure operations using NVIDIA tools, including installation and deployment of NVIDIA software stacks with Base Command Manager, administration of Slurm and Kubernetes GPU clusters with user access and resource management, workload management for AI training and inference job scheduling, and troubleshooting and optimization of GPU performance and connectivity. The exam covers four domains: Installation and Deployment (31%), Administration (23%), Workload Management (23%), and Troubleshooting and Optimization (23%). Format: 70-75 multiple-choice questions, 120 minutes, proctored online.
Sample questions
A free preview of 15 source-grounded questions from this exam — answers and explanations included.
- Q1Installation and Deploymenthard
An architect is designing the storage layout for a DGX SuperPOD high-availability configuration and must place user home directories and shared data so the head nodes can fail over cleanly. What does NVIDIA require for these directories?
- A.They must live on local NVMe inside each head node and sync over the RDMA failover link.
- B.They must be stored in GPU memory and mirrored by GPUDirect Storage between head nodes.
- C.They must sit on a shared NFS filesystem made available to all of the head nodes for HA.Correct answer
- D.They must be packed into the DGX OS image and re-flashed onto the secondary head node.
Sources
Questions are grounded in 150 references from official and authoritative materials.