- 71 Actual Exam Questions
- Compatible with all Devices
- Printable Format
- No Download Limits
- 90 Days Free Updates
Get All AI Infrastructure Exam Questions with Validated Answers
| Vendor: | NVIDIA |
|---|---|
| Exam Code: | NCP-AII |
| Exam Name: | AI Infrastructure |
| Exam Questions: | 71 |
| Last Updated: | October 8, 2026 |
| Related Certifications: | NVIDIA-Certified Professional |
| Exam Tags: |
Looking for a hassle-free way to pass the NVIDIA AI Infrastructure exam? DumpsProvider provides the most reliable Dumps Questions and Answers, designed by NVIDIA certified experts to help you succeed in record time. Available in both PDF and Online Practice Test formats, our study materials cover every major exam topic, making it possible for you to pass potentially within just one day!
DumpsProvider is a leading provider of high-quality exam dumps, trusted by professionals worldwide. Our NVIDIA NCP-AII exam questions give you the knowledge and confidence needed to succeed on the first attempt.
Train with our NVIDIA NCP-AII exam practice tests, which simulate the actual exam environment. This real-test experience helps you get familiar with the format and timing of the exam, ensuring you're 100% prepared for exam day.
Your success is our commitment! That's why DumpsProvider offers a 100% money-back guarantee. If you don’t pass the NVIDIA NCP-AII exam, we’ll refund your payment within 24 hours no questions asked.
Don’t waste time with unreliable exam prep resources. Get started with DumpsProvider’s NVIDIA NCP-AII exam dumps today and achieve your certification effortlessly!
During a multi-day NeMo burn-in, intermittent "GPU fell off bus" errors occur. Which diagnostic approach isolates hardware faults?
The error 'GPU fell off bus' is a critical failure where the PCIe link between the GPU and the CPU/PCIe Switch has collapsed, often due to thermal stress, power instability, or physical hardware defects. To isolate the root cause during an intensive workload like NVIDIA NeMo (Large Language Model framework), the administrator must collect high-fidelity telemetry. DCGM (Data Center GPU Manager) diagnostics are designed for exactly this scenario. By running dcgmi diag -r 3 (a comprehensive hardware stress test) or monitoring health via dcgmi health --check concurrently with the workload, the system can capture the exact moment parameters like PCIe replay counts, temperature spikes, or XID errors occur. This data allows the engineer to determine if a specific H100 module is faulty or if the issue is systemic (e.g., a failing PCIe switch on the motherboard). Lowering the workload (Option C or D) might hide the symptom, but it does not diagnose the hardware's inability to handle peak power and data throughput.
A customer has just completed the first boot of their DGX system and is prompted to create an administrative user. What is the correct approach for setting up this user to ensure secure BMC and GRUB access?
During the initial 'first boot' setup of an NVIDIA DGX system (such as the DGX H100 or A100), the installation wizard requires the creation of a primary administrative user. This account is pivotal because it is used not only for local OS login but is also synchronized to provide access to the Baseboard Management Controller (BMC) and the GRUB bootloader. NVIDIA best practices emphasize security by mandating the use of a unique, strong password. Using a lower-case username is a standard Linux convention that ensures compatibility across various authentication services. By setting this up correctly during the first boot, the system ensures that 'out-of-band' management (via BMC) and 'pre-boot' configuration (via GRUB) are protected from unauthorized access. Relying on default credentials (Option B) or weak passwords (Option D) is a significant security risk in AI infrastructure, as the BMC often has high-level control over power, firmware, and remote console access.
A system engineer needs to set the vGPU scheduling behavior for all GPUs to share the scheduling equally with the default time slice length. What command should be used?
When deploying NVIDIA vGPU on VMware ESXi, the NVIDIA driver provides several scheduling policies to determine how GPU physical resources are shared among multiple virtual machines. The default behavior is often the 'Best Effort' scheduler, but for environments requiring predictable performance across all users, the 'Equal Share' scheduler is preferred. This scheduler gives each vGPU an equal 'time slice' of the physical GPU's engines. The configuration is managed via module parameters passed to the nvidia kernel driver during host boot. The specific registry key for this behavior is RmPVMRL. Setting RmPVMRL=0x01 enables the Equal Share scheduler (Option A). Conversely, 0x00 would revert to the default time-sliced behavior. It is critical to use system module parameters set to ensure the setting persists across reboots and is applied globally to the NVIDIA driver stack. This ensures that no single 'noisy neighbor' VM can monopolize the GPU cycles, which is a common requirement in shared AI research labs or virtual desktop infrastructures where consistency is more important than raw peak throughput of a single task.
Refer to the output:
~ $ sudo nvsm show healthinfo
---Timestamp: Sat Dec 16 16:26:32 2017 -0800
Checks---BIOS Revision [5.11].........................
DGX Serial Number [YSY72800016)..................
Verify installed DIMM memory sticks........................Healthy
...[output truncated)
Verify Ethernet controllers...........................Healthy
Verify installed GPU's..............................Unhealthy
Checking output of 'lspci' for expected GPU's
Missing GPU at PCI address '07:00.0'
Verify installed InfiniBand controllers....................Healthy
Verify PCIe switches..................................Healthy
...[output truncated)
What insights can a system administrator gain regarding the DGX system's health?
The output provided is a result of the NVIDIA System Management (NVSM) tool, specifically the nvsm show healthinfo command. NVSM is an essential diagnostic framework for NVIDIA DGX systems that monitors hardware health, identifies faults, and helps ensure the system remains within its validated operational state.
In this specific diagnostic trace, the system reports that the 'Verify installed GPU's' check has returned a status of Unhealthy. To provide a root cause, NVSM cross-references the live hardware enumeration from the lspci command against the system's known 'Golden Configuration' (the hardware manifest defined in the firmware). The explicit error message, 'Missing GPU at PCI address '07:00.0'', indicates that the system expects a GPU module to be present at that specific PCIe bus address, but the hardware is not responding or visible to the bus.
This insight allows a system administrator to conclude that a GPU is missing from the logical perspective of the system. This is a critical hardware fault rather than a software or driver issue. In a DGX H100 or A100 system, this could be caused by a physical module failure, a power delivery issue to that specific segment of the GPU baseboard, or a failure in the PCIe switch fabric. Because the DGX relies on a full set of 8 GPUs for high-speed collective communications (NCCL), a single missing GPU will prevent the node from participating in large-scale training jobs, requiring physical inspection or a GPU tray replacement (RMA).
During BCM cluster setup, an engineer must configure bonded network interfaces on DGX nodes for high availability. Which cmsh command sequence properly configures a bond0 interface with two physical NICs?
In NVIDIA Base Command Manager (BCM), the management and storage traffic often requires redundancy via NIC bonding (Link Aggregation). The cmsh utility uses a hierarchical command structure to modify node configurations. To create a bond, the administrator must first navigate to the specific node's interface configuration (device use <node>). The correct sequence involves adding a new interface of type 'bond' (interfaces add bond bond0). After the bond object is created, the physical slave interfaces must be associated with it using the append interfaces command. Finally, the bonding mode (e.g., Mode 1 for Active-Backup or Mode 4 for LACP) and the logical network assignment must be defined. Option B correctly follows this logic. Assigning networks directly to physical ports (Option C) would prevent the use of a unified bond IP, and Option A incorrectly attempts to add a VLAN before the underlying bond is established. This configuration is essential for ensuring the control plane remains reachable even if a single management cable or switch port fails.
Security & Privacy
Satisfied Customers
Committed Service
Money Back Guranteed