Implementing Cisco Data Center AI Infrastructure — Free Practice Questions
10 free sample questions from a bank of 44, with the correct answers and explanations. No signup required — start practising right now.
1Drag and Drop QuestionRefer to the exhibit. When a new AI backend network that supports Cisco UCS-based GPU servers with NVIDIA Connect-X network adapters is benchmarked, the traffic distribution observed has unacceptable variability across the interfaces. Drag and drop the code snippets from the bottom onto the boxes in the code to configure the switch to resolve the issue.
Answer:
The short version
The fix maps flowlet mode plus aging plus DLB uplink membership plus retained PFC with ECN into the four blanks. The reconstructed exhibit shows an AI backend for UCS GPU servers with NVIDIA adapters where few large RoCEv2 flows polarize classic five tuple ECMP and benchmark variance is unacceptable.
Key concepts in this question
Few elephant RoCEv2 flows defeat static ECMP hashing. Dynamic load balancing watches real transmit utilization and reselects the least used link at each flowlet boundary so uplinks share evenly while PFC with ECN preserves the lossless fabric that RoCEv2 requires.
Why this mapping is correct
Each snippet answers exactly one blank by function. The flowlet mode line enables utilization aware spraying without the reordering risk of per packet mode on ConnectX adapters and the aging line sets the idle gap that starts a new flowlet and the membership line puts the spine-facing ECMP ports under DLB control and the QoS stanza keeps PFC with ECN so losslessness is not sacrificed for balance.
Why the others are wrong
Any cross mapping breaks mode or scope or losslessness. Per packet mode in the mode blank risks reordering that ConnectX benchmark flows do not expect and a static ECMP hash in the mode blank preserves the polarization under test and a server-facing or breakout member in the membership blank steers the wrong links and dropping PFC or ECN from the QoS blank trades congestion loss for balance.
300-640 exam tip
For AI backend variance pick DLB flowlet with aging on spine uplinks and keep PFC with ECN.
2A medium-sized breadmaking operation must deploy a fleet of GPU servers that meet these requirements for their flagship application:- policy managed through cloud-based interface- low latency between specific pairs of GPUs- at least 2 GPUs per server- all GPUs should be same model- Intel CPU- well suited for Inferencing- customer is power consciousWhich server meets the requirements?
C245 M8
X210C M8 with x580 PCIe node
C845A M8
C885A M8
Answer: B
The short version
B — The X210C M8 with x580 node fits the inferencing brief. It combines cloud management, Intel CPUs, matched GPUs, low-latency pairing, and power efficiency.
Key concepts in this question
Intersight management: cloud-based policy, inventory, and lifecycle for UCS servers.
GPU pairing: low-latency operation favors directly linked identical GPUs in one node.
Inferencing profile: favors efficient PCIe GPUs and CPUs rather than maximum training scale.
Why B is correct
The requirements demand cloud-managed policy, at least two identical GPUs with low-latency pairing, Intel CPUs, inferencing suitability, and power awareness. The X210C M8 compute node with the x580 PCIe GPU node satisfies every clause: Intersight management, Intel platform, paired identical GPUs linked for low latency, and an efficiency-oriented inferencing footprint.
Why the others are wrong
A. The C245 M8 platform does not match the specified PCIe GPU node combination and pairing profile.
C. The C845A M8 targets a different large-scale GPU architecture, exceeding the power-conscious inferencing brief.
D. The C885A M8 is an 8-GPU training-class system, oversized for the stated inferencing and efficiency needs.
300-640 exam tip — memory hook
Inferencing plus power-conscious: choose paired PCIe GPUs, not 8-way training chassis.
3Which components are part of the HyperFabric AI solution?
Cisco UCS servers combined with high-speed networking switches and third-party storage
integrated compute and networking hardware optimized for AI workloads, with optional cloud management
Cisco UCS servers with NVIDIA GPUs, switches, optional VAST Data Storage and cloud- managed networking
combination of Cisco networking hardware and software-defined storage designed for AI scalability
Answer: C
The short version
C — HyperFabric AI bundles UCS NVIDIA compute, switching, VAST storage, and cloud networking. That integrated set is the documented solution composition.
Key concepts in this question
HyperFabric AI: Cisco validated fabric pairing compute and network for AI clusters.
UCS NVIDIA nodes: GPU servers supplying the AI compute capacity.
VAST plus cloud management: optional scalable storage with cloud-operated networking.
Why C is correct
The solution is defined as UCS servers with NVIDIA GPUs for compute, Cisco switches for high-throughput east-west fabric, optional VAST Data storage for unstructured AI datasets, and cloud-managed networking for operations. Option C is the only choice naming all four elements together, so it correctly identifies the HyperFabric AI components.
Why the others are wrong
A. Generic third-party storage omits the specified NVIDIA GPU and VAST plus cloud-managed elements.
B. Integrated compute and networking alone is too vague and omits GPUs, VAST, and cloud management.
D. Networking plus software-defined storage omits the UCS NVIDIA compute core of the solution.
300-640 exam tip — memory hook
HyperFabric AI: UCS plus NVIDIA, switches, VAST optional, cloud-managed.
4A Cisco AI infrastructure requires fast convergence after a spine switch failure. Which routing design principle best supports this objective?
Single-path forwarding
Static route redistribution
Equal-cost multipath routing
Layer 2 loop dependency
Answer: C
The short version
C — Equal-cost multipath gives fast spine-failure recovery. Multiple active Layer 3 paths let traffic reconverge immediately when one spine is lost.
Key concepts in this question
Spine-leaf fabric: every leaf connects to every spine for redundant forwarding.
ECMP: installs several equal-cost next hops so one failure leaves others usable.
Convergence: routing withdraws only the failed path instead of recomputing a single path.
Why C is correct
AI fabrics need both high bisection bandwidth and rapid recovery from spine loss. ECMP advertises each leaf prefix through all spines, so every router already holds alternate next hops. When a spine fails, the control plane withdraws that next hop and data continues over survivors, delivering the fastest convergence among the listed designs.
Why the others are wrong
A. Single-path forwarding creates a single point of failure and forces slow recomputation after loss.
B. Static redistribution is manual, error-prone, and converges slower than dynamic ECMP.
D. Layer 2 loop dependency increases blast radius and relies on spanning tree, which converges slower.
300-640 exam tip — memory hook
Spine failure plus fast recovery: think ECMP with all spines active.
5A customer uses Cisco AI Products within Intersight and observes packet drops for RoCEv2 traffic. Which System QoS policy configuration must be deployed to ensure that RoCEv2 traffic is classified as lossless?
Disable Platinum system class.
Enable all QoS system classes.
Disable Gold system class.
Enable Platinum class and uncheck Allow Packet Drops.
Answer: D
The short version
D — RoCEv2 needs the Platinum class set lossless. Enabling Platinum and clearing packet-drop behavior gives RDMA the no-drop lane it requires.
Key concepts in this question
RoCEv2: RDMA over Converged Ethernet, highly sensitive to loss and delay.
System QoS classes: Platinum, Gold, and others map traffic to buffering and scheduling behavior.
Lossless class: configured with no-drop and typically paired with priority flow control.
Why D is correct
Packet drops show RoCEv2 is currently in a lossy class. Intersight System QoS designates Platinum as the priority lossless class for storage and RDMA traffic. Enabling Platinum and unchecking Allow Packet Drops marks it no-drop, so RoCEv2 frames receive lossless treatment and RDMA performance stabilizes.
Why the others are wrong
A. Disabling Platinum removes the very lossless lane RoCEv2 should use.
B. Enabling every class spreads buffers without guaranteeing RoCEv2 lands in a no-drop queue.
C. Disabling Gold does not by itself make Platinum lossless or classify RoCEv2 correctly.
300-640 exam tip — memory hook
RDMA drops: put RoCEv2 in Platinum and make Platinum no-drop.
6A network architect wants to provide high availability for AI compute nodes connected to dual Cisco Nexus switches. Which technology enables active-active uplink forwarding without STP blocking?
vPC
GLBP
VRRP
PVST+
Answer: A
The short version
A — vPC gives dual-homed active-active forwarding without STP blocking. Both Nexus uplinks forward while the peer link handles loop prevention.
Key concepts in this question
vPC: virtual port channel pairing two Nexus switches into one logical endpoint.
Active-active uplinks: server traffic uses both links instead of one blocked port.
STP independence: loop avoidance moves to vPC peer mechanisms rather than port blocking.
Why A is correct
AI compute nodes need full bandwidth and no single-switch outage. vPC lets the node port-channel across two physical Nexus switches that appear as one logical switch, so both uplinks forward simultaneously. Peer-link coordination prevents loops without placing either uplink in spanning-tree blocking, meeting the stated requirement exactly.
Why the others are wrong
B. GLBP provides gateway redundancy for hosts, not loop-free active-active server uplinks.
C. VRRP supplies first-hop router redundancy, not multi-chassis link aggregation.
D. PVST+ is a spanning-tree variant that by design blocks redundant Layer 2 uplinks.
300-640 exam tip — memory hook
Dual Nexus with no blocked uplink: answer vPC.
7An engineer is deploying a Cisco AI POD environment to support large-scale AI workloads. The solution must provide a software-defined storage platform with a Disaggregated Shared- Everything architecture, allow compute nodes to run on Cisco UCS C225 M8 servers managed through Cisco Intersight, and deliver exabyte-scale, all-flash performance with integrated high- speed RDMA and GPU Direct Storage capabilities. Which solution meets the requirements?
NetApp All-Flash FAS (AFF)
Pure Storage FlashBlade
Local NVMe SSDs within each GPU server
VAST Data Universal Storage
Answer: D
The short version
D — VAST Universal Storage meets the exabyte RDMA POD brief. Its disaggregated all-flash design with GPU Direct fits C225 M8 Intersight compute.
Key concepts in this question
Disaggregated Shared-Everything: every controller accesses all media for scale and resilience.
GPU Direct Storage: GPUs read datasets without CPU bounce buffers for AI throughput.
Intersight compute: C225 M8 nodes managed through cloud policy in the AI POD.
Why D is correct
The brief requires software-defined exabyte-scale all-flash storage, Disaggregated Shared-Everything architecture, high-speed RDMA, and GPU Direct Storage alongside Intersight-managed C225 M8 compute. VAST Data Universal Storage is the platform matching that exact combination, while the other options lack the stated architecture or RDMA plus GPU Direct integration.
Why the others are wrong
A. NetApp AFF is enterprise flash but not the Disaggregated Shared-Everything RDMA plus GPU Direct POD component specified.
B. Pure FlashBlade is scale-out flash yet does not match the named architecture and integration brief.
C. Local NVMe in each GPU server is neither shared nor exabyte-scale software-defined storage.
300-640 exam tip — memory hook
AI POD plus DSE plus GPU Direct: choose VAST.
8An engineer must monitor congestion, flow latency, and traffic drops using Cisco Nexus Dashboard and chooses to use the Traffic Analytics feature. Which Nexus Dashboard feature is automatically disabled when Traffic Analytics is enabled?
Delta Analysis
Flow Telemetry
Connectivity Analysis
SNMP Traps
Answer: B
The short version
B — Enabling Traffic Analytics turns off Flow Telemetry. The two collection modes are mutually exclusive on Nexus Dashboard.
Key concepts in this question
Traffic Analytics: aggregated congestion, latency, and drop analysis from switch telemetry.
Flow Telemetry: per-flow export path using the same hardware telemetry resources.
Feature exclusivity: the platform supports one heavy telemetry pipeline at a time.
Why B is correct
Both features consume the same underlying flow telemetry hardware and pipeline capacity. Cisco documents that activating Traffic Analytics for congestion, latency, and drop insight automatically disables standalone Flow Telemetry to avoid resource conflict. Selecting Flow Telemetry therefore identifies the feature sacrificed when Traffic Analytics is enabled.
Why the others are wrong
A. Delta Analysis compares snapshots and is not the conflicting export pipeline.
C. Connectivity Analysis validates reachability paths and remains available alongside Traffic Analytics.
D. SNMP traps use a separate notification path and are not disabled by analytics mode.
300-640 exam tip — memory hook
Analytics on means Flow Telemetry off — one pipeline at a time.
9A Cisco UCS C885A M8 server contains a GPU sled configured with six power supply units (PSUs) in an N+2 redundancy configuration. The workload on the GPU sled is currently drawing a load of 4500W. Using Power Save Mode, how many PSUs can be placed into standby mode to improve power efficiency without compromising redundancy?
0
1
2
3
Answer: C
The short version
C — Two PSUs can standby while keeping N+2 cover for 4500W. Four active supplies carry the load with two redundant units still available.
Key concepts in this question
N+2 redundancy: the sled tolerates two PSU failures without dropping the load.
Power Save Mode: parks unneeded supplies in standby to raise efficiency at light load.
Load math: active supplies must cover 4500W plus the two-failure reserve.
Why C is correct
With six installed supplies and an N+2 policy, four units must remain active to carry 4500W while preserving two redundant units. That leaves exactly two supplies eligible for standby under Power Save Mode. Parking two improves efficiency because the remaining four run at a higher, more efficient utilization without compromising the required redundancy.
Why the others are wrong
A. Zero standby wastes efficiency when redundant capacity is clearly available.
B. One standby is overly conservative and leaves achievable efficiency gains unused.
D. Three standby units would leave only three active, breaking the N+2 reserve for this load.
300-640 exam tip — memory hook
N+2 math: active equals need plus two; standby is the rest.
10Refer to the exhibit. Users report delayed response times and network latency when running certain AI workloads on a specific Cisco UCS Intersight Managed Mode Domain. Based on the fault message and reported symptoms, which troubleshooting action must be taken first to address the network latency issues?
Review the AI workload configuration to optimize resource allocation (CPU, memory, GPU) for the affected applications.
Examine the UCS fabric interconnect configuration for port-channel 11 on Fabric Interconnect B.
Check the physical cabling and SFP+ transceivers on port 11 of switch B for potential hardware failures or incorrect configurations.
Review the pause frame counters for signs of congestion on Fabric Interconnect B Port 11.
Answer: C
The short version
C — Start with cables and optics on Fabric Interconnect B port 11. Physical faults best explain the reported latency plus the port fault message.
Key concepts in this question
Fault-first triage: address the alarmed object before tuning workloads.
Optics and cabling: marginal SFPs and fibers cause errors, pauses, and latency.
Port-channel context: one bad member degrades the whole channel and application response.
Why C is correct
Users report latency on the domain while the fault points at port 11 of Fabric Interconnect B. Layer 1 faults create exactly that pattern through errors and retransmissions. Inspecting cabling and SFP transceivers first is the fastest way to confirm or rule out the alarmed hardware before deeper configuration review.
Why the others are wrong
A. Workload CPU and GPU tuning cannot fix latency caused by a faulted fabric port.
B. Reviewing the port-channel config skips the more likely physical fault indicated by the alarm.
D. Pause counters are useful follow-up data, but they come after verifying the physical path.
300-640 exam tip — memory hook
Fault plus latency on a port: check Layer 1 before Layer 7.