Fifteen infrastructure devices across two data centers (the Akron and NDC sites) were assessed against published vendor minimums for AI training and inference. The estate is a healthy, current-generation enterprise virtualization platform — but it was built for transactional and file workloads, not for accelerated computing. No GPU compute, no RDMA/RoCE fabric, and no all-flash AI data path are present today.
The four Ready devices are the two Nutanix HCI clusters (current AOS 7.5.1, AHV GPU-passthrough capable) and the two Rubrik backup clusters (9.4.3) — the platforms a phased AI build would actually land on. The four Needs-Upgrade devices are the two PowerScale (Isilon) NAS clusters and the two PowerProtect Data Domain appliances: their software is supportable but they sit below the AI data-path floor. The seven Blocked devices are the three legacy Fibre Channel switches (no RDMA path, end-of-life switch OS) and four storage systems that are not AI data platforms — three Dell Unity arrays and a QNAP management NAS.
This is not “rip and replace.” ACME Widget already owns a GPU-ready control plane in Nutanix and a current data-protection tier in Rubrik. The gap to an AI-capable estate is three additive investments — GPU compute nodes, an RDMA-capable east-west / storage fabric, and an all-flash hot data tier — not a wholesale refresh.
Visual One Intelligence captures storage, fabric, and HCI telemetry; it does not enumerate server CPU SIMD level, RAM, GPU inventory, or NIC RDMA capability for this estate. GPU and host-OS readiness in this report is assessed at the platform-capability level (what the observed software can support), not from a per-server hardware bill of materials. A targeted compute-collection pass is recommended before procurement — see Section 07.
Every active device observed in the reporting period, with its firmware/OS, dominant media type, capacity utilization, and AI-readiness verdict. Storage utilization and media are drawn from the enterprise device and disk-type telemetry; fabric OS from the SAN switch summary.
| Device | Class | Firmware / OS | Media | Util. | Verdict |
|---|---|---|---|---|---|
| AKR-Isilon | Dell PowerScale (Isilon) NAS | OneFS 9.10.1.4 | SATA | 46% | NEEDS-UPGRADE |
| NDC-Isilon | Dell PowerScale (Isilon) NAS | OneFS 9.10.1.6 | SATA | 63% | NEEDS-UPGRADE |
| Nutanix-AKR | Nutanix HCI (AHV/Prism) | AOS 7.5.1 | SSD | 18% | READY |
| Nutanix-NDC | Nutanix HCI (AHV/Prism) | AOS 7.5.1 | SSD | 28% | READY |
| AKR-Rubrik | Rubrik (data protection) | 9.4.3-31180 | SSD | 43% | READY |
| NDC-Rubrik | Rubrik (data protection) | 9.4.3-31180 | SSD | 56% | READY |
| AKR-DD6800 | Dell PowerProtect Data Domain | DDOS 7.3.0.5 | SATA | 63% | NEEDS-UPGRADE |
| NDC-DD6800 | Dell PowerProtect Data Domain | DDOS 7.3.0.5 | SATA | 56% | NEEDS-UPGRADE |
| AKR-U300 | Dell EMC Unity (block/unified) | Unity OE 5.2.2 | SATA | 88% | BLOCKED |
| NDC-U300 | Dell EMC Unity (block/unified) | Unity OE 5.2.2 | SATA | 90% | BLOCKED |
| AKR-U650 | Dell EMC Unity (block/unified) | Unity OE 5.2.1 | SSD | 11% | BLOCKED |
| NDC-QNAP-MGMT | QNAP QuTS hero (mgmt NAS) | QuTS hero h5.2.6 | SATA | 77% | BLOCKED |
| ACME-DS5300B-SW1 | Brocade FC switch | FOS v7.1.1c | — | — | BLOCKED |
| ACME-DS5300B-SW2 | Brocade FC switch | FOS v7.1.1c | — | — | BLOCKED |
| AKR-FCSW-1 | Cisco MDS FC switch | NX-OS 6.2(11c) | — | — | BLOCKED |
Verdict definitions — Ready: software ≥ AI floor and no missing prerequisite. Needs-Upgrade: below the AI floor but the vendor publishes an upgrade/expansion path on supported hardware. Blocked: no AI data-path role, end-of-life, or a missing prerequisite that cannot be added in place.
Each device class is audited against the attributes that determine whether it can sit on an AI training or inference data path: software floor, media tier, accelerator/GPU path, and RDMA fabric capability. Green passes the floor, red fails it, gray is a legacy/non-AI role.
Fig. 3.1 Nutanix is the strongest AI landing zone in the estate: current AOS, GPU-passthrough/vGPU capable on AHV, and ample free flash. The two gaps — GPU nodes and an RDMA fabric — are additive, not architectural.
Fig. 3.2 OneFS 9.10 is GPUDirect-Storage-capable by version, but Dell publishes 200GbE / InfiniBand front-end and NFS-over-RDMA support only on all-flash F-series nodes (F710/F910) with Mellanox ConnectX adapters. ACME Widget’s SATA capacity nodes make these clusters an excellent archive / data-lake tier, not a hot training target.
Fig. 3.3 The fabric is the hardest blocker. AI east-west (GPU-to-GPU) and GPUDirect Storage require RDMA — NDR InfiniBand or lossless RoCEv2 Ethernet — which this all-Fibre-Channel estate cannot provide. The FC switches are also running end-of-life switch OS (Brocade FOS 7.1, Cisco NX-OS 6.2) and report very high link-reset and loss-of-sync counters.
Fig. 3.4 A training corpus is intellectual property and must be protected. Rubrik (9.4.3) clears the floor and provides an immutable cyber-recovery vault today; the Data Domain pair sits one major version below the DDOS 8.0 floor but is a supportable in-place upgrade.
Three additive investments move ACME Widget from “no AI data path” to “training-capable” — GPU compute, an RDMA fabric, and an all-flash hot tier — plus two supporting hygiene items on OS/driver standardization and corpus protection.
Observed: No GPU compute is present in the telemetry. The two Nutanix clusters run current AOS 7.5.1 with ample free flash (18% and 28% utilized) and are the natural AI landing zone, but carry no accelerators today.
Observed: The entire estate fabric is Fibre Channel: two Brocade switches on FOS v7.1.1c and one Cisco MDS on NX-OS 6.2(11c). No RDMA, RoCE, or InfiniBand path exists. Both switch families run end-of-life switch OS and report very high link-reset and loss-of-sync counters (tens of millions of events), indicating an aging, error-prone fabric.
Observed: Both Isilon clusters run a current, GPUDirect-capable OneFS (9.10.1.4 / 9.10.1.6) but are built entirely on SATA capacity nodes. The high-throughput AI front-end (200GbE / InfiniBand, NFS-over-RDMA) is published only for all-flash F710/F910 nodes with Mellanox ConnectX adapters.
Observed: VOI does not capture guest-OS, NVIDIA driver, or CUDA versions for this estate, so no GPU host OS could be verified. There is no observed Linux GPU-host baseline to assess against.
Observed: Both Data Domain appliances run DDOS 7.3.0.5 — one major version below the DDOS 8.0 AI-archive floor. Deduplication is excellent (17–19× reduction) and capacity is healthy (56–63% used). Rubrik (9.4.3) already provides a current, immutable protection tier.
Observed: Nutanix AOS 7.5.1 and Rubrik 9.4.3 are both current, AI-relevant releases. The HCI clusters are lightly loaded (18% / 28%) with substantial free flash, and the protection tier is immutable-capable today.
Capacity headroom is the immediate operational story. Most of the estate has room, but both Unity OE 5.2.2 arrays are at or near the 80% capacity-alert threshold — a pre-existing risk independent of AI. The Nutanix and Unity all-flash pools are the only flash capacity with meaningful headroom; the NAS and backup tiers are SATA.
Fig. 5.1 Capacity utilization by device. Green = all-flash pools with headroom (the candidate AI flash capacity). Hatched = above the 80% alert threshold (dashed line): both Unity OE 5.2.2 arrays need capacity relief regardless of AI plans.
Observed workload concentrates almost entirely on the Nutanix HCI clusters — 104 distinct hosts/VMs drive roughly 685,000 IOPS in aggregate, dominated by the NDC Nutanix cluster (~652K IOPS, led by SQL Server VMs). That is transactional database I/O, not AI streaming throughput. There is no observed high-bandwidth sequential-read pattern characteristic of a training data path, consistent with the absence of GPU compute. The flash headroom on Nutanix (and the lightly-used all-flash Unity AKR-U650 at 11%) is the only existing flash capacity that could seed an AI proof-of-concept before the F-series tier in R3 lands.
Zero RDMA capability exists in the estate today. This is the single hardest blocker to AI readiness — and unlike compute or storage, it cannot be addressed by upgrading what is installed.
All three SAN switches are Fibre Channel. AI training fabrics require either NDR InfiniBand or lossless RoCEv2 Ethernet for GPU east-west traffic and GPUDirect Storage; Fibre Channel carries block storage only and has no RDMA-to-GPU path. Beyond the architectural mismatch, the installed switches are aging:
| Switch | Vendor / OS | Active Ports | Link-Reset (cum.) | Loss-of-Sync (cum.) | AI Fabric Verdict |
|---|---|---|---|---|---|
| AKR-FCSW-1 | Cisco MDS • NX-OS 6.2(11c) | 14 / 48 | 5,227 | 39,154,720 | BLOCKED |
| ACME-DS5300B-SW1 | Brocade • FOS v7.1.1c | 7 / 80 | 45,612,355 | 2,846 | BLOCKED |
| ACME-DS5300B-SW2 | Brocade • FOS v7.1.1c | 10 / 80 | 50,839,785 | 646,236 | BLOCKED |
The FC switches are (1) the wrong technology for AI — no RDMA/RoCE/IB — and (2) running end-of-life switch OS (Brocade FOS 7.1 dates to the early 2010s; Cisco NX-OS 6.2 similarly) with cumulative link-reset and loss-of-sync counters in the tens of millions. The AI path needs a net-new RDMA Ethernet or InfiniBand fabric (R2); the existing FC, if retained for block storage, warrants its own modernization and error-remediation track.
Visual One Intelligence does not enumerate GPU presence, server CPU model/SIMD level, or RAM for this estate. The compute readiness verdict below is assessed at the platform-capability level (what the observed AOS can support), not from a hardware bill of materials. Do not interpret the absence of GPU rows as confirmation that zero accelerators exist — confirm with a targeted compute-collection pass before procurement.
At the platform level, the Nutanix HCI clusters are the credible AI compute foundation. AOS 7.5.1 is current; AHV supports NVIDIA GPUs in both passthrough and vGPU modes, and Nutanix publishes validated reference designs (GPT-in-a-Box, Enterprise Edge AI) that run distributed PyTorch training and Kubernetes-served inference on the same platform via the NVIDIA GPU Operator. What is missing is the accelerator hardware itself (R1) and the RDMA fabric to feed it (R2).
There is no evidence of CPU-only AI suitability data either (AVX-512 coverage, ≥ 512GB RAM hosts) because that telemetry is not collected. CPU inference for small models remains a fallback, but it cannot be confirmed from VOI data and should be validated during the compute-collection pass.
For mixed inference plus VDI/general workloads, NVIDIA vGPU on AHV maximizes accelerator sharing; for training, GPU passthrough (as in Nutanix’s A100 reference design) gives full-device performance. AHV uses depth-first GPU scheduling and locks all guest memory when a GPU is attached — size cluster RAM accordingly. Note NVIDIA’s platform-wide limitation that virtualization-based security is unsupported with vGPU.
A phased path that sequences the gaps by dependency: prove the concept on existing flash, then add the fabric and storage tier that production training requires, then scale.
Run the compute-collection pass (R4) to establish a true hardware baseline. Stand up an inference/PoC on existing Nutanix flash with a small number of L40S/A100 GPU nodes (R1, initial tranche) and vGPU. Schedule the DDOS 8.x upgrade (R5) and define corpus protection policy. Remediate the worst FC link-error counters and relieve the two Unity arrays at 88–90%.
Deploy the dedicated RDMA fabric — RoCEv2 100/200GbE (or NDR InfiniBand for larger training) (R2). Add all-flash PowerScale F710/F910 node pools with ConnectX front-end NICs and enable NFS-over-RDMA (R3). This is the inflection point: it converts the estate from “inference-PoC-capable” to “training-capable.”
Expand GPU node count for production training, tier the SATA Isilon as the archive/data-lake behind the F-series hot tier, and operationalize GPT-in-a-Box / Enterprise AI on NKP with the NVIDIA GPU Operator, observability, and token governance. Decide the long-term role of the legacy FC SAN (modernize to 64G FC-NVMe or retire as workloads consolidate onto HCI).
Fig. 8.1 Impact / effort placement of the remediation items. R1 (GPU) and R2 (fabric) are the high-impact major projects that gate training; R4 (OS baseline) is a low-effort quick win to do first; R5 (DDOS) and Unity capacity relief are low-effort hygiene.
Same checks, your estate
Every remediation item above started as a line in a configuration file nobody had read.
Infrastructure AI-Readiness Assessment, starts at $2,500. No agents, vendor agnostic.
Inventory, firmware/OS, media type, and capacity utilization were pulled from Visual One Intelligence for client 3360, period P4165 (telemetry collected 2026-05-27): enterprise_storage_all_devices_proc, enterprise_disk_type_by_device_name_proc, enterprise_devices_with_hosts_proc (104 hosts / 534 attachments, IOPS aggregation), and enterprise_switch_proc (fabric OS and error counters). Vendor floors were cross-referenced against a vendor documentation corpus.
Storage: Dell PowerScale OneFS 9.7+ for GPUDirect Storage, with all-flash F-series + ConnectX 200GbE/IB for the RDMA front-end; Dell PowerProtect DDOS 8.0+ for the corpus-archive tier; Rubrik CDM/RSC 9.0+. Compute: Nutanix AOS 6.7+ with AHV GPU passthrough/vGPU; NVIDIA driver 545+/CUDA 12.4+ for H100-class (535/12.2 for A100/L40S). Fabric: NDR InfiniBand or RoCEv2 100/200GbE lossless for training; Fibre Channel is block-storage-only and does not satisfy the AI east-west or GPUDirect path. OS: Ubuntu 22.04 / RHEL 9.2+.
GPU inventory, server CPU/SIMD level, RAM, and NIC RDMA capability are not collected by VOI for this estate. Compute and OS readiness are assessed at platform-capability level only; a targeted compute-collection pass is recommended before procurement.
The host capacity-modelling proc (enterprise_host_capacity_modelling_list_proc) returned upstream errors on repeated attempts; host data-path metrics were derived from enterprise_devices_with_hosts_proc instead.
Dell Unity arrays are not represented in the AI-readiness storage matrix; they are bucketed Blocked for the AI data path on the basis that Unity provides no GPUDirect Storage or RDMA target stack — this is a capability statement, not a fault. Unity remains fully fit for its current transactional role.
One recommendation (R5, DDOS 8.x) is marked [Citation pending] because a DDOS 8.x release-notes source was not present in the corpus at report time; the floor is taken from the internal AI-readiness matrix and should be confirmed against Dell release notes.
No cloud/hyperscaler compute (Azure/AWS GPU SKUs) was observed in VOI for this client; cloud AI readiness was therefore out of scope this period.
Device fully-qualified domain suffixes and WWNs were removed at collection time. For publication, the client company name is fictional; site prefixes (AKR = Akron, NDC) are retained for readability. All findings, software versions, utilization figures, and error counters are reproduced unmodified.
——— END OF ASSESSMENT ———
Visual One Intelligence • Infrastructure AI-Readiness Assessment • ACME Widget • Period P4165 • Telemetry collected 2026-05-27.
Verdicts reflect telemetry and published vendor minimums as of the collection date and should be re-validated at procurement time.