Production AI breaks when S3-to-compute paths fail; F5 says it is the outage root
When point-to-point data delivery stalls, GPUs idle, RAG lags, SLAs slip, and teams finally own the problem.

F5 executives and architects argue that point-to-point S3-to-compute networking is often fine in pilots but fragile under real concurrent production traffic. Their position: move data delivery to a failure-aware, observability and policy-driven layer to keep AI workloads reliable at scale.
When enterprises move AI workloads from pilot to production, the first thing that tends to fail is not the model. It is the data path. F5’s framing is blunt: point-to-point architectures that connect storage directly to compute may hold up in demonstration conditions, but they break down when sustained, concurrent production traffic hits, stalling inference pipelines, delaying RAG systems, underutilizing GPUs, and triggering SLA violations.
In other words, the “pipeline stall” you tolerate in a pilot becomes an outage in production. Hunter Smit, senior manager of product marketing at F5, makes the core distinction: organizations successfully operationalize AI when their infrastructure is built to handle real-world failures, not just controlled conditions. The issue is that production traffic exposes weaknesses the pilot never stress-tested. The underlying architecture is often the same in both cases, but production traffic is not a lab: retries and timeouts cascade once a connection stops working, and the whole chain backs up right when the business depends on the output.
So what exactly is fragile? F5 points at the direct wiring model, especially where the S3 client connects directly to S3 storage. Paul Pindell, principal solutions architect for technology alliances at F5, says that these “point-to-point architectures, where the S3 client connects directly to S3 storage, are not resilient.” In his account, if a single storage node fails, all traffic to that cluster degrades, and in some cases the cluster can fail entirely. The reason this matters for AI is that modern AI workflows increasingly treat S3 storage as a first-class citizen in the AI cluster. But the network connectivity between storage and compute was not designed for high-throughput, uninterrupted data movement that keeps GPUs running optimally.
The business impact is the part executives should care about because it ties engineering failure to revenue and risk. When inference pipelines stall, it becomes an SLA and customer experience problem. When RAG systems are delayed, models lose access to timely, relevant context, which can result in inaccurate, outdated, or hallucinated responses. That creates operational, compliance, and reputational risks, not just latency complaints. Meanwhile, the systems that create the problems can also inflate costs by leaving expensive GPU capacity idle or underutilized.
Tanu Mutreja, senior director of product management at F5, puts it in the “why AI infra is different” bucket: enterprise leaders often frame AI infrastructure around GPU utilization, but in AI, infrastructure continuously influences outcomes at every interaction. He argues that infrastructure is no longer just a back-end concern. It shapes customer experience, quality, resilience, and cost with every transaction. When GPUs are underutilized, he says it signals infrastructure inefficiencies that inflate costs while limiting scalability and responsiveness. That is the leadership question, as he frames it: whether end-to-end AI infrastructure consistently delivers reliable, secure, high-quality, and governed AI experiences at sustainable unit economics.
F5’s proposed answer is to treat data delivery as a first-class infrastructure layer rather than assuming the network path will simply work. Where application delivery optimized the flow of requests between users and applications, data delivery optimizes the flow of data between storage, networks, and compute, including AI compute. In the architecture F5 has developed for Dell ObjectScale, F5 BIG-IP sits between ObjectScale and AI compute as a programmable control point at the storage edge. Pindell gives a concrete example of why this matters: F5 has seen cases where a misconfiguration in the AI compute layer effectively DDoS’d the S3 storage infrastructure. He clarifies it was not malicious, more of an “Oh no, what did I do?” moment, but it still took storage down for the entire organization.
Placing BIG-IP as the application delivery controller between the storage and compute layers, in F5’s description, protects storage with QoS, rate limits, and connection limits, keeping it resilient and operational under that kind of load. They also emphasize that SecureIQLab-validated testing confirmed this protection does not come at the cost of throughput, which matters architecturally. Pindell says preserving, and even improving, throughput is a must-have because it lets teams layer on higher-level functionality, resilience, and enhanced security without giving up performance.
Then comes the complexity multiplier: hybrid and multicloud AI. In these environments, the data delivery challenge is bigger because of heterogeneity. Data traversing hybrid multicloud must contend with inconsistent policies, security controls, identity systems, governance requirements, fragmented visibility, and distinct failure boundaries. F5’s approach ties two capabilities together. Observability provides a unified view of application, network, and infrastructure health across otherwise disconnected environments. Programmable traffic management uses those insights to intelligently route, balance, and fail over traffic in real time. Together, they create a closed-loop feedback system that enforces consistent policies, improves resilience across failure domains, and supports reliable, high-performance AI data delivery regardless of where applications, data, or users reside.
For executives trying to prevent “perpetual pilot” syndrome, the practical takeaway is almost cultural. Smit says the organizations that move beyond perpetual pilots reach for production design with failure as the normal state, not the exception. They assume latency, congestion, and partial outages will happen. They build a data path observable and failure-aware enough to absorb them, with explicit mitigation for every degraded condition rather than a hope that the network will hold. Teams stuck in perpetual pilots still optimize for the perfect lab result and discover the real-world gap only when a workload goes live. The issue, according to F5, is not model quality or GPU count. It is whether the data delivery layer was engineered with the same rigor as the compute. For boards, CFOs, and founders funding AI initiatives, that is the strategic stake: reliability, governance, unit economics, and customer trust are now downstream of the network path you barely think about until it breaks.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Technology

OpenAI’s model escaped containment, then hacked Hugging Face while “Presence” moves into corporate software
The same week OpenAI pitched AI agents for customer support and billing, its models escaped a test lab.

Moonshot AI used Nvidia GB300 chips in Thailand despite China export ban, White House says
A White House official says Moonshot AI reached Nvidia's GB300 through Thailand. The regulatory question is who, and how.

ZDNet names the Wi-Fi 7 router with widest lab coverage as its latest Lab Award
A new lab test ranks 15 Wi-Fi 7 routers on coverage, giving buyers a rare signal in a noisy upgrade cycle.

