AI-RAN performance gains won’t justify the investment on their own
In sum – what we know:
- The performance gap – Hughes accepts that ecosystem work is delivering “substantial improvements” in areas like cell-edge uplink throughput, but says operators still struggle to fund a build on those gains alone.
- Priority before profit – Running enterprise AI alongside the RAN requires policy-driven orchestration that reserves capacity for the network first, then partitions, monitors and bills what’s left.
- A differentiator hyperscalers lack – Operators can pair AI services with connectivity, spectrum, mobility and thousands of distributed edge sites — but scaling standardized deployments is the unsolved part.
Operators are being told AI will transform the radio access network, with the idea including things like spectral efficiency, energy savings, and automation. But the operators hearing that pitch are stuck on a simpler question. Do today’s performance gains actually justify the investment? At the Intelligent RAN Forum, Rob Hughes, head of wireless marketing at 1Finity, a Fujitsu company, argued they should stop waiting for the head-to-head numbers to mature and instead chase AI-RAN deployments that already have a business case attached — and a customer paying for it.
It’s a sequencing argument more than a technology one, and it slightly changes how operators might think about monetizing infrastructure they’ve already built.
The business case gap
The AI-RAN Alliance draws a distinction between two modes. There’s AI for RAN, which means using AI to optimize the network itself — improving cell-edge uplink throughput is his example. Then there’s AI RAN, where RAN infrastructure doubles as a platform to serve AI workloads for enterprises and other customers. Hughes is focused on the second, though he doesn’t dismiss the first.
1Finity has worked with ecosystem partners for years on RAN optimization use cases and, according to Hughes, has seen “substantial improvements.” But he concedes that operators are “struggling to justify the investments based on the performance improvements that are available today.” The gains are real, in other words, but not yet large enough to carry the capital case on their own. His advice is blunt. Performance improvements are great, “but where are they now?” Rather than waiting for that side of the equation to mature, MNOs should start learning and engaging with AI-RAN through opportunity-driven deployments — ones where there’s “an immediate business case and more importantly a paying customer.”
There’s a structural reason this makes sense. RAN networks are largely designed for peak capacity, which means the compute sitting behind them is heavily underutilized most of the time. That’s wasted headroom. Hughes’s argument is that pooling and reselling that excess capacity gives operators a revenue stream from infrastructure they’re already powering and maintaining. “Don’t plan to put GPUs everywhere day one. We’ll get there,” Hughes said. The shift he’s advocating isn’t a speculative nationwide GPU rollout. It’s deploying where demand already justifies the spend.
Orchestration, not hope
If operators are going to run enterprise AI workloads alongside the RAN on shared infrastructure, the obvious question is how you keep those workloads from degrading network performance. Hughes rejects the idea of throwing everything onto one server pool and hoping for the best.
What’s required, he argues, is policy-driven orchestration that treats the RAN as the priority workload at all times. That means continuous resource monitoring, reservation, and prioritization — so the network meets its performance objectives before any other workload consumes capacity. It’s a straightforward principle, but the implementation isn’t trivial. Beyond the technical layer, operators need to partition between tenants, monitor consumption, and bill accordingly. Not every enterprise customer will want the same tier of service. One might pay for dedicated priority capacity, while another takes best-effort or off-peak access. Hughes argues you need both options to actually soak up the excess compute.
The more interesting part of the pitch is the contrast with hyperscalers. AWS, Azure, and Google Cloud can obviously provide AI services at scale, and they’ve spent billions building out the infrastructure to do it. But Hughes names a set of advantages MNOs hold that the hyperscalers don’t. Last-mile connectivity and transport to the device. Spectrum and wireless connectivity. Mobility and flexibility for moving devices and sensors. Reliability for business-critical AI applications. Data sovereignty, which he framed as mattering where data has to stay on site or in country. And — probably the most tangible differentiator — thousands of distributed edge locations with existing real estate and power, set against the regional cloud locations hyperscalers build on.
“The hyperscalers can provide the AI services, but MNOs can combine the AI services with the connectivity, the mobility, spectrum, transport, distributed edge,” Hughes said. The bundling argument isn’t new, but in this context it carries a specific implication. If an operator can attach high-value AI services to the connectivity it’s already selling, it improves stickiness for the existing business. That’s a different value proposition than trying to out-compete a hyperscaler on raw compute.
Whether MNOs can actually execute on that bundling at scale is a separate question. The telecom industry’s track record with edge computing hasn’t exactly been spotless, and hyperscalers have the software ecosystems, developer relationships, and go-to-market muscle that operators generally lack. Still, the physical infrastructure argument is hard to dismiss entirely.
Mass customization
Hughes is candid about what killed earlier mobile-edge-computing efforts. Customization. Every deployment required bespoke integration work. “You need to customize every single time. That’s far too time consuming,” he said. If AI-RAN enterprise services follow the same path, they won’t scale.
1Finity’s answer comes in two parts. The simpler play is GPU-as-a-service — selling access to compute capacity on shared infrastructure, which requires relatively little customization on the operator’s side. For smaller enterprises that also need wireless connectivity and AI agents configured, 1Finity works with partner Aible on agent templates. Hughes claims Aible offers over 100 standard application agents that enterprises can adapt through guided workflows requiring no coding. The idea is that the enterprise customizes for itself rather than loading that work onto the operator — a self-service model that reduces deployment friction.
In a demo, Hughes walked through a concrete example. A shrinkage and shoplifting detection template connects store CCTV to the AI platform. The hardware chain runs from camera to CPE to an Open RAN-compliant radio, through a switch, and into a pool of three NVIDIA GPUs on Quanta servers. Those servers simultaneously run the RAN stack (CU, DU, and 5G core), the Aible AI platform and application, and a third workload reserved for AI-as-a-service. Hughes made a point of noting that the radio “doesn’t have to be 1Finity” — an open-ecosystem signal aimed at operators wary of vendor lock-in.
One detail worth flagging is the orchestration layer underneath. Partner Armada’s resource management reaches into the CU and DU to pull RAN-specific KPIs — connected UEs, throughput per UE — rather than just reporting GPU utilization. Generic compute metrics won’t tell you whether adding another AI workload is about to degrade a subscriber’s experience. RAN-aware monitoring gives policy engines the data to anticipate that impact before it hits.
Hughes pointed to two further ways to sell that capacity beyond the direct enterprise deployment. The first is the GPU-as-a-service play from earlier, aimed at large enterprises that design for nominal load and need burst capacity — or teams that run out of AI compute budget mid-cycle and need a bridge. The second is AI-as-a-service for developers and academics who need compute briefly, with GPU slicing and controls on tokens per minute and price per token.
The strategic framing is that all of this is available now. 1Finity’s pitch isn’t about waiting for standards bodies or next-generation silicon. It’s about giving operators a practical on-ramp — monetize shared infrastructure, build in-house operational expertise with real deployments, and then extend toward AI-native public networks as the technology and demand curve mature.
It’s a compelling sequence, at least on paper. AI-RAN doesn’t need a nationwide build to be real. It needs one enterprise with a check.
Table of Contents