Physical AI is forcing the technology industry to rethink your entire computing stack.
Robots, autonomous systems and intelligent devices need economical inference, secure data access and infrastructure that works beyond conventional clouds. Rafay Systems is addressing those demands through orchestration that lets providers offer open models without dedicating entire GPU systems to individual customers, potentially lowering enterprise AI costs while protecting sensitive information, in keeping with Haseeb Budhani (pictured, right), co-founder and chief executive officer of Rafay Systems Inc.
“The serverless part on the infrastructure level, we’ve solved that … we’ve done for some time,” Budhani said. “Now we’ve got a slice of essentially a confidential GPU. I can now have higher security on top of the GPUs that I have already got, which implies as a provider, I get to monetize every second of my infrastructure. But, for the enterprise, it’s actually a greater deal. They like that because they give the impression of being at the whole cost of ownership, they give the impression of being on the capex they usually don’t have the large money. It’s too expensive, otherwise. This enables them to have a secure solution, which actually delivers a greater price point. That’s what everybody wants.”
Budhani, Eiman Ebrahimi (left), CEO of Protopia AI Inc., and other industry leaders spoke with theCUBE’s John Furrier and guest host Howie Xu, chief AI & innovation officer of Gen Digital Inc., as a part of theCUBE + NYSE Wired: Robotics & AI Infra Leaders event, during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They examined the infrastructure, economics and technology required to move AI into the physical world. The event builds on theCUBE’s Machina AI Summit coverage of robotics moving from demonstrations into production. (* Disclosure below.)
Physical AI changes the infrastructure equation
The economics behind physical AI begin with tokens. Agentic applications eat much more tokens than easy chat interactions because every latest motion may require previous context to be processed again. That makes token volume, model selection and workload design increasingly essential measures of infrastructure demand, explained Max Kan, tokenomics technical lead at SemiAnalysis LLC.
“On the whole, agentic workloads eat exponentially more tokens than the chat workloads everyone was using a yr ago since it’s all inherently multi-turn,” he said. “And each time you ask your agent a follow-up or it uses some tool like web search or perhaps it searches your code base, it has to reprocess all previous tokens in your conversation, which really causes total tokens produced/processed to go up exponentially.”
The inference market is splitting into specialized workloads, with prefill demanding compute, while decode is determined by memory capability and bandwidth. Positron AI Inc. is targeting those requirements through systems optimized for tokens per dollar and tokens per watt, including air-cooled designs suited to enterprise data centers that can’t support liquid-cooled racks, in keeping with Darren Chien, managing director of APAC at Positron AI.
“We consider that inference is fundamentally an economics problem,” he said. “We’d like to ensure that these tokens are generated as cheaply and as widely available to users as possible. The best way we compete is on optimizing the 2 metrics of tokens per dollar and tokens per watt.”
As physical AI moves into factories, vehicles and distributed environments, infrastructure providers need greater than faster accelerators. Axiado Corp. addresses the resulting demands for power, security and system control through silicon-based platform management that adjusts cooling, frequency and voltage by workload while reducing operational pressure on CPUs and GPUs at scale, noted Gopi Sirineni, founder, president and CEO of Axiado.
“We’re bringing that intelligence right into a silicon level, and it would be offloaded off of your GPUs, CPUs,” he said. “We’re monitoring the system itself, including the fans, liquid cooling and all that. Primary, we manage the efficiency by managing the fans and fan controls and liquid cooling … after which, two, we do dynamic frequency and voltage scaling for the platform.”
Intelligence moves closer to the physical world
Physical AI increasingly is determined by compact models that may run locally despite limits on power, memory and connectivity. Liquid AI Inc. is developing customizable foundation models for laptops, vehicles and other devices, allowing specialized intelligence to cut back inference costs, improve privacy and deliver faster responses without relying entirely on centralized data centers, identified Ramin Hasani, co-founder and CEO of Liquid AI Inc.
“One small model will not be going to be generally intelligent,” he told theCUBE. “What they’ll do, they might be fine-tuned. You may customize them. The good thing about this thing is that customizing a small model, the price of it will be shockingly low.”
As intelligence moves onto devices, networking becomes a part of the physical AI application architecture. Aria Networks Inc. applies specialized AI to high-resolution telemetry, helping operators detect failures and performance issues before they disrupt costly workloads. The goal is to strengthen human decision-making with timely context, not remove people from network operations, in keeping with Mansour Karam, founder and CEO of Aria Networks.
“It starts with collecting telemetry. There are two points to collecting telemetry,” he said. “You should be collecting telemetry on the microsecond resolution. That’s primary. You wish it at the proper resolution. Number two, the network, and we will speak about all different networks in an AI factory, nevertheless it spans very different domains.”
Enterprise AI is moving beyond chatbots toward computer-use agents that may operate existing software, capture workflows and execute routine tasks without lengthy integrations. H Company is targeting that shift by placing a human-centered intelligence layer over legacy systems, helping organizations automate repetitive work while keeping consequential actions under human control, emphasized Gautier Cloix, CEO of H Company.
“What’s hard is to do it with the security part, without the danger,” he said. “That’s the identical thing for enterprise. The best way we built our agents is that they can’t execute irreversible tasks – sending an email, deleting something, ordering something – and not using a human validation at first … that’s where we focus lots of energy, ensuring that deterministically our agents cannot do the flawed things.”
Latest markets emerge around compute capability
AI’s rapid growth is creating demand for financial tools that may manage volatile GPU pricing and availability. Silicon Data and The Compute Exchange Inc. are developing compute benchmarks and instruments that would help infrastructure operators hedge falling prices while protecting major AI consumers against rising costs as demand shifts over time, explained Carmen Li, founder and CEO of Silicon Data and CEO of The Compute Exchange.
“I’m helping probably a dozen market participants arrange their compute desks. Persons are actually helping their clients to hedge,” she said. “Take into consideration in case your client, in case you’re a bank, your client might be the hyperscaler or neoclouds, they’ve a protracted exposure, they struggle to assist them manage the futures and volatility for his or her revenues. Or … your client might be the huge AI company, they’re going to eat lots of GPU tokens. They assist them manage the short exposure.”
Compute markets have gotten more distributed as organizations mix owned infrastructure with capability from neoclouds and regional providers. San Francisco Compute Co. is constructing a marketplace for GPU capability that supports this shift, using secure multi-tenancy to match infrastructure with workload, location and availability needs without exposing customer data, emphasized Alan Butler, chief business officer of San Francisco Compute Co.
“That was the genesis of getting a compute layer where you may rent out on short-term provisioning that became then the premise of, in essence, a marketplace,” he said. “There’s lots of marketplace corporations which are flipping GPUs over a fence with no SLAs behind it. And the whole lot we do has an SLA standing behind it with an information center provider.”
Edge deployments require infrastructure that runs models near where data is created without forcing enterprises to rebuild data centers. Axelera AI B.V. is extending its edge-efficient architecture into servers and cloud systems, using familiar software frameworks so developers can adopt latest hardware while avoiding proprietary application redesigns and reducing deployment barriers.
“They’re designed together very intentionally,” said Alexis Crowell, chief marketing officer and general manager of the Americas at Axelera AI. “We have now as many of us working on the software side of our business as we do on the hardware side. It’s not that I’m attempting to monetize the software. I’m attempting to … ensure that developers have a frictionless experience.”
To observe more of theCUBE’s coverage of theCUBE + NYSE Wired: Robotics & AI Infra Leaders event, here’s our complete video playlist:
https://www.youtube.com/watch?v=videoseries
(* Disclosure: Neither ScaleFlux Inc., the presenting sponsor of theCUBE + NYSE Wired: Robotics & AI Infra Leaders event, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)
Photo: SiliconANGLE
Support our mission to maintain content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.
- 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
- 11.4k+ theCUBE alumni — Connect with greater than 11,400 tech and business leaders shaping the longer term through a novel trusted-based network.
About SiliconANGLE Media
Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our latest proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to assist technology corporations make data-driven decisions and stay on the forefront of industry conversations.

