The Price of an Unchanged World
As Physical AI gets smarter, when is it cheaper to change the machine - and when is it cheaper to change everything around it?
Frontier technology is making the machine capable of absorbing increasingly more variability. The Second Order question is what that means for everything around the machine - infrastructure, APIs, standards, safety, procurement, asset design, and deployment costs. Better Capability is the fastest race in Physical AI right now - better perception, better hands, better reasoning, bigger data sets, and increasingly general models that learn across tasks and robot bodies.
Open any report or tech post and it likely will be heading to - if we keep making the machine smarter, it will eventually cope with the variable world as it is.
There’s less debate about what had to change around the robot to make its performance survive a normal shift, at a real site, with real edge cases.
But the reality is that every autonomous system is paying for variability somewhere.
There are only a few places where this variability can go
1/ We make the machine absorb it with better perception, reasoning, manipulation and learning. 2/ We make the environment remove it through fixtures, lanes, standard containers or controlled workflows 3/ We clearly define the operation through maps, APIs, permissions, orchestration and human exception handling.
Either a robot can get better at finding an object, or the object can arrive in a known place. We make a vehicle can handle arbitrary traffic, or operate on controlled roads. A mobile robot can learn every lift and doorway, or the building can expose a common interface…We need to figure out where that variability is cheapest, and which is worth turning into machine capability
Here’s examples from two opposite ends of this choice: Autonomous mining and Robotaxis
Autonomous Mining: Komatsu has commissioned 1,000 autonomous ultra- class haul truck with customers moving 11.5B tonnes. BHP’s Escondida Norte operation has 33 autonomous trucks and 11 autonomous drills are operating in an autonomous pit (30% of Escondida’s production), moving 350K +tonnes daily).
While mining isn’t an easy environment to operate in, it’s controllable. The trucks don’t need to solve driving in the general sense. A mining operator controls access, routes, traffic behaviour, fleet interactions, maintenance and exceptions. The mine has narrowed the problem.
Autonomous Driving: Waymo (500K+ trips or 4M miles/week across 10 US cities). is betting on almost the opposite thing. It can’t redesign public streets around its Waymo Driver, so a lot more variability and complexity must sit inside the vehicle and its operating system.
In Warehousing, we see both strategies in one sector. Locus Robotics’ Pick and Piece AMR proposition is based on fitting into existing operations with limited redesign (17,000+ robots across 360+ sites, 7B picks and 190M hours of autonomous navigation). On the other hand, Symbotic redesigns more of the warehouse around repeatable storage and movement of cases/pallets, and uses a system-deployment model (70 deployed systems with over $22.7 billion in contracted backlog).
Locus is absorbing more variation through mobile autonomy and software. Symbotic redesigns more of the warehouse. Both are scaling.
This isn’t a smart robot versus not situation - it’s different ways of distributing complexity. A standard container can simplify manipulation. Controlled traffic systems can reduce planning. Geofences can shrink the behaviour that must be validated…
If changing an environment is cost effective, durable and shared across multiple actions, removing the variability once can be a very good investment decision. If the environment is expensive to alter, changes often, or sits outside the operator’s control, it makes more sense to pay for adaptability in the machine.
Robotics is missing some key metrics: benchmarks versus deployed variability cost
I see advertised a lot of robot counts, autonomous miles, benchmark success, model parameters and pre-training hours. I haven’t found much advertised on what it really takes to convert those capabilities into customer accepted output at a real site.
Variability Tax: Cost created by the real-world variation the system has not designed away. (site adapting + commissioning + task-specific data + human intervention + exception downtime + validation and revalidation + site-specific support) (/) (accepted outputs meeting quality, availability and safety KPIs)
Most of those indicators to the right are private, or scattered between different parties, or not measured in a consistent way.
But the balance is moving: Frontier Physical AI is attacking the machine side of the Variability tax.
The cost of AI capability can fall very quickly. Check Stanford’s 2025 AI Index - it found that the inference cost of GPT-3.5-level performance fell 280x in 18 months.
EgoScale trained on 20K+ hours of egocentric human video and is claiming a 54% per cent improvement over a no pretraining baseline on dexterous manipulation. NVIDIA DreamZero, a World Action Model, is claims over 2X the generalization of VLA baselines on new tasks and environments, and that it can adapt to a new embodiment with 30 minutes of play data. NVIDIA’s Cosmos brings physical reasoning, world generation and action generation into the same model family…
But it’s still early days - new limits being discovered to push through to scale - it’s worth keeping an eye on both - the developments and the constraints – e.g. Arxiv’s research paper highlights inference latency can create stale actions and discontinuities; another of their papers defines a “prediction-deliberation gap - finite-horizon prediction and action generation are not enough for to handle complex embodied tasks that require global planning, cross-stage state maintenance, execution verification, and failure recovery. The pace of research and experimentation is definitely accelerated, but not yet overcome.
In the future will the balance move towards the machine more than today's deployments suggest though? Think of it this way - concrete, steel, conveyor geometry and facility layouts don’t have anything like a software learning curve. It’s not going to get dramatically cheaper to rebuild a loading area or change workstations. Software however, does get cheaper
Now while robotics doesn’t cleanly inherit that - still needs sensors, actuators, batteries, real-time control and much higher reliability than a chatbot (this is a discussion for another day). But the difference matters when considering trade-offs.
So if frontier Physical AI delivers what it is aiming for at scale, the machine side of the equation can change very rapidly. Skills transferring across embodiments, learning new hardware in hours not months, learns exception-handling in generated worlds and plans through unfamiliar, complex multi-step tasks - then we will need far less environment engineering than now. It can wipe out a whole category of demonstrations, local tuning, bespoke integration. A multi-year site redesign may become obsolete before it wears out. Today’s cheapest architecture may not necessarily be tomorrow’s. What will survive that however will still be unified operating state, permissions, common interfaces, , exception handling…
Less physical constraint, more defined operating context
If you’ve been following VDA 5050, it’s Version 3.0 is useful as direction of travel. It now supports more AMRs to plan their own routes. The shared system will still define restricted areas, one- way rules and places where a controller must grant permission but the robot gets more freedom & the environment becomes clearer about the rules.
RoMi-H by Changi General Hospital in Singapore is another example. The hospital is operating 80+ robots across its campus. RoMi-H is a proprietary robotics Middleware that provides a common layer through which robots, building infrastructure, software and IoT systems communicate. It’s robot agnostic and Changi Hospital has made parts of the environment machine readable so every robot does not have to master every lift or door on its own.
BMW’s Figure 02 pilot is in the same vein - to deploy it, BMW also needed production IT, occupational safety, process management and shop-floor logistics. it changed safety arrangements, improved 5G coverage and created standard interfaces.
The direction I see as most important is this shift from hard constraints to softer operating contracts. A defined painted lane tells a robot exactly where it can move. A semantic zone tells it where it may move. A physical barrier prevents an action. A machine- readable rule says an action is not permitted.
But what about Humanoids? A bet against redesign
Humanoids are the strongest counterargument to needing extensive physical redesign - but let’s discuss this.
Many developers are highlighting physical compatibility of the human-type form factor. And in a way, the human body plan is to some extent an installed interface standard.
This does remove some physical retrofits and point integrations. But I want to highlight that physical compatibility is not operational compatibility. Humanoids will still need to rely on mapping, workflow integration, monitoring and fleet management. Being able to press the lift button is not the same as knowing whether you are allowed on the floor.
From a cost perspective, it’s not about Humanoid versus traditional robot - it’s the premium you are paying for adaptability versus the cost of redesign. The more expensive an environment is to change, the more valuable a human compatible machine becomes. The higher the throughput, stability and operator control, the stronger the case for removing the repeating variability once instead of asking every machine to solve it repeatedly.
So where should we be spending the next dollar on Autonomy?
Should we buy a better model? A better sensor or hand? A standardized tote? A redesigned workstation? A semantic map?A common lift interface? Remote assistance? Better validation?
I’ve started to build out a Variability Capture Worksheet (see warehouse example below) that I think through when i get asked this question - where’s the most cost friendly way to handle a variability, and will it survive the future?
Here’s an example for a Brownfield Warehouse Robotics Deployment for Pick and Sort into existing racks, totes, carts, aisles and WMS/WES workflows. The assumption is that the existing warehouse stays mostly intact. We want useful output without rebuilding the site around one robot generation.
The key points are: Don’t ask the robot to solve variability that is cost effective and pointless. Don’t redesign the warehouse to remove variability that is expensive, valuable or likely to become cheaper for the robot to handle. And don’t ask perception to rediscover facts the WMS already knows.
As I come across more of these, I’ll be building out more examples - and I’m looking for more validation and discussion from folks in the space.
This is a co-design problem
I’ve worked with both operators and robotic providers in deployments. These are not decisions operators can make on one side of a procurement table and robot companies will make on the other.
The operator knows which variability is valuable, what the downtime costs are, and what can change. The robot provider knows intervention rates, data requirements and what the next capability roadmap can be. The Integrator knows the real commissioning effort. Safety teams will know the validation needs - And that’s one reason it is so hard to understand the full burden of the variability tax - the data is split across the relationship.
Designing for robots that haven’t been built yet
Machines getting better at coping with physical variability gives asset owners another design headache - It may be a mistake to build tomorrow’s facility too strongly around today’s robot but it may also be a mistake to assume future intelligence will solve every interface on its own - so where is the right middle - ground?
From the earlier point in this article on direction of travel with VDA 5050, I’d focus on autonomy-readiness. Connectivity, digital representations, programmable access, standard equipment interfaces, and machine-readable restrictions that will enable future machines to enter an operation without rebuilding the asset around any one vendor. Being Robot-Ready is about designing an environment that future robots can understand: what state it is in, what capabilities are available, where they may go and what they are allowed to do.
The two trend curves to watch -
How quickly is the cost of machine adaptability falling (very visible)
How quickly the overall Variability Tax of deployment falls (marginal visibility right now)
And that gives us an understanding of which strategies to adopt:
I’m still building my thinking on this topic and I want to hear from folk building, integrating or operating these systems.
What kind of variability is costing you the most today, and how are you handling it? If you are trialing or running these systems in production, what indicators are you using to see if the economics work?
Pallavi Chari Moving Parts






