the numbers first :
every number below is linked to its primary source, because this corner of the internet is full of statistics that evaporate when you chase them (more on that in a minute).
- unplanned downtime costs the world's 500 biggest companies $1.4 trillion a year, 11% of their revenues, per Siemens' True Cost of Downtime 2024. in automotive, one unproductive hour now costs $2.3 million, more than $600 a second.
- more than 70% of industrial companies are stuck in "pilot purgatory", technology deployed experimentally at reduced scale that never reaches production, per the World Economic Forum with McKinsey.
- the canonical peer-reviewed field study (Carlson & Murphy, IEEE Transactions on Robotics) measured fielded mobile robots at 6 to 20 hours mean time between failures, with availability below 50%. fielded robots spent half their lives broken.
- data from more than 400 factories shows robot cells stop on average every 87 minutes, and 80% of those stops are not the robot's fault : conveyors, sensors, peripherals (The Robot Report, citing IJPE data).
- the spec-sheet-versus-floor gap in one pair of numbers : FANUC rates its hardware at 100,000 hours MTBF; a measured automotive paint line with 156 robots found its worst performer at 5,280 hours. both numbers are true. they measure different worlds.
- and even the success stories fail constantly : a museum robot that ran 446 days and 1,300 km autonomously still logged failures its authors called "virtually impossible to avoid."
as a16z frames the learned-policy version of the same story : a policy at 95% in the lab might drop to 60% in deployment, against production expectations above 99.9%. at 95%, the robot fails 50 times a day, and every failure summons a human.
about that study everyone quotes :
while researching this article we chased a statistic that circulates in robotics decks and blog posts : that an academic study found "45 of 75 real-world robot deployments failed."
we could not find it. no paper, no dataset, no primary source. our best guess is that it's a garbled mashup of two unrelated percentages from the 400-factory dataset above. if you have seen this number cited, check the citation : it points at another blog post, which points at another blog post.
this matters beyond pedantry. an industry that quotes made-up failure numbers is an industry that has not measured its real ones. the real, verifiable numbers above are bad enough.
failure cause 1 : the field is not the lab :
Zume raised over $445 million to make pizza with robots, and the robots made pizza. then the pizzas got in the delivery trucks, and the cheese slid off when the trucks turned. baking in a moving vehicle, the entire premise, had never been validated against the physics of an actual delivery route.
the Henn na Hotel in Japan, the Guinness-certified first robot-staffed hotel, fired half of its 243 robots in 2019. its in-room assistant interpreted one guest's snoring as voice commands and kept waking him up to ask what he wanted. an unbounded input distribution, real humans, doing human things, met a system validated on a narrow one.
this is the same failure at two price points : the deployment environment was outside the distribution the system was tested in. it is the oldest finding in field robotics, Carlson & Murphy flagged hardware "designed and tested for a narrow range of environmental conditions" twenty years ago, and it is the default failure mode of learned policies today.
failure cause 2 : reliability economics :
Walmart's shelf-scanning robots worked. they roamed roughly 500 stores auditing inventory, on contract toward 1,000. then the pandemic filled the aisles with human workers picking online orders, and Walmart noticed those workers could check shelves as a byproduct. the contract ended in november 2020, with Walmart citing "different, sometimes simpler solutions that proved just as useful."
october 2022 delivered the same lesson twice in one month : Amazon shut down its Scout sidewalk delivery program (~400 people reassigned) and FedEx killed its Roxo delivery bot, stating it "did not meet necessary near-term value requirements." when capital got expensive, the programs that could not show a reliability-adjusted return died first.
the lesson is uncomfortable but useful : a robot does not compete with "no robot", it competes with the cheapest alternative under whatever conditions exist at renewal time. the only defense is a defensible number, this reliability, at this cost, versus that alternative, which is exactly the number most deployments never measure.
failure cause 3 : the demo writes checks the fleet can't cash :
Pepper, SoftBank's humanoid greeter, was the most charismatic demo in robotics. it scaled to about 27,000 units before production was quietly halted, with Reuters citing weak demand, limited functionality, and reliability issues. it never had one job it did reliably enough to be rehired for.
the pattern generalizes : a demo is a claim about the best case, and a deployment is a promise about the average case, including the bad days. every gap between those two is paid back with interest in the field, in front of the customer.
at the Henn na Hotel, the one robot that kept its job was the boring one : a mechanical luggage-storage arm doing a narrow, structured, bounded task. there's a whole deployment philosophy in that single staffing decision.
failure cause 4 : the hidden payroll :
every deployed robot carries an invisible payroll : the humans who reset it, patch around it, and answer its calls for help. the Henn na staff said the quiet part out loud after the robot layoffs : "it's easier now that we're not being frequently called by guests to help with problems with the robots."
the 400-factory data makes the same point quantitatively : a stop every 87 minutes, with 80% of stops caused by the systems around the robot. nobody budgets for that at pilot time, because at pilot scale an engineer is standing right there and the interventions feel free. at fleet scale they are the single biggest line item nobody measured.
intervention rate, how often a human steps in, and what it costs when they do, is the metric that decides whether a deployment survives its first contract renewal. we make the case for tracking it from day one in robot deployment reliability.
the one that worked :
DHL started with a single-site Locus Robotics pilot in 2017. nine years later the partnership passed one billion picks, with thousands of robots across 40+ facilities and an agreement for 5,000 more. DHL's supply chain CEO summarized the whole discipline in one line : "an idea is only a good idea if it can scale."
notice what the success has that the failures lacked : a bounded environment (a warehouse, not a sidewalk), human-in-the-loop task design, scaling gated on measured productivity at every step, and a fleet-management layer that treats reliability as the product. none of that is luck. all of it is measurable.
the common thread :
read the post-mortems again and the pattern is hard to unsee : the robots mostly worked. what failed was the absence of honest measurement, of field conditions before deploying (testing), of whether a new version still clears the bar (evaluation), and of what it actually costs to keep the fleet running (the deployment gap).
the companies above could absorb a failed robotics program. the robotics teams reading this mostly cannot. measuring reliability honestly, per condition, from the first pilot, is the cheapest insurance in the industry.
frequently asked questions :
Why do robot deployments fail ?
Documented post-mortems cluster around four causes : the deployment environment differs from the conditions the robot was validated in (distribution shift), the robot works but does not beat the cheapest alternative on cost (reliability economics), the demo sets expectations the fleet cannot sustain, and the hidden human labor of keeping robots running erases the savings. Almost none fail because the core robotics research was wrong.
What percentage of robotics pilots reach production ?
The World Economic Forum, in collaboration with McKinsey, found that more than 70% of industrial companies are stuck in 'pilot purgatory' : technology deployed experimentally at reduced scale that never rolls out at production scale. Robotics pilots follow the same pattern, which is why gated, measured scaling is the discipline that separates the survivors.
How reliable are industrial robots really ?
The spec sheets and the floor tell different stories. Manufacturers rate robot hardware at up to 100,000 hours MTBF, and one measured automotive paint line found even its worst robot at 5,280 hours. But at the cell level, data from more than 400 factories shows a stop on average every 87 minutes, and 80% of those stops are caused by things around the robot, conveyors, sensors, peripherals, not the robot itself. Reliability is a property of the system, not the arm.
What did failed robot deployments have in common ?
The robots mostly worked. Walmart's shelf scanners scanned shelves, Zume's robots made pizzas, Pepper greeted customers. What failed was everything around the demo : unvalidated field conditions, economics that stopped clearing the bar when conditions changed, and maintenance burdens nobody had measured. The common fix is measuring reliability honestly, per condition, before and during deployment.
deployments don't fail because robotics is hard. they fail because nobody was measuring the things that kill deployments.
we are building the infrastructure that measures them : capture every run, score reliability per condition, and know what changed before the customer does.
Get a free audit today