top of page

The body listed on the exchange; the mind delivered the report card.

Sep 3
10 min read

For Physical AI to touch the factory floor, it requires an operating system fundamentally different from robot VLAs.

150.80 yuan per share. That was the IPO price set by China's Unitree Robotics on Shanghai’s STAR Market. On its August 19 debut, the stock surged over 600%—the moment the public market put a price tag on the "body" of humanoids and quadruped robots.

Less than a month later, on July 30, Google DeepMind unveiled the "mind" to drive those bodies, split into three parts: Gemini Robotics 2, Embodied Reasoning Model ER 2, and the On-Device 2 running directly on the robot hardware.

As bodies become market valuations and minds become quantitative metrics, the question on the factory floor is changing. It is no longer "Can the robot do it?" but "How do we deploy that model onto our equipment, our compliance standards, and our isolated internal networks?"

That is precisely where Mithril’s answer lies, built across more than 25 real-world job sites over the past two and a half years.




1. Valuing Bodies and Minds: Two Markets Unveiled in the Same Month

Unitree’s IPO marked the public market debut of mainland China’s first pure-play humanoid and Physical AI company. With 2025 revenue reaching approximately 1.7 billion yuan and net profit around 280 million yuan, Unitree set a financial benchmark for the capital markets—a standout in an industry where most humanoid startups remain unprofitable.


That same summer, capital flooded not just into hardware bodies, but into robotic brains as well. Figure AI bet big on full-stack integration at a valuation of roughly $39 billion, while Physical Intelligence and Skilled AI pushed hardware-agnostic general-purpose robot models. Xpeng Robotics raised $900 million in a single funding round in August 2026. DeepMind took a third path, simultaneously releasing models and benchmarks to present its report card upfront.


Three distinct tracks have emerged: scaling the hardware body, commercializing general-purpose brains, and opening benchmark ecosystems. Mithril stands on a fourth path: superimposing an intelligence-and-decision layer directly onto operational factory infrastructure—MES, SCADA, FDC, and PLC—without replacing a single piece of existing equipment.




2. Lightbulbs at 92%, Dustpans at 32% — And the Factory's Action Space Is Not Joint Space

The most significant takeaway from DeepMind’s announcement isn't the slick demo videos, but their own transparent success rate matrix:

  • Whole-Body Manipulation (Apollo 2, Inspire Hand)

    • Table pick-and-place: 68.4%

    • Shelf pick-and-place: 76.3%

    • Floor pick-and-place: 45.7%

  • Multi-Fingered Dexterous Manipulation (Apollo 2, Shadow Hand)

    • Unscrewing a lightbulb: 92%

    • Screwing in a lightbulb: 36%

    • Tying a trash bag: 44%

    • Ziploc bag sealing: 40%

    • Sweeping into a dustpan: 32%

  • Gripper Manipulation (Franka Duo)

    • Pick-and-place transfer: 74.2%

    • Tool kitting: 78.9%

    • Precision insertion: 89.6%

Same hand, same model—yet unscrewing a lightbulb succeeds nine times out of ten, while screwing one back in succeeds only once every three tries.

A 56-percentage-point gap. The pattern is clear: success rates stay high in predictable environments with deterministic object locations, but drop sharply when tasks demand substantial posture shifts, fine force control, and precise real-time alignment.


Source: Google DeepMind deepmind.google
Source: Google DeepMind deepmind.google

Gemini Robotics 2 brings whole body intelligence to robots — Google DeepMind

DeepMind's Gemini Robotics 2: Unifying Whole-Body Manipulation and Dexterous Hand Control in a Single Policy


Source: Google DeepMind deepmind.google
Source: Google DeepMind deepmind.google

Picking off a shelf, playing chess, typing on a keyboard, organizing cups, precision insertion—success rates vary wildly from task to task.


This table is honest. But when you try to translate it directly onto the factory floor, a fundamental question emerges:

A robot VLA (Vision-Language-Action model) outputs continuous control commands for a 6-DoF end-effector. In contrast, the commands a manufacturing plant actually receives are temperature, gas flow, RF power, and Recipe IDs.

These four rows represent the structural gap Mithril has mapped out while designing enterprise-grade Manufacturing AX (AI Transformation):

Key Factor

Robot VLA (RT-2 / $\pi$ / GR00T)

Industrial Manufacturing Requirements

Action Space

Continuous control of 6-DoF end-effectors

Recipe parameters, equipment control commands

Sensor Modality

RGB + Depth + Joint state

FDC time-series + Vision + Acoustics / Vibration / Thermal

Control Frequency

30–200 Hz continuous motor commands

Wafer- and batch-level event triggers

Safety & Audit

Collision avoidance, black-box traceability

Explicit rationale behind every decision; audit trails for yield, defects, and occupational safety

This is not because general-purpose robot models are weak; it is because their objects of control are fundamentally different.

The intelligence required to move arm joints and the intelligence required to adjust chamber pressure within strict tolerances may share the label "VLA," but they operate on vastly different inputs, outputs, and operational liabilities.

DeepMind released On-Device 2 because physical bodies cannot afford to wait for cloud latency.

Factories reject the cloud for an additional, even more decisive reason: process data cannot leave the plant perimeter.


Source: VAHLE vahle.com
Source: VAHLE vahle.com

Power & Data Solutions for Cleanroom Automation | VAHLE USA | VAHLE

A Semiconductor Fab: Where Commands Aren't Joints, but Recipes and Interlocks.




3. The Mithril Architecture: One Core, Four Adapters

Seeing Mithril purely as an industrial safety company means seeing only half the picture. We offer four distinct product lines—powered by a single engine.

  • NENYA: A specialized VLA (sVLA) core trained directly on real-world industrial video, sensor metrics, manufacturing processes, and field operations.

  • Guardian-Alpha: Proactive industrial safety—monitoring personnel, hazard zones, and IoT systems.

  • Optima-Alpha / Optima-Vision: Manufacturing operations, equipment anomaly detection, quality control, and defect inspection.

  • Robotics & Physical OS: The physical action layer for robots, drones, and humanoids.

Moving from Guardian to Optima, and from Optima to Robotics, is not a business pivot. It is a systematic expansion—layering domain-specific adapters on top of the same NENYA core.

Perception -> Context -> Evidence -> Action. 

This four-word framework is shared across all four products.

Five structural pillars make this framework succeed on the factory floor:


A Core Built on Industrial Data

NENYA is not merely a fine-tuned chatbot built on an open-source VLM. While it absorbs the latest LLM/VLM methodologies, it remains completely backbone-agnostic. Instead, it stacks proprietary layers—industrial encoders, multimodal fusion engines, inference modules, XAI (Explainable AI), validation gates, and PLC adapters—into a specialized foundation model designed for physical facilities. Where general-purpose models solve "What do I see?", NENYA answers "Given this recipe, this chamber, and this SOP, what is authorized right now?"

Superimposed on Existing Infrastructure

We do not replace MES, SCADA, FDC, or PLC systems. Instead, we layer decision, recommendation, and diagnostic capabilities directly over existing infrastructure—drastically lowering adoption barriers and operational risk. NENYA proposes, deterministic rules validate, and the PLC executes. It is not AI blindly operating industrial machinery.

Runs on Existing Cameras and Air-Gapped Networks

Safety monitoring leverages pre-installed CCTV infrastructure without requiring secondary sensor rollouts, while manufacturing modules ingest live FDC and sensor logs natively. Inference runs entirely on on-premise edge hardware—delivering under 200 ms latency measured on a Jetson Orin NX. Your data never leaves the plant perimeter.


Source:  AG Neovo agneovo.com
Source:  AG Neovo agneovo.com

Choosing The Right Monitor For CCTV And Location-Critical Video Surveillance | AG Neovo Global

In factories with pre-existing CCTV setups and control displays, Mithril's starting point is never to replace those screens—it is to superimpose intelligence directly on top of them.


Decisions Come with Evidence.

Optima's operational closed loop consists of four stages:

Perception (Video · Sensor · Log) -> Decision (Root Cause Reasoning · Action Recommendation)

-> Action (Execution within an Authorized Control Envelope) -> Audit (Evidence Chain)

An alert screen is not the product. The product is the architecture where data flows and decision rationales close the loop. On the factory floor—where an incorrect action leads directly to yield loss and severe industrial accidents—an unexplained shutdown is not a solution; it is a liability.


Validated in the Field, Not in Research Papers.

Two and a half years since our founding in November 2023: deployed across 25+ industrial sites with 23 million cumulative training frames.

  • Hazardous Behavior Detection Recall: 98.7%

  • False Positive Rate: 0.8%

  • Safety Incident Reduction: 94% (Compared to pre-deployment baseline, internal operational metric)


Key Enterprise Clients:

Hanil Hyundai Cement, Sampyo, Sungshin Cement, SPC, AJU, Dongwon Systems, and Hankook Tire.

While our domain expanded from cement to food & beverage, semiconductor equipment, and secondary batteries, the core engine remained unchanged.


Source: Domestic Industrial Helmet Manufacturing Line youtube.com
Source: Domestic Industrial Helmet Manufacturing Line youtube.com

Process of Making Hard Hats. Safety Product Manufacturer in Korea

One of the first things Guardian-Alpha observed: hard hats are not equipment—they are compliance.




4. Between DeepMind’s 200 Examples and Mithril’s 25 Deployment Sites

DeepMind stated that On-Device 2 requires fewer than 200 examples to adapt to a new robot.

200 is a small number—but it is not zero. Even foundation models demand site-specific training examples when confronted with a new physical body.

On the factory floor, this reality hits even harder. Every process and machine variation comes with distinct parameters and failure modes, while operational data remains strictly non-exportable. This is why general-purpose models cannot instantly comprehend industrial environments upon deployment, and why domain adaptation requires lightweight adapters rather than broad zero-shot transfer.

Consequently, Mithril’s scaling strategy is not about building new models from scratch; it is about mounting domain-specific adapters onto the exact same core engine:

  • Guardian Adapter: Monitors hard hats, zone violations, and near-miss collisions, closing the loop via PTZ camera controls and industrial safety signals.

  • Optima-Alpha Adapter: Tracks process drift and equipment anomalies through FDC, PLC, and sensor logs. On semiconductor cleaning and etching lines, it unifies chamber pressure, temperature, gas flow, RF power, and Recipe IDs onto a single interface.

  • Optima-Vision Adapter: Classifies surface scratches, particulate contamination, and pattern defects.

  • Robotics Adapter: Translates this identical decision logic into patrol and manipulation routines for quadruped robots, drones, and humanoids.

In a battery coating deployment, client feedback confirmed a 26% reduction in defect rates, a 31% decrease in equipment downtime, and a 380-basis-point (+3.8%p) yield improvement. In a blind test on semiconductor manufacturing equipment, NENYA Gen1 achieved a perfect 5/5 detection score, outperforming existing systems which scored 2/5.

These metrics originate from actual production lines—not laboratory leaderboards.


Source: TYCORUN tycorun.com
Source: TYCORUN tycorun.com

Lithium ion batteries production - How does it work – TYCORUN

Secondary Battery Production Line: Process drift on a single coating line translates directly into yield.


As you move toward an Autonomous OS, the execution loop becomes even tighter.

An edge GPU embedded directly within the equipment processes sensor feeds, vision data, and system logs locally, targeting end-to-end response times under two seconds—from anomaly detection to mitigation. Raw data never leaves the machine perimeter; only signed model weights and metadata synchronize with internal registries. Federated learning and air-gapped deployment form the core operational framework for this architecture—providing a practical, secure method for updating models in semiconductor fabs where cloud-based OTA (Over-The-Air) updates are strictly prohibited.

NENYA’s model generations directly mirror this liability-conscious architecture:

  • Gen 1 (Perception & Explanation): Focuses on situational understanding and rationale generation. (Currently in production)

  • Gen 2 (Action & Validation Gates): Introduces execution controls, passing strictly validated Typed Actions within authorized control envelopes.

  • Gen 3 (PLC Integration & Feedback Retraining): Enables full-loop PLC interoperability and continuous learning.

While the ultimate goal is fully autonomous manufacturing intelligence, what actually runs on the factory floor today is a tripartite framework: a proposing AI, a validating rule set, and an executing PLC.

Skipping this sequence yields faster demos—and unmitigated liabilities.


Source: Korea JoongAng Daily koreajoongangdaily.com
Source: Korea JoongAng Daily koreajoongangdaily.com

A robot dog spends its days patrolling Posco's steel mill in Pohang, North Gyeongsang, on the lookout for hazards that humans once checked for — and were endangered by. It is just

A Quadruped Robot Patrols a POSCO Steelworks: Once the Hardware Enters the Field, the Remaining Question Is Not the Patrol Itself—It Is Decision and Rationale.


Source: Boston Dynamics bostondynamics.com
Source: Boston Dynamics bostondynamics.com

Industrial Inspection Solutions | Boston Dynamics

An Industrial Inspection Quadruped: Where the Physical OS Attaches Is Not the Showroom, but This Passage.




5. Three Takeaways for the Enterprise RFP Today

1) Shift the Unit of Adoption from Robots to Tasks and Equipment.

DeepMind publishing a benchmark task with a 32% success rate is a step forward. However, procurement documents should never read "Deploy Humanoids." Instead, they must explicitly define: "This specific task in this process, target success rate, interlock triggers upon failure, and who logs the decision rationale." Crucially, validation datasets must remain in the hands of the buyer—not the vendor.

2) Build Evaluation Criteria Around Internal SOPs First.

While benchmarks like ASIMOV-Agentic are necessary, if they fail to test for confined space entry permits, Lockout/Tagout (LOTO), fall protection, conveyor entanglement, or IATF quality traceability, those scores mean nothing on your factory floor. This is precisely why Mithril placed the Evidence Chain as the final, non-negotiable quadrant of its operational closed loop.

3) Stop Treating Legacy Systems as Enemies.

As general-purpose models grow more capable, the enterprise value of pre-installed FDC, CCTV, and PLC infrastructure actually increases. In capital-intensive plants where replacement costs stall AI adoption, retrofit-capable decision layers unlock immediate ROI. This is why Mithril first opened 25 sites using Guardian—and why Optima seamlessly stacks on top of it.




6. Beyond the Numbers

Robot foundation models are now knocking on factory doors with report cards in hand.

Humanoid manufacturers have already proven their valuation in public markets.

The remaining issue,

however, is the void left in between—not the spectacle, but the operation.

The success rate table released alongside Google DeepMind’s Gemini Robotics 2 in July clearly illustrates this void. Using the same robot and the same hands, unscrewing a lightbulb yielded a 92% success rate, while sweeping with a dustpan came in at 32%. Without changing the model or the hardware, changing a single task caused a 56-percentage-point gap.

This number doesn't merely highlight model limitations:

it proves that a single benchmark score cannot guarantee performance in any real-world field.  

The question must be reframed:

"When placing our site's recipes and safety interlocks next to this table, where do we start?"

"For the remaining percentage where failures occur, who stops the system,

and what evidence records that decision?"

Without a decision and control layer to answer this, 92% remains marketing copy,

while 32% becomes the reality on the shop floor.

Mithril’s response is straightforward:

  • Decision Core: NENYA

  • Safety: Guardian

  • Yield: Optima

  • Movement: Physical OS

A unified decision framework spans from cameras beside operators to chamber sensors and patrol robots. This architecture is backed by joint R&D with NVIDIA Inception, TIPS, Chungnam C-STAR, and Korea University's Pattern Recognition & Machine Learning Lab, with international access established via Sunway, TOA, and Elsewedy/AOI.

Between 92% and 32%: Where Does Your Site Stand?

DeepMind's table is a model report card. What it proves is that performance varies by up to 56 percentage points when tasks shift. Factories need answers for the blank space beside that table: the real success rate for a specific process, equipment, and SOP, and who halts operations with what evidence during a failure.

Deploying manufacturing AI extends beyond model selection.

Realizing benchmark numbers requires aligning workflows, telemetry, quality standards, personnel, security policies, and maintenance frameworks.

Otherwise, 92% stays in the lab, and 32% becomes the reality on the line. This is why physical systems that cannot wait for cloud latency and factories that cannot export data co-exist.

Mithril fills this blank space starting from task definition. We design how to layer NENYA's decision engine onto legacy CCTV, FDC, and PLC infrastructure, whether to enforce safety via Guardian or optimize yield via Optima, and how to structure PoC execution and site expansion tailored to local data and regulations.

Existing logs, video feeds, targeted safety risks, and yield goals are all that are needed. Start by assessing deployment points with these assets.

Discuss your manufacturing AI deployment strategy with Mithril.

 
 
 

Comments


Commenting on this post isn't available anymore. Contact the site owner for more info.
bottom of page