XPENG

How XPENG IRON works

XPENG IRON runs VLM and VLT on top of its own sensing and whole-body control — the parts the 4 public sources on this page actually name.

IRON is described with three named model layers: one that understands what it sees and hears, one that decides the task steps, and one that predicts the physical world and turns that into actions.

07 named models · 4 sources · updated 9 Aug 2026

XPENG IRON humanoid robot
Image: XPENG

What XPENG IRON uses to sense, think, move and learn

Senses

Turns cameras, audio and touch into one picture

  • VLM

Understands

Works out what the job is

  • VLM
  • VLT: Vision-Language-Task

Predicts

Anticipates what happens next

  • XPENG VLA 2.0
  • X-World
  • X-Foresight
  • X-Mind

Acts

Chooses the movement and the grasp

  • XPENG VLA 2.0
  • VLA 2.0 cross-domain model

Controls

Keeps the body balanced while it works

  • Model not named

Learns

Improves the models between runs

  • X-World
  • X-Foresight
  • X-Mind

What it learns feeds back into the models before the next run.

Greyed cards are parts the vendor hasn't named a model for.

01

How the robot is built to work

XPENG also has separate car-side world-model work that is not confirmed to run on IRON itself.

Technical read

XPENG states IRON combines VLM (vision-language understanding), VLT (Vision-Language-Task reasoning and decision-making) and VLA 2.0, which XPENG describes as both an action-generative model and a physical-world model spanning cars, humanoids and flying vehicles.

  • Named VLM/VLT/VLA 2.0 stack
  • Cross-domain physical-world model
  • High DoF body
12
System parts on the public record

4 confirmed as running on the robot

09
Parts with a public model name

0 supplied by a partner

04
Primary architecture sources

13 open questions left unanswered

XPENG is architecturally interesting because it discloses VLM/VLT/VLA 2.0 around IRON and separately discloses world-model work. Keep X-World/X-Mind as XPENG physical-AI context unless directly tied to IRON runtime.

02

How it senses, reasons and moves

Four-level chain

Understanding, task reasoning, prediction and action are named as separate models handing off in sequence.

Skin and body perception

It has a skin-like surface and AI to understand people and space.

Model not named

Run the models

IRON likely has strong onboard compute, but exact TOPS should be quoted carefully.

  • Turing AI chips

Contact safety

Skin could help contact awareness, but hard safety proof is not public.

Model not named

See and understand language

This is the part that understands what it sees and what people say.

  • VLM

Decide the task steps

VLT is the robot task brain: it decides what job steps make sense.

  • VLT: Vision-Language-Task

Predict the world and choose movement

XPENG is trying to make the AI predict the physical world before acting, and this same model turns that understanding into actions.

  • XPENG VLA 2.0

Move around

The same physical-world AI thinking from cars may help the robot move around.

  • VLA 2.0 cross-domain model

Hands and manipulation

The body and hands are highly articulated and skin-covered for human-like interaction.

Model not named

Whole-body walking

The body is designed to move more like a person, including spine and foot motion.

Model not named

Physical body

We know the body is complex; motor details are less public.

Not disclosed

Feeds back into the models

Around the stack

Named on the public record, but not part of the runtime chain above.

Car-side world-model context

XPENG may reuse prediction ideas across cars and robots, but we should not claim IRON definitely runs X-Mind or X-World.

  • X-World
  • X-Foresight
  • X-Mind

Cross-domain learning

XPENG can reuse data and model ideas from cars and robotics to improve physical AI, but this is platform-level evidence, not IRON-specific proof.

  • X-Mind

03

What sensors and hardware it has

XPENG IRON — IRON hands close-up
IRON hands close-up · XPENG

Parts sit on the body only where the public record places them. 1 of 8 entries are still not publicly disclosed.

Vision and audio

01
  • Electronic skin; exact camera/tactile sensor array not fully listed in text.

    Skin/contact is public; full sensor list is not.

    Reported for this robot · 1 source

Touch and force

01
  • Seamless electronic skin and full-body soft exterior.

    Skin-like covering may help safety/contact and human-facing design.

    Confirmed on this robot · Across the robot · 1 source

Body position and balance

02
  • Official XPENG IRON page lists 173 cm and 65 kg.

    Human-sized but relatively light for its DoF.

    Confirmed on this robot · Across the robot · 1 source

    Still open: Some secondary pages list different early prototype specs; use official page for current IRON.

  • Bionic lumbar spine and bionic forefoot.

    Spine and feet are designed to support more natural motion.

    Confirmed on this robot · Legs · 1 source

Joints and actuators

01
  • 82 DoF full body; 22 DoF per hand.

    Very high articulation, especially in the hands.

    Confirmed on this robot · Across the robot · 1 source

Compute and connectivity

02
  • VLM + VLT + VLA 2.0 described for IRON.

    It has named model layers from seeing/talking to task/action.

    Confirmed on this robot · 1 source

  • XPENG Turing AI compute is part of the broader physical-AI story, but exact IRON compute values need careful versioning.

    Don't overquote TOPS unless the specific IRON page states it.

    Conflicting public disclosures · 1 source

    Still open: Need exact product spec.

Battery and runtime

01
  • Not clearly specified in reviewed official IRON page text.

    We do not know runtime.

    Not publicly disclosed · 1 source

04

How it does one real job

Vendor-stated scenario

Retail/service task: guide a visitor to a display and hand over a small product sample

This walkthrough is a reasoned synthesis of publicly disclosed architecture pieces, not a confirmed end-to-end demo transcript.

6 moments · uses 5 of 6 parts of the system

  • 01

    Understand visitor request

    The robot interprets what the visitor sees and says.

    Technical detail

    VLM processes visual/language context from the visitor.

    • Senses
    • Understands
    • Predicts
    • Acts
    • Controls
    • Learns
    • VLM

    Confirmed on this robot · 1 source

  • 02

    Choose task sequence

    It converts the request into an ordered task plan.

    Technical detail

    VLT converts the request into a robot task plan: guide, stop, pick product, hand over.

    • Senses
    • Understands
    • Predicts
    • Acts
    • Controls
    • Learns
    • VLT: Vision-Language-Task

    Confirmed on this robot · 1 source

  • 03

    Predict and select action

    It predicts likely outcomes and picks actions accordingly.

    Technical detail

    VLA 2.0 generates actions and includes physical-world understanding/prediction.

    • Senses
    • Understands
    • Predicts
    • Acts
    • Controls
    • Learns
    • XPENG VLA 2.0

    Confirmed on this robot · 1 source

  • 04

    Walk to display

    The robot walks across the space toward the display.

    Technical detail

    Locomotion stack moves IRON through the space.

    Autonomous navigation evidence should be tracked.

    • Senses
    • Understands
    • Predicts
    • Acts
    • Controls
    • Learns
    • VLA 2.0 cross-domain model

    Confirmed on this robot · 1 source

  • 05

    Pick sample

    The hand grasps the item using vision, action and touch.

    Technical detail

    Hand/action system grasps item; electronic skin/contact may support safe interaction.

    Grip/tactile resolution not public.

    • Senses
    • Understands
    • Predicts
    • Acts
    • Controls
    • Learns

    No named model for this moment.

    Confirmed on this robot · 2 sources

  • 06

    Hand over safely

    The robot coordinates approach and handover while watching for contact.

    Technical detail

    Task/action layers coordinate approach and handover, while the contact/safety layer should detect interaction.

    Safety details not public.

    • Senses
    • Understands
    • Predicts
    • Acts
    • Controls
    • Learns
    • VLT: Vision-Language-Task
    • XPENG VLA 2.0

    Cyborgs interpretation · 2 sources

Parts are lit where the vendor's own description puts them to work. Unlit parts still run on the robot; they are simply not what this job turns on.

05

How it learns and improves

4 of 4 stations on the public record · 3 named training models

  1. 01

    The robot works

    Runs the parts of the system that later receive updates (1)

  2. 02

    Experience is captured

    Vehicle sensor/simulation data

  3. 03

    Car-side world-model context

    Vendor-stated, pre-production

  4. 04

    Updates go back on the robot

    Predicted future video/action, vehicle-side

Station 04 returns to station 01 — the loop repeats.

XPENG may reuse prediction ideas across cars and robots, but we should not claim IRON definitely runs X-Mind or X-World.

Technical detail

X-World is a controllable generative world model for future video/action simulation; X-Mind/X-Foresight are predictive world-model reasoning for vehicle-side driving. Strongest public evidence is in the vehicle/autonomy context, not confirmed as IRON runtime.

Vendor-stated, pre-production

  • X-World
  • X-Foresight
  • X-Mind

Does IRON directly use X-Mind/X-World?

What the updates land on

  • Predict the world and choose movementPredicts

06

What the vendor has shared

11
Vendor-confirmed

Confirmed on this robot

03
Reported or contested

Conflicting public disclosures · Reported for this robot

04
Stated, not shown

Vendor-stated, pre-production

02
The record is silent

Not publicly disclosed

A map of the public record, not a verdict. A silent line means the vendor has not said — it never moves the ranking.

What each band means
  • Vendor-confirmed

    The vendor publicly describes this for the exact robot named on this page.

  • Reported or contested

    Public sources disagree; we show the conflict rather than pick one.

  • Stated, not shown

    The vendor has stated the intent; it has not been shown in a shipped configuration.

  • The record is silent

    The public record does not say. We leave it visible as unknown.

Still open · 03

Does IRON directly use X-Mind/X-World?

XPENG has world-model work; direct IRON runtime linkage is not fully established.

Why it matters

Avoids confusing vehicle world model with humanoid runtime.

Vendor-stated, pre-production

What is the exact compute/battery?

Current official IRON page has body/DoF details but limited compute/runtime disclosure.

Why it matters

Affects autonomy and deployment practicality.

Not publicly disclosed

What tasks has IRON done autonomously?

Public claims and demos need evidence classification separately from architecture disclosure.

Why it matters

Architecture is impressive but ranking needs proof.

Vendor-stated, pre-production

07

Technical details

The record

Every row is the public position. Undisclosed rows stay in the list.

03 of 04 rows have a public answer

08

Sources

The ranking measures demonstrated capability. This page explains the public system behind it. The system description does not affect the score.

Return to the XPENG IRON report →