At a conference in Denver, XPENG wrote down how its driver assistance is built

In June 2026 XPENG was invited to CVPR for the third time. Dr Xianming Liu, head of the company’s General Intelligence Center, gave a keynote at the inaugural workshop on deploying foundation models for embodied AI, alongside Tesla, Nvidia and Waymo and academics from the University of California and the University of Toronto. The keynote was titled Building the World Model for Autonomous Driving, and an XPENG paper on driving-scene generation was accepted at the conference.

Composite image of a white humanoid robot walking beside a dark XPENG X9, with a small aircraft hovering and a city skyline behind

It is rare for XPENG to explain its own engineering in English outside a press release, which is the reason to read this one.

The division of labour

XPENG puts it this way: VLA 2.0 learns from how humans drive, while the world model learns the laws of physics by predicting what happens next. VLA 2.0 teaches the model how to act, and the world model gives it an understanding of how the surroundings evolve after each action.

The world model’s three capabilities are deliberative reasoning, controllable generation and long-horizon forecasting. Three papers are published with links in the release: X-World, which produces physically plausible future video from specified motion inputs and is used in closed-loop simulation and to generate training data; X-Foresight, which is built into VLA 2.0 and jointly predicts future multi-view imagery and the car’s own actions; and X-Cache, which XPENG says cuts roughly 70 per cent of redundant computation and speeds the denoising backbone by up to around 2.7 times. A fourth, X-Mind, is announced but not published.

The numbers

VLA 2.0 is in formal mass production, and XPENG says that in its first month after rollout the system accounted for more than 50 per cent of assisted driving mileage. Over what fleet, in which markets and how that share is measured, the release does not say.

The model has billions of parameters, is trained on hundreds of millions of video clips, and consumes more than four trillion tokens per model iteration. In the twelve months to March 2026, per-GPU training efficiency rose 1,010 per cent and single-job efficiency 4,360 per cent, and hardware utilisation went from 40 to 90 per cent. There are no absolute baselines behind those percentages.

The robotaxi, built on the GX platform, has rolled off the line with 3,000 TOPS on board.

The date for the robot

IRON is, the release says, entering the phase where hardware and software are integrated, with formal mass production by the end of 2026. And the robots are to start working as in-store shopping guides at XPENG’s own retail outlets from the first quarter of 2027.

That date does not appear in the story of the robots’ production line from September 2026. This release is where it is.

Nothing in the text is European. The VLA 2.0 rollout, the robotaxi and the coverage are Chinese, and the release does not say what of it reaches a European car. Turing and the agent is the thread that has come closest.

Images: XPENG

Sources

XPENG: the world model presented at CVPR in Denver · 3 June 2026