VLA: the model behind NGP

VLA 2.0 is a unified model that takes the camera images and an instruction and returns the driving decision. What the second generation adds, and what XPENG has stated about Europe.

The number of Turing chips determines which version of VLA the car can run. In Europe, the system must also be approved before it can be activated.

NGP is the name of what you switch on. VLA is the model underneath it. The two names therefore describe different layers of the system.

This guide explains what the three letters mean, what the second generation adds, the research XPENG has published, and which cars have been announced with the model. Which chips sit in which car is covered in the assistance hardware guide.

Hourglass of light particles on an XPENG VLA 2.0 poster

Images: XPENG

Jump to: What VLA is · Second generation · The research · VLA in China and Europe · The XPACE model in IRON · Approval in Europe · What XPENG has not said · Q&A · Sources

No European XPENG runs second-generation VLA today. The equipment list for L03 Ultra states VLA 2.0 NGP from the first quarter of 2027. G9L has its European premiere in October 2026, but XPENG has not yet stated which Chinese chip tiers Europe will get. The dates are under Approval in Europe.

What VLA is, and what it replaced

VLA stands for vision-language-action. One model takes the camera images and a text instruction and returns the driving decision. It replaces a chain of separate parts, where one registered the surroundings, the next predicted what they would do, the third laid a path, and the fourth steered the car. XPENG has not explained publicly why the chain was dropped.

VLA 2.0 converts camera signals directly into vehicle actions without first producing a textual explanation. That is the difference from the earlier V-L-A structure, which used language as an intermediate step. Language still enters when the driver gives an instruction and through the data used to train the model.

On XPENG’s EU website, XPILOT ASSIST 2.5 is still the product name. It covers the assistance package, not the model.

XPENG uses three names for the model itself. VLA 2.0 is the model from 2025. XPENG presented second-generation VLA on 27 August 2026. The English material calls it the first major upgrade of VLA 2.0, while the Chinese online store uses the name second generation. Posters shorten it to the 630 version after XOS 6.3.0, the release it ships with. The rollout in China began in September 2026.

What the second generation adds

Infini-VLA lets up to 30 seconds of history count in the decision. The model can therefore continue to account for a cyclist who disappeared behind a van ten seconds ago. XPENG’s Xianming Liu has written that 30 seconds is the window they chose for the car, not a limit in the model. When the car receives a new frame, a cache reuses earlier calculations. The model therefore does not have to process the full history again, so extending the window adds little computing work.

Timeline of 30 seconds of surroundings behind a moving XPENG

X-Foresight sketches where the other road users are heading, around six seconds ahead. The name is on XPENG’s Chinese G9L material but not on the equipment list, and whether an owner can switch it on and off has not been stated.

Predicted paths for road users around an XPENG at a junction

Streaming inference lets detection, reasoning and action run at the same time instead of one after another. XPENG states a 300 percent faster response from sensor to action. That figure is company-reported.

Master Agent combines the driving model with the cabin’s vision-language model (VLM). When the driver asks the car to pull over, the model distributes the command among the driving, chassis, cabin and body systems. XPENG calls the method Mixture of Tasks. It is built on Omni, the company’s multimodal foundation model.

Master Agent splitting one spoken command across the car's systems

VLA 2.0 Lite is designed for one Turing chip. P7+ and X9 each have one. HybridViT is the Lite version’s vision encoder and performs the same task as TuringViT, explained under The research. It sends 67 percent fewer image tokens to the language model, reducing the computing requirement. He Xiaopeng explained the choice in a video on Weibo on 29 August 2026. According to him, halving the parameters would also reduce the model’s capabilities.

The twentyfold safety figure on all the posters is not one measurement. A footnote in the G9L launch material shows that the figure combines results across seven areas: risk recognition, path planning, roadworks, delayed reaction, future reasoning, generalisation and robustness.

The phrase “Robotaxi-like L4 experience” is on the posters too. That is marketing. As of September 2026, none of these cars has L4 approval anywhere.

The research behind it

TuringViT is the part of the model that converts camera images into data the rest of the system can process. A conventional attention layer compares every part of an image with every other part. If the image resolution doubles, the number of calculations roughly quadruples. That is too demanding when the car processes several cameras at once. TuringViT therefore uses five linear-attention layers followed by one conventional attention layer. The five layers reduce the computing work, while the final layer retains details that linear attention can otherwise lose. The computing work consequently grows roughly in step with image resolution instead of with its square.

Architecture diagram for TuringViT with five linear-attention layers and one conventional attention layer
TuringViT - the architecture is available with 18 or 24 layers.

X-World generates new training situations. From one camera recording and a text description, it creates seven camera views of a simulated sequence on which the model can be trained and tested. By the end of April 2026, XPENG had run more than 500,000 simulated scenarios against 30,000 the year before. The company puts the daily simulation at the equivalent of 30 million kilometres driven.

X-Mind reduces the delay between a world model and VLA. If the models run in sequence, the car must first wait for the world model’s prediction and then for VLA’s decision. In X-Mind, the driving model performs both calculations in the same process. The prediction is a top-down representation with twelve future steps encoded as 96 tokens, not photorealistic images. XPENG does not state how much time those twelve steps cover.

Architecture diagram for X-Mind with prediction and driving decision in the same model

X-Cache saves compute. XPENG puts it at cutting around 70 percent of the redundant work and making the model run up to about 2.7 times faster.

At CVPR 2026, an AI research conference focused on image analysis and pattern recognition, XPENG described the difference between VLA and world models. A VLA model learns from what people did behind the wheel. That provides fewer examples, but each one is a complete decision. A world model learns from the footage itself, frame by frame, how cars, cyclists and traffic lights moved.

VLA in China and Europe

In China, G9L is the first car delivered with second-generation VLA. The full model runs on Ultra SE, Ultra and the Ultra flagship, the distilled Lite version on Max, and the cabin VLM is only on Ultra and the Ultra flagship. What that takes in hardware is in the assistance hardware guide.

In China, the customer chooses the number of Turing chips as part of the equipment tier. The Chinese G9L is therefore offered with between one and three chips and with different versions of VLA. In Europe, each XPENG model has so far had one fixed chip configuration. XPENG is now changing that pattern with L03, which is offered as L03 and L03 Ultra with different chip configurations. This may be the first sign that part of the Chinese tier structure is coming to international markets, but XPENG has not announced a general strategy.

Three XPENG Turing AI chips on a poster from the G9L launch

On the European market, L03 Ultra is currently the only model with two Turing chips assigned to driving. It has three Turing chips in total; the third is assigned to the cabin. The hardware configuration can run the full VLA model, but the equipment list does not state VLA 2.0 NGP until the first quarter of 2027. XPENG has also said that current cars with Orin-X cannot be upgraded to Turing hardware.

G9L has its European premiere at the Paris Motor Show in October 2026. XPENG has not yet stated which of the Chinese variants Europe will get. For L03, the company chose fewer international variants than it offers in China. It may make the same choice for G9L.

Outside XPENG, Volkswagen is the company’s first external customer for VLA 2.0 and has chosen the Turing chip for its own cars at the same time. Which model, which market and when, XPENG has not stated.

XPENG has not said whether P7+ and X9 will get NGP or VLA 2.0 Lite. Both have one Turing chip, for which Lite is designed, but that hardware detail does not confirm a future software rollout.

The XPACE model in IRON

XPENG builds the humanoid robot IRON with three Turing chips. On 15 September 2026, XPENG Robotics published a report on arXiv about XPACE, the model that controls the robot. It selects the robot’s movements and predicts the video those movements produce. Part of the training material is generated in a simulator.

XPENG IRON humanoid robot in a test rig in front of three bowls on a white table

XPACE is not the car’s VLA. Both models connect video data with actions, but XPACE controls a robot while VLA controls a car.

At the G9L launch, XPENG said that the cars’ VLA and VLM would be used in IRON. The company also said that IRON’s multilingual dialogue would later be used in cars. XPENG had, however, already shown multilingual VLM dialogue in a car video from AI Day in November 2025. The September statement therefore cannot be read as the first introduction of multilingual in-car conversations. XPENG did not state what specifically would be transferred from IRON to the cars.

The numbers in the XPACE report are laboratory results. After extra training the robot’s success rate went from 61.7 to 86.7 percent across three table tasks, measured with 20 attempts per task on one unit. The whole story is in the news item on XPACE.

Approval in Europe

The first quarter of 2027 is on L03 Ultra’s equipment list against VLA 2.0 NGP. A list is not an approval.

The first half of 2027 is what XPENG is aiming at for European approval. The approval is sought under UN Regulation 171 on Driver Control Assistance Systems, and the company has already run local acceptance testing in Germany. A local test is not type approval.

Whether a model trained in China can drive in Europe is still unsettled. XPENG stated at Brand Day in Munich that VLA 2.0 accounted for 50.44 percent of the kilometres the equipped cars in China drove. That is XPENG’s own figure, with no method stated, and it covers a Chinese fleet. By XPENG’s own account the China-trained model came close to home-market level on German city roads with almost no local training data. XPENG has not stated how much of the training comes from simulation.

The rest of the differences between a Chinese and a European XPENG, from chips to map data, are in the assistance hardware guide.

What XPENG has not said

XPENG’s own materials do not assign the Turing chip and TuringViT to the same products. The company says the Turing chip is used in cars, humanoid robots and flying vehicles. The TuringViT report mentions only driving, cabin interaction and humanoid robots. The report therefore does not establish whether TuringViT is also used in the flying vehicles.

How large the model in the car is, XPENG has never stated. Only the cloud model’s 72 billion parameters are named, and for the car there is only a ratio: 3.5 times the version running today.

How long X-Mind’s twelve steps cover is not in the report.

When Lite comes to European cars, XPENG has said nothing. In China the company is testing a distilled version internally on cars with two Orin-X chips. Liu says adapting it crosses chip and model architectures. It needs optimisation down to individual compute operations.

Q&A

No. No European XPENG runs second-generation VLA. L03 Ultra has three Turing chips in total, two of which are assigned to driving, and its equipment list states VLA 2.0 NGP from the first quarter of 2027. XPENG has not confirmed VLA 2.0 for other European models.

NGP is the product you switch on. VLA is the model running underneath it. Same system, two levels: one is the name on the screen, the other is the thing that decides.

VLA 2.0 Lite uses the same foundation model with a smaller vision encoder. XPENG says Lite retains most of the capability and can compare with leading Chinese L2 systems. That is the company’s own assessment. No independent measurement is known.

Because the company says the cars’ VLA and VLM will be used in the IRON robot. XPENG has also talked about transferring multilingual dialogue from IRON to cars, but it had already shown multilingual VLM dialogue in a car video in November 2025. It has not stated what will be transferred or when.

Sources

  • XPENG AI Day, 5 November 2025 - VLA 2.0; the 72 billion parameter cloud model; Volkswagen as first external customer
  • XPENG Foundation Model Team: TuringViT, 23 June 2026 - the vision encoder
  • XPENG research presented at CVPR, Denver, June 2026 - the split between VLA and world model; X-World, X-Foresight and X-Cache
  • X-Mind, 29 June 2026 - the prediction folded into the driving model
  • Brand Day, Munich, 16 July 2026 - the 50.44 percent in China
  • XPENG’s corporate release on second-generation VLA, 27 August 2026 (published 1 September) - Infini-VLA, X-Foresight, streaming inference, Master Agent, Lite on G9L Max, acceptance testing in Germany
  • Xianming Liu on the architecture behind XOS 6.3.0, 15 September 2026 - the Infini-VLA cache and adaptation to cars with two Orin-X chips
  • He Xiaopeng on Weibo, 29 August 2026 - HybridViT and the 67 percent fewer image tokens
  • XPENG’s own Weibo posts from the G9L launch, 17 September 2026 - variants and models, the twentyfold figure with its footnote, Master Agent, the use of VLA and VLM in IRON
  • XPENG Robotics: XPACE, arXiv 2609.17372, 15 September 2026 - the model behind IRON

Supplementary video: XPENG Future Intelligent Cabin, November 2025 - demonstration of multilingual VLM dialogue in the car.

Related: Driver assistance hardware · Software · The XPACE news item

Changes to this page (1)

All changes across the site