The model has leapfrogged from the 3D era to the 4D era.

At the Xpeng Physics‑AI Sharing Session & 2nd‑Generation VLA New Version Experience Day on August 27, Xpeng Motors gave a detailed briefing on the technical roadmap for its brand‑new second‑generation VLA. Lei Tech / EV Insight was invited to attend the on‑site event.


According to Liu Xianming, Head of the General Intelligence Center at Xpeng Group, the core upgrade of the new second‑generation VLA lies in incorporating time into the model for the first time. This advances the AI’s dimension for understanding the world from 3D space to 4D space‑time. Space tells the model what the world looks like, while time tells it what is happening in the world.


IMG_3822.jpg

(Photograph source: LeiTech/EV Pass, captured onsite)




It features the brand‑new Infini‑VLA long‑temporal architecture, which can retain memories of the world over the preceding 30 seconds. When making decisions, the model can leverage valid information from much longer‑range context.


The X‑Foresight world‑prediction model can forecast what will unfold over the next six seconds. For instance, it can infer whether a vehicle in an adjacent lane will cut in based on its movement, or judge whether the car ahead is going to stop according to changes in its speed.


Built upon Infini‑VLA and X‑Foresight, Xpeng’s new second‑generation VLA has evolved from single deterministic outputs to multi‑solution outputs. With further reinforcement learning, trajectories generated by the model become more natural and human‑like.


Conventional intelligent driving models follow the logic of “perceive‑compute‑output‑re‑perceive”. By contrast, the new second‑generation VLA operates on a “perceive‑think‑act simultaneously” paradigm. To ensure the model’s decision‑making speed keeps pace with environmental changes, it also integrates streaming autoregressive inference, delivering a 300 % improvement in end‑to‑end response speed.


Second-Generation VLA Advantages.jpg

(Photograph source: XPeng)


Liu Xianming also mentioned that on‑device computing power is inherently limited, yet a model’s generalization ability is closely tied to its parameter count. While the new second‑generation VLA boosts model parameters by 3.5 times — over 1.5 times those of mainstream VLAs — it also adopts the MoT hybrid architecture.


The MoT architecture splits complete Transformer Blocks while retaining globally shared joint attention. Different Transformer sub‑towers can exchange information with one another, bringing lower load‑balancing pressure compared with MoE.


Training data for Xpeng’s new second‑generation VLA comes from its million‑vehicle fleet and billions‑scale data assets. Single‑run training data throughput has increased ten‑fold compared with six months ago. Moreover, the model can proactively optimize data distribution, identify anomalous data, continuously mine long‑tail data, and raise the upper bound of model capabilities.


Beyond that, Xpeng has re‑architected the in‑vehicle brain with a robotics‑oriented approach and rolled out Master Agent, which fuses VLA and VLM. Built on the Omni multimodal model, Master Agent can interpret natural language and ambiguous semantics more accurately, and proactively break down users’ voice commands into individual tasks. Via voice commands, users can instruct the vehicle to stop, perform parking maneuvers, start navigation, and add waypoints.


11.jpg

(Photograph source: XPeng)



This complete foundational base spans L2 through L4. It can scale upward to support L4‑level autonomous driving, and also be distilled downward for compatibility with a wider range of vehicle models. Furthermore, this foundation bridges automobiles and robots. Xpeng’s humanoid robot Iron is equipped with three Turing AI chips. With the model deployed on‑device, it can proactively complete tasks without remote control.


In addition, the Turing AI chip supports the XLLM general large‑model computing framework, enabling independent inference tasks on a single chip with an average inference speed of over 20 tokens per second.


At this point, Xpeng owners are most likely wondering when they will get access to the new second‑generation VLA. Xpeng has not kept us waiting long. Both the new second‑generation VLA and its distilled second‑generation VLA Lite variant will start rolling out in September. The full version is for Ultra and Ultra SE trims, while VLA Lite targets Max‑trim vehicles fitted with a single Turing AI chip. The Xpeng G9L will be the first vehicle to come with both the new second‑generation VLA and the distilled second‑generation VLA Lite.


IMG_3847.jpg

(Photograph source: LeiTech/EV Pass, captured onsite)


In Q1 of this year, Xpeng Motors rebranded itself as Xpeng Group, rolling out its full lineup of businesses including new‑energy vehicles, flying cars, Robotaxis and humanoid robots. At this stage, Xpeng required an AI foundation capable of unifying all scenarios and product categories.


It is clear that the new second‑generation VLA serves as exactly such a foundation. It delivers a leap in model comprehension from spatial to temporal understanding. It can retain valid long‑term temporal memory spanning 30 seconds and predict the motion trajectories of surrounding objects for the next six seconds. This safeguards the safety of autonomous driving for its vehicle products and enables humanoid robots to complete tasks autonomously.


The arrival of the new second‑generation VLA marks Xpeng’s genuine evolution from an automaker into a “Physics‑AI company”.


Latest With Watermark.png


Gaoding Design-14.png