Xpeng
Machine Learning Engineer - Ai Foundation
Overview
We are looking for a full-time Machine Learning Engineer - AI Foundation, with deep knowledge and strong enthusiasm towards establishing a state-of-art ML infrastructure for training very large foundation model and accelerating model training/inference.
About Xpeng
XPENG is a leading smart technology company at the forefront of innovation, integrating advanced AI and autonomous driving technologies into its vehicles, including electric vehicles (EVs), electric vertical take-off and landing (eVTOL) aircraft, and robotics. With a strong focus on intelligent mobility, XPENG is dedicated to reshaping the future o
Requirements & Eligibility
- Master's Degree in CS/CE/EE, or equivalent, in industry experience.
- Deep knowledge of PyTorch.
- Knowledge of model inference framework (e.g. vLLM, SGLang)
- In-depth knowledge of transformer architecture and ways to accelerate the training and inference of transformer models.
- Experience of performing large scale distributed training of models.
- A track record of profiling model and doing detective work to improve model training and inference speed.
Key Responsibilities
- Design and implement training data pipeline that streams data from hundreds of petabytes of labeled and unlabeled data from a fleet of over a million vehicles.
- Implement training framework for all physical AI foundation models in XPeng, including VLA 2.0, XWorld, Robotics.
- Accelerate training with state of the art parallelisms, e.g., FSDP, Expert Parallel, Context Parallel, and data types.
- Accelerate model inference on the cloud for closed-loop simulation, reinforcement learning, and enterprise LLM/VLM applications.
Disclaimer: Trace Hiring is an independent job board. We are not directly affiliated with Xpeng. Please verify all details on the official company application portal.