Senior Machine Learning Engineer - LLM Quantization & Deployment
xpeng motorsSan Jose, CA
today
Occupations
Software DevelopersData ScientistsComputer and Information Research ScientistsIndustries
Custom Computer Programming ServicesComputer Systems Design ServicesSoftware PublishersAbout the role
Overview
In this role you will advance VLA inference for XPENG’s next‑gen AI and autonomous driving stack, focusing on robust LLM quantization and deployment. You will work with cross‑functional teams to productionize PTQ/QAT, manage mixed‑precision inference, and ensure numerical consistency with training models. You’ll build end‑to‑end pipelines for export, calibration, benchmarking, and validation, while shaping performance estimates and sign‑offs for field testing and on‑vehicle use. You will contribute to cutting‑edge AI research translation and have a meaningful impact on autonomous mobility.
Compensation / Benefitsfun, supportive and engaging environmentinfrastructures and computational resourcescutting‑edge technologies with top talentsimpact on transportation revolution through autonomous drivingcompetitive compensation packagesnacks, lunches, dinners, and fun activities
Responsibilities Develop VLA inference models and productionize LLM quantization methods (PTQ, QAT, mixed‑precision, INT8, FP4)Write production‑quality Python with strong testing, observability, reproducibility, and failure handling Build export, calibration, benchmarking, validation, and deployment pipelines Collaborate with VLA model research to estimate performance and feasibility Curate evaluation datasets and establish a comprehensive metric suite for benchmarking VLA performance Analyze numerical errors, accuracy regressions, and performance trade‑offs Develop PTQ and QAT orchestration workflows Interface with field‑testing and simulation teams for autonomous driving performance sign‑off Collaborate with in‑vehicle software on latency analysis and issue triage Collaborate with training infrastructure on QAT and model distillation
Key requirements Master in CS/CE/EE, or equivalent, 1–3 years industry experience (Open to new graduates)Strong understanding of Transformer architectures and LLM inference Hands‑on experience quantizing or deploying DL models in production Proficiency with PyTorch and at least one inference or compilation stack Strong Python programming and software engineering skills Ability to work across research, systems, infrastructure, and product teams Excellent communication and problem‑solving skills in a fast‑paced, collaborative environmentexcellent communicationproblem solvingcollaborative Transformer architectures and LLM inference Quantization/deployment of DL models in production PyTorch
Matching similar jobs
JOB OVERVIEW
Experience level
Lead
Location
San Jose, CA
Occupation
Software Developers
Industry
Custom Computer Programming Services
Posted
today
Tired of running searches?
Rank the roles you'd take once, and matches like these arrive on their own.
CREATE PROFILE