Source description
About the role
Deep expertise in model quantization (PTQ, QAT) and mixed-precision inference frameworks (INT8, FP8, FP4, BF16/FP16).
Proven experience optimizing large-scale models (Multi-Modal Sensor Fusion models, LLMs, VLMs/VLAs) utilizing Efficient Attention mechanisms (e.g., FlashAttention, Linear Attention), KV-cache optimization (e.g., PagedAttention) and Speculative Decoding.
Extensive experience with model conversion/compilation pipelines (e.g., ONNX, TensorRT, torch.compile) and performing rigorous latency benchmark and model quality parity valuation.
Proficiency in low-level programming for AI accelerators, specifically developing and optimizing custom ML OPs and TensorRT Plugins with efficient CUDA kernel implementations.
Production-level C++ (14/17/20) and Python programming skills, with experience developing concurrent, memory-safe, real-time inference code for edge devices.
More at Zoox
Related open roles
Staff ServiceNow Platform Engineer
San Francisco Bay Area · Hybrid
Embedded Software Engineer - Power Systems
San Francisco Bay Area · San Diego · Hybrid
Senior Data Engineer – Enterprise, Data & AI
San Francisco Bay Area · Hybrid
AI Engineer – Enterprise, Data & AI
San Francisco Bay Area · Hybrid
Senior Full-stack Engineer - Mapping Web Services
San Francisco Bay Area · Hybrid
Part-Time Student Worker – AI Validation and Benchmarking Engineer
San Francisco Bay Area · Hybrid