Padmi

AI/ML Engineer Model Optimization & Acceleration

BangalorePosted 1 month ago
Software engineeringSeniorFull Time; Regular
Apply at Programming.Com

Opens the source posting on shine.com

Source description

About the role

View original

AI/ML Engineer Model Optimization & Acceleration (810 Years) Location: Bengaluru, IndiaExperience: 810 Years ( If you have experience more than 8 years only apply then)Open Positions: 3We are looking for an experienced AI/ML Engineer to optimize and deploy machine learning models across heterogeneous platforms (CPU, GPU, and NPU). If you're passionate about building high-performance, production-ready AI systems and working on cutting-edge technologies, we'd love to hear from you! Key Responsibilities Optimize AI models including LLMs, Diffusion Models, CNNs, Computer Vision, Multi-modal, and Speech Models.Port models across frameworks (PyTorch ONNX Runtime).Deploy and optimize models on GPU/NPU hardware accelerators.Improve inference latency, throughput, and memory efficiency.Implement quantization, model compression, and performance tuning.Profile, benchmark, and debug AI system performance.Required Skills Strong expertise in PyTorch and ONNXProficiency in Python and C++Experience with CUDA, ROCm, or GPU accelerationStrong understanding of Transformers, CNNs, Deep LearningHands-on experience in Model Optimization, Quantization, Inference Optimization, and Performance TuningGood to Have Edge AI / Embedded AI deploymentGenerative AI or Multi-modal AIDistributed inference or streaming pipelinesTensorRT, OpenVINO (preferred)Compensation: 2,500,000.00 - 3,000,000.00 per year Experience: AI/ML Engineer Model Optimization & Acceleration: 8 years (Preferred)Work Location: In person AI/ML Engineer Model Optimization & Acceleration (810 Years) Location: Bengaluru, IndiaExperience: 810 Years ( If you have experience more than 8 years only apply then)Open Positions: 3We are looking for an experienced AI/ML Engineer to optimize and deploy machine learning models across heterogeneous platforms (CPU, GPU, and NPU). If you're passionate about building high-performance, production-ready AI systems and working on cutting-edge technologies, we'd love to hear from you! Key Responsibilities Optimize AI models including LLMs, Diffusion Models, CNNs, Computer Vision, Multi-modal, and Speech Models.Port models across frameworks (PyTorch ONNX Runtime).Deploy and optimize models on GPU/NPU hardware accelerators.Improve inference latency, throughput, and memory efficiency.Implement quantization, model compression, and performance tuning.Profile, benchmark, and debug AI system performance.Required Skills Strong expertise in PyTorch and ONNXProficiency in Python and C++Experience with CUDA, ROCm, or GPU accelerationStrong understanding of Transformers, CNNs, Deep LearningHands-on experience in Model Optimization, Quantization, Inference Optimization, and Performance TuningGood to Have Edge AI / Embedded AI deploymentGenerative AI or Multi-modal AIDistributed inference or streaming pipelinesTensorRT, OpenVINO (preferred)Compensation: 2,500,000.00 - 3,000,000.00 per year Experience: AI/ML Engineer Model Optimization & Acceleration: 8 years (Preferred)Work Location: In person

One address, no account. We’ll tell you when matching roles go live.

More at Programming.Com

Related open roles

View all roles