Source description
About the role
AI/ML Engineer Model Optimization & Acceleration (810 Years) Location: Bengaluru, IndiaExperience: 810 Years ( If you have experience more than 8 years only apply then)Open Positions: 3We are looking for an experienced AI/ML Engineer to optimize and deploy machine learning models across heterogeneous platforms (CPU, GPU, and NPU). If you're passionate about building high-performance, production-ready AI systems and working on cutting-edge technologies, we'd love to hear from you! Key Responsibilities Optimize AI models including LLMs, Diffusion Models, CNNs, Computer Vision, Multi-modal, and Speech Models.Port models across frameworks (PyTorch ONNX Runtime).Deploy and optimize models on GPU/NPU hardware accelerators.Improve inference latency, throughput, and memory efficiency.Implement quantization, model compression, and performance tuning.Profile, benchmark, and debug AI system performance.Required Skills Strong expertise in PyTorch and ONNXProficiency in Python and C++Experience with CUDA, ROCm, or GPU accelerationStrong understanding of Transformers, CNNs, Deep LearningHands-on experience in Model Optimization, Quantization, Inference Optimization, and Performance TuningGood to Have Edge AI / Embedded AI deploymentGenerative AI or Multi-modal AIDistributed inference or streaming pipelinesTensorRT, OpenVINO (preferred)Compensation: 2,500,000.00 - 3,000,000.00 per year Experience: AI/ML Engineer Model Optimization & Acceleration: 8 years (Preferred)Work Location: In person AI/ML Engineer Model Optimization & Acceleration (810 Years) Location: Bengaluru, IndiaExperience: 810 Years ( If you have experience more than 8 years only apply then)Open Positions: 3We are looking for an experienced AI/ML Engineer to optimize and deploy machine learning models across heterogeneous platforms (CPU, GPU, and NPU). If you're passionate about building high-performance, production-ready AI systems and working on cutting-edge technologies, we'd love to hear from you! Key Responsibilities Optimize AI models including LLMs, Diffusion Models, CNNs, Computer Vision, Multi-modal, and Speech Models.Port models across frameworks (PyTorch ONNX Runtime).Deploy and optimize models on GPU/NPU hardware accelerators.Improve inference latency, throughput, and memory efficiency.Implement quantization, model compression, and performance tuning.Profile, benchmark, and debug AI system performance.Required Skills Strong expertise in PyTorch and ONNXProficiency in Python and C++Experience with CUDA, ROCm, or GPU accelerationStrong understanding of Transformers, CNNs, Deep LearningHands-on experience in Model Optimization, Quantization, Inference Optimization, and Performance TuningGood to Have Edge AI / Embedded AI deploymentGenerative AI or Multi-modal AIDistributed inference or streaming pipelinesTensorRT, OpenVINO (preferred)Compensation: 2,500,000.00 - 3,000,000.00 per year Experience: AI/ML Engineer Model Optimization & Acceleration: 8 years (Preferred)Work Location: In person
More at Programming.Com