Padmi

AI Implementation Strategist

BangalorePosted 1 month ago
Software engineeringSeniorFull Time; Regular
Apply at Silicon Patterns

Opens the source posting on shine.com

Source description

About the role

View original

Work Location: Bangalore Job Title: IT - AI/ML Engineer Experience: 8-10 Years Note :- Apply only if you are immediate joiner Job Description AI model optimization & acceleration. Seeking an AI Engineer to optimize and deploy ML models across heterogeneous platforms (CPU, GPU, NPU). Work on scalable, production-ready AI systems across domains like robotics, healthcare, and automotive. Job Responsibilities / Day-to-Day Activities Optimize diverse models: generative (LLMs, diffusion), vision (classification, detection, segmentation), multi-modal, and speech Port models across frameworks (e.g., PyTorch ONNX runtimes) Deploy on hardware accelerators (GPU/NPU) and optimize performance Improve inference latency, throughput, and memory (batching, caching, parallelism, fusion) Apply quantization and model compression (FP32 lower precision) Profile and debug system and model performance Required Skills Strong in PyTorch (or similar), ONNX (or equivalent) Proficient in Python and C++ Experience with GPU/hardware acceleration (CUDA/ROCm or similar) Solid understanding of deep learning models (transformers, CNNs) Knowledge of optimization, quantization, and performance tuning Good to Have Edge AI or embedded deployment Generative or multi-modal AI systems Distributed inference or streaming pipelines #Hiring #AIML #AIEngineer #MLEngineer #ArtificialIntelligence #MachineLearning #Python #DeepLearning #GenerativeAI #LLM #BangaloreJobs #HiringNow #MachineLearning #AI #PyTorch #ONNX #CUDA #GPU #NPU #Inference #EdgeAI #TechJobs #ROCm #GPUComputing Work Location: Bangalore Job Title: IT - AI/ML Engineer Experience: 8-10 Years Note :- Apply only if you are immediate joiner Job Description AI model optimization & acceleration. Seeking an AI Engineer to optimize and deploy ML models across heterogeneous platforms (CPU, GPU, NPU). Work on scalable, production-ready AI systems across domains like robotics, healthcare, and automotive. Job Responsibilities / Day-to-Day Activities Optimize diverse models: generative (LLMs, diffusion), vision (classification, detection, segmentation), multi-modal, and speech Port models across frameworks (e.g., PyTorch ONNX runtimes) Deploy on hardware accelerators (GPU/NPU) and optimize performance Improve inference latency, throughput, and memory (batching, caching, parallelism, fusion) Apply quantization and model compression (FP32 lower precision) Profile and debug system and model performance Required Skills Strong in PyTorch (or similar), ONNX (or equivalent) Proficient in Python and C++ Experience with GPU/hardware acceleration (CUDA/ROCm or similar) Solid understanding of deep learning models (transformers, CNNs) Knowledge of optimization, quantization, and performance tuning Good to Have Edge AI or embedded deployment Generative or multi-modal AI systems Distributed inference or streaming pipelines #Hiring #AIML #AIEngineer #MLEngineer #ArtificialIntelligence #MachineLearning #Python #DeepLearning #GenerativeAI #LLM #BangaloreJobs #HiringNow #MachineLearning #AI #PyTorch #ONNX #CUDA #GPU #NPU #Inference #EdgeAI #TechJobs #ROCm #GPUComputing

One address, no account. We’ll tell you when matching roles go live.

More at Silicon Patterns

Related open roles

View all roles