Source description
About the role
Work Location: Bangalore Job Title: IT - AI/ML Engineer Experience: 8-10 Years Note :- Apply only if you are immediate joiner Job Description AI model optimization & acceleration. Seeking an AI Engineer to optimize and deploy ML models across heterogeneous platforms (CPU, GPU, NPU). Work on scalable, production-ready AI systems across domains like robotics, healthcare, and automotive. Job Responsibilities / Day-to-Day Activities Optimize diverse models: generative (LLMs, diffusion), vision (classification, detection, segmentation), multi-modal, and speech Port models across frameworks (e.g., PyTorch ONNX runtimes) Deploy on hardware accelerators (GPU/NPU) and optimize performance Improve inference latency, throughput, and memory (batching, caching, parallelism, fusion) Apply quantization and model compression (FP32 lower precision) Profile and debug system and model performance Required Skills Strong in PyTorch (or similar), ONNX (or equivalent) Proficient in Python and C++ Experience with GPU/hardware acceleration (CUDA/ROCm or similar) Solid understanding of deep learning models (transformers, CNNs) Knowledge of optimization, quantization, and performance tuning Good to Have Edge AI or embedded deployment Generative or multi-modal AI systems Distributed inference or streaming pipelines #Hiring #AIML #AIEngineer #MLEngineer #ArtificialIntelligence #MachineLearning #Python #DeepLearning #GenerativeAI #LLM #BangaloreJobs #HiringNow #MachineLearning #AI #PyTorch #ONNX #CUDA #GPU #NPU #Inference #EdgeAI #TechJobs #ROCm #GPUComputing Work Location: Bangalore Job Title: IT - AI/ML Engineer Experience: 8-10 Years Note :- Apply only if you are immediate joiner Job Description AI model optimization & acceleration. Seeking an AI Engineer to optimize and deploy ML models across heterogeneous platforms (CPU, GPU, NPU). Work on scalable, production-ready AI systems across domains like robotics, healthcare, and automotive. Job Responsibilities / Day-to-Day Activities Optimize diverse models: generative (LLMs, diffusion), vision (classification, detection, segmentation), multi-modal, and speech Port models across frameworks (e.g., PyTorch ONNX runtimes) Deploy on hardware accelerators (GPU/NPU) and optimize performance Improve inference latency, throughput, and memory (batching, caching, parallelism, fusion) Apply quantization and model compression (FP32 lower precision) Profile and debug system and model performance Required Skills Strong in PyTorch (or similar), ONNX (or equivalent) Proficient in Python and C++ Experience with GPU/hardware acceleration (CUDA/ROCm or similar) Solid understanding of deep learning models (transformers, CNNs) Knowledge of optimization, quantization, and performance tuning Good to Have Edge AI or embedded deployment Generative or multi-modal AI systems Distributed inference or streaming pipelines #Hiring #AIML #AIEngineer #MLEngineer #ArtificialIntelligence #MachineLearning #Python #DeepLearning #GenerativeAI #LLM #BangaloreJobs #HiringNow #MachineLearning #AI #PyTorch #ONNX #CUDA #GPU #NPU #Inference #EdgeAI #TechJobs #ROCm #GPUComputing
More at Silicon Patterns