Padmi

AI ML Engineer

BangalorePosted 2 months ago
Software engineeringSeniorFull Time; Regular
Apply at Oolka

Opens the source posting on shine.com

Source description

About the role

View original

Responsibilities: Design and implement asynchronous multi-agent orchestration.Own end-to-end latency from user message to AI response.Build resilient inference pipelines that gracefully degrade under load.Implement intelligent request routing and load balancing for AI workloads.Migrate critical AI conversation flow from monolith to dedicated services.Implement WebSocket/streaming infrastructure for real-time chat.Design circuit breakers and fallback strategies for AI model failures.Build comprehensive observability for AI system performance.Optimise credit data retrieval and caching strategies. Requirements: 6+ years building production systems handling > 10k concurrent users.Proven experience with async/event-driven architectures (not just REST APIs).Hands-on experience scaling ML/AI inference in production.Deep understanding of caching strategies (Redis, in-memory, CDN).Experience with message queues and real-time communication protocols.Built systems integrating multiple LLM/AI models in production.Experience with AI model serving frameworks (TensorFlow Serving, Triton, etc. )Understanding of AI inference optimisation (batching, caching, model quantisation).Knowledge of conversation state management and context handling.Has debugged production issues under high AI inference load. Responsibilities: Design and implement asynchronous multi-agent orchestration.Own end-to-end latency from user message to AI response.Build resilient inference pipelines that gracefully degrade under load.Implement intelligent request routing and load balancing for AI workloads.Migrate critical AI conversation flow from monolith to dedicated services.Implement WebSocket/streaming infrastructure for real-time chat.Design circuit breakers and fallback strategies for AI model failures.Build comprehensive observability for AI system performance.Optimise credit data retrieval and caching strategies. Requirements: 6+ years building production systems handling > 10k concurrent users.Proven experience with async/event-driven architectures (not just REST APIs).Hands-on experience scaling ML/AI inference in production.Deep understanding of caching strategies (Redis, in-memory, CDN).Experience with message queues and real-time communication protocols.Built systems integrating multiple LLM/AI models in production.Experience with AI model serving frameworks (TensorFlow Serving, Triton, etc. )Understanding of AI inference optimisation (batching, caching, model quantisation).Knowledge of conversation state management and context handling.Has debugged production issues under high AI inference load.

One address, no account. We’ll tell you when matching roles go live.

More at Oolka

Related open roles

View all roles