Padmi

Senior AI Engineer

BangalorePosted 1 month ago
Software engineeringSeniorFull Time
Apply at Recro

Opens the source posting on foundit.in

Source description

About the role

View original

AI Infrastructure Engineer (Backend Platform) Experience: 5–8 Years Employment Type: Full-time About the Role We are looking for an experienced AI Infrastructure Engineer to build and scale the backend systems that power production-grade AI applications. You will work on high-performance, distributed systems that orchestrate AI workloads, optimize inference performance, and deliver reliable, low-latency conversational experiences for users at scale. This role is ideal for engineers who have built scalable backend platforms, integrated Large Language Models (LLMs) into production systems, and have hands-on experience with asynchronous architectures, real-time communication, and distributed systems. Key Responsibilities AI Infrastructure & Orchestration Design and implement asynchronous, event-driven AI orchestration systems. Build and optimize multi-agent workflows for complex AI interactions. Own end-to-end latency from user request to AI response. Develop resilient AI inference pipelines with retries, circuit breakers, and graceful fallback mechanisms. Implement intelligent request routing and load balancing across multiple AI models and providers. Build scalable services for AI conversation orchestration and migrate monolithic components into distributed microservices where required. Backend Platform Engineering Design and develop highly scalable backend services capable of handling high concurrent traffic. Build real-time communication infrastructure using technologies such as WebSockets or Server-Sent Events (SSE). Optimize backend performance through caching, asynchronous processing, and efficient data retrieval. Implement distributed messaging using Kafka, RabbitMQ, or similar event-streaming technologies. AI Platform & Performance Integrate production-grade LLM APIs and model serving platforms. Optimize inference performance through batching, response caching, prompt optimization, and context management. Design conversation state management for multi-turn AI interactions. Implement intelligent fallback strategies across multiple AI providers. Reliability & Observability Build monitoring and observability solutions for latency, throughput, error rates, and AI performance metrics. Monitor production systems using tools such as Grafana, Prometheus, Datadog, or OpenTelemetry. Troubleshoot production issues related to AI inference, distributed systems, and backend scalability. Required Skills & Experience Experience 3–5 years of experience building scalable backend systems. Proven experience developing production systems serving high concurrent user traffic (10,000+ concurrent users preferred). Experience designing distributed systems and microservices architectures. Backend Technologies Strong programming skills in one or more of the following: Python Go Java Node.js Experience with backend frameworks such as FastAPI, Spring Boot, Express/NestJS, or equivalent. Strong understanding of asynchronous programming concepts. Distributed Systems Hands-on experience with event-driven architectures. Experience using Kafka, RabbitMQ, AWS SQS, Google Pub/Sub, NATS, or similar messaging platforms. Strong understanding of distributed system design principles. AI & LLM Integration Experience integrating Large Language Models into production applications. Hands-on experience with one or more of the following: OpenAI Anthropic Claude Google Gemini Azure OpenAI vLLM Triton Inference Server TensorFlow Serving Experience implementing retry logic, rate limiting, fallback strategies, and AI provider failover. Understanding of prompt optimization, inference latency, batching, and conversation context management. Caching & Performance Strong experience with Redis for caching, session management, and rate limiting. Experience optimizing backend latency and throughput. Real-Time Systems Experience building streaming applications using: WebSockets Server-Sent Events (SSE) Event streaming platforms Observability Experience with one or more of: Grafana Prometheus Datadog OpenTelemetry Good to Have Experience building multi-agent AI systems using LangChain, LangGraph, CrewAI, or custom orchestration frameworks. Experience implementing Retrieval-Augmented Generation (RAG) pipelines. Experience migrating monolithic applications to microservices. Exposure to Kubernetes and Docker. Experience working in cloud environments such as AWS, Azure, or GCP. Background in fintech, payments, or other high-scale consumer platforms. Startup experience with ownership of end-to-end product features. What We're Looking For Strong problem-solving and system design skills. Passion for building scalable, resilient backend platforms. Ability to work independently in a fast-paced environment. Strong communication and collaboration skills. Ownership mindset with a focus on performance, reliability, and continuous improvement. Why Join Us Build next-generation AI infrastructure powering real-world applications. Solve complex distributed systems and scalability challenges. Work with cutting-edge AI technologies and modern cloud-native architectures. Take ownership of mission-critical systems with significant business impact. Collaborate with a high-performing engineering team focused on innovation and technical excellence.

One address, no account. We’ll tell you when matching roles go live.

More at Recro

Related open roles

View all roles