Source description
About the role
Role: Senior AI/ML Engineer – Agentic AI & LLM Systems Function: Artificial Intelligence / Machine Learning Engineering Location: India (Remote) Type: Full-time Industry: Information Technology & Services, Management Consulting About Company The company is a digital engineering firm founded in 2020, headquartered in Tampa, Florida. It specializes in AI-driven digital transformation for enterprises. With 450+ professionals across seven global offices, it operates in 25+ countries. The company has completed 55+ engagements spanning platform engineering, data analytics, quality engineering, and supply chain transformation. In 2023, it expanded its product engineering capabilities through an acquisition. The culture is ownership-driven, fast-paced, and built for engineers who want their work to ship and matter. Position Overview This role sits at the technical core of the company's AI practice. The engineer will own the architecture and hands-on development of enterprise-scale agentic AI systems, advanced RAG pipelines, and multi-agent orchestration platforms built to handle millions of daily interactions. The role also involves driving LLM selection, cost optimization, and observability standards, while mentoring a team of AI/ML engineers. The work spans the full stack — from vector retrieval and knowledge graphs to containerized cloud deployment and CI/CD for AI systems. Role & Responsibilities Architect end-to-end agentic AI platforms supporting millions of daily interactions with sub-100ms latency, including multi-agent orchestration frameworks coordinating 15+ specialized agents with state persistence, memory management, and error recovery Design and implement advanced RAG systems combining vector search, semantic search, knowledge graphs (Neo4j), and hybrid retrieval strategies using Pinecone, Weaviate, Milvus, Chroma, and FAISS Build production agentic AI applications using LangChain, LangGraph, CrewAI, AutoGen, and Azure ADK with ReAct patterns, Chain-of-Thought reasoning, and advanced function calling Develop and maintain MCP (Model Context Protocol) servers for tool standardization and agent-to-system communication; implement prompt optimization, semantic caching, and token-level cost reduction strategies Establish evaluation frameworks using RAGAS and custom metrics; implement observability via LangSmith and OpenTelemetry for hallucination tracking, retrieval quality, and model performance monitoring in production Deploy and manage containerized AI workloads on Kubernetes and Docker across AWS, Azure, or GCP; design CI/CD pipelines with automated testing and canary releases for AI model deployment Mentor junior and mid-level AI engineers, lead architectural reviews, and define technical standards for AI system reliability, observability, and cost governance Must Have Criteria 5+ years of backend software development with Python required; Java, Node.js, Go, or Rust acceptable as a secondary language — with strong REST API and microservices experience 2+ years of hands-on production experience building LLM and agentic AI systems at scale using LangChain and LangGraph Production experience with at least one agentic AI framework: CrewAI, AutoGen, or Azure ADK Hands-on experience with vector databases (Pinecone, Weaviate, Milvus, Chroma, or FAISS) and RAG architecture design including hybrid retrieval strategies Experience working with multiple LLM providers — OpenAI GPT-4, Anthropic Claude, Meta Llama, or Google Gemini — including model routing and trade-off evaluation Advanced Docker and Kubernetes expertise with production deployment experience on AWS, Azure, or GCP Experience with LLM observability platforms (LangSmith or equivalent) and OpenTelemetry-based distributed tracing Nice to Have Experience developing Model Context Protocol (MCP) servers and familiarity with emerging AI tooling standards Advanced LLM knowledge: fine-tuning, quantization, or knowledge distillation at scale Experience building knowledge graphs and graph-based retrieval systems using Neo4j or GraphQL Expertise in multimodal AI systems (text, image, or audio processing pipelines) Contributions to open-source AI/ML projects, published research, or AWS/Azure/GCP cloud certifications Track record of MLOps implementation and cost optimization at enterprise scale What We Offer Technical ownership of cutting-edge agentic AI and LLM systems with direct impact on millions of end users Access to the latest LLM models, AI research, and a professional development budget covering conferences and certifications Clear career trajectory toward Staff Engineer or Technical Leadership roles within a fast-growing global AI practice Collaborative engineering culture with a strong emphasis on ownership, speed, and delivering measurable client outcomes
More at recrew ai