Source description
About the role
Company overview: TraceLinks software solutions and Opus Platform help the pharmaceutical industry digitize their supply chain and enable greater compliance, visibility, and decision making. It reduces disruption to the supply of medicines to patients who need them, anywhere in the world. Founded in 2009 with the simple mission of protecting patients, today Tracelink has 8 offices, over 800 employees and more than 1300 customers in over 60 countries around the world. Our expanding product suite continues to protect patients and now also enhances multi-enterprise collaboration through innovative new applications such as MINT. Tracelink is recognized as an industry leader by Gartner and IDC, and for having a great company culture by Comparably. We are seeking a highly experienced Senior ML Engineer GenAI & ML Systems to lead the design, architecture, and implementation of advanced agentic AI systems within our next-generation supply chain platforms (SCP) This role is hands-on and execution-focused. You will design, build, deploy, and maintain large-scale multi-agent systems capable of reasoning, planning, and executing complex workflows in dynamic, non-deterministic environments. You will also own production concerns, including context management, knowledge orchestration, evaluation, observability, and system reliability. This position is ideal for a strong ML Engineer or Software Engineer with deep practical exposure to GenAI, data science, and modern ML systems, who is comfortable working end-to-endfrom architecture through production deployment. Experience in life sciences supply chain or other regulated environments is a strong plus. Key Responsibilities - Architect, implement, and operate large-scale agentic AI / GenAI systems that automate and coordinate complex supply chain workflows. - Design and build multi-agent systems, including agent coordination, planning, tool execution, long-term memory, feedback loops, and supervision. - Develop and maintain advanced context and knowledge management systems, including: - RAG and Advanced RAG pipelines - Hybrid retrieval, reranking, grounding, and citation strategies - Context window optimization and long-horizon task reliability - Own the technical strategy for reliability and evaluation of non-deterministic AI systems, including: - Agent evaluation frameworks - Simulation-based testing - Regression testing for probabilistic outputs - Validation of agent decisions and outcomes - Fine-tune and optimize LLMs/SLMs for domain performance, latency, cost efficiency, and task specialization (strong plus). - Design and deploy scalable backend services using Python and Java, ensuring production-grade performance, security, and observability. - Implement AI observability and feedback loops, including agent tracing, prompt/tool auditing, quality metrics, and continuous improvement pipelines. - Apply and experiment with reinforcement learning or iterative improvement techniques within GenAI or agentic workflows where appropriate. - Collaborate closely with product, data science, and domain experts to translate real-world supply chain requirements into intelligent automation solutions. - Guide system architecture across distributed services, event-driven systems, and real-time data pipelines using cloud-native patterns. - Mentor engineers, influence technical direction, and establish best practices for agentic AI and ML systems across teams. Required Qualifications - 6+ years of experience building and operating cloud-native SaaS systems on AWS, GCP, or Azure (minimum 5 years with AWS). - Strong ML Engineer / Software Engineer background with deep practical exposure to data science and GenAI systems. - Expert-level, hands-on experience designing, deploying, and maintaining large multi-agent systems in production. - Proven experience with advanced RAG and context management, including memory, state handling, tool grounding, and long-running workflows. - 6+ years of hands-on Python experience delivering production-grade systems. - Practical experience evaluating, monitoring, and improving non-deterministic AI behavior in real-world deployments. - Hands-on experience with agent frameworks such as LangGraph, AutoGen, CrewAI, Semantic Kernel, or equivalent. - Solid understanding of distributed systems, microservices, and production reliability best practices. Big Plus / Preferred Qualifications - Hands-on experience fine-tuning LLMs or SLMs for domain-specific tasks (training, evaluation, deployment). - Experience designing and deploying agentic systems in supply chain domains (logistics, manufacturing, planning, procurement). - Solid knowledge of knowledge organization techniques, including RAG, Advanced RAG, hybrid search, and reranking. - Experience applying reinforcement learning, reward modeling, or iterative .
More at Important Group
Related open roles
RTS- Kafka Engineer (Chennai)
Chennai
Lead Assistant Manager - Business Intelligence - Urgent Position (Gurugram)
Delhi NCR
Staff CPU Power Management Firmware Developer - Limits Management (Hyderabad)
Hyderabad
Machine Learning Developer (Freelance) (Pune)
Mumbai
REACT DEVELOPER - Start Now (Delhi)
Delhi NCR
Powerplatform Copilot Developer Product Owner (Pune)
Mumbai