Source description
About the role
Role: AI Platform Engineer – GenAI / LLM Infrastructure Location: Hyderabad Experience: 1–4 years Job Description: Role Summary We are hiring AI Platform Engineers to design, build, and scale the core AI infrastructure and platform capabilities that power enterprise-grade AI solutions. This role focuses on: Building reusable AI infrastructure and pipelines Enabling scalable deployment of LLM-based systems Ensuring reliability, observability, and governance of AI systems
Key Responsibilities: 🔹 1. AI/ML Infrastructure & Pipeline Engineering Design and build end-to-end AI/ML pipelines , including: Data ingestion Data transformation Model interaction (LLMs / APIs) Output processing Implement scalable and modular pipelines for enterprise AI systems
-
LLM Platform & System Design Build foundational components for LLM-based systems , including: Prompt orchestration frameworks Retrieval pipelines (RAG infrastructure) Context management layers Enable standardized patterns for integrating multiple LLM providers (OpenAI, Azure, HuggingFace)
-
MLOps & AI Deployment Develop and maintain CI/CD pipelines for AI systems Enable: Automated deployment of AI models Version control for models and workflows Continuous evaluation and monitoring
Build systems for: Experiment tracking Model lifecycle management
-
System Reliability, Scalability & Performance Ensure AI systems meet enterprise-grade requirements for: Scalability High availability Low latency Optimize infrastructure for: Cost efficiency Performance of LLM-based workloads Implement failover, retry, and resilience mechanisms
-
Observability & Governance Design and implement monitoring systems for: Model performance Drift detection Latency and usage metrics Build guardrails for: Responsible AI usage Output validation and traceability
Ensure compliance with enterprise-grade governance requirements
- Reusable Platform Components Build reusable platform modules such as: AI service layers Model serving endpoints Workflow orchestration frameworks
Enable internal teams to build AI applications on top of standardized platform capabilities
-
Integration with Enterprise Ecosystems Enable AI systems to integrate seamlessly with: Enterprise applications Insurance platforms (e.g., Duck Creek ecosystem) Support “no data leaves environment” principles and secure deployment architectures
-
Collaboration & Platform Enablement Work closely with: AI Application Engineers Product Managers (Flarre) DevOps and Cloud teams
Enable broader engineering teams to build and deploy AI solutions on the platform
Qualifications: Core Engineering Strong Python programming
-
Experience with: Backend systems / APIs Data pipelines (ETL / processing frameworks)
-
AI Platform & MLOps Understanding of: ML lifecycle management CI/CD pipelines Model deployment strategies
-
Exposure to: LLM ecosystems (OpenAI / Azure / HuggingFace) API-based AI integration
-
Systems & Infrastructure Knowledge of: Distributed systems concepts System design fundamentals Familiarity with: Containerization (Docker) Orchestration tools (Kubernetes)
-
Good to Have Skills Experience with: Vector databases (Pinecone, FAISS) Workflow orchestration tools
-
Exposure to Cloud platforms (Azure / AWS / GCP) Understanding of Observability tools (monitoring/logging systems)
-
Domain Expertise (Preferred) Exposure to enterprise systems in: Insurance / BFSI domain Understanding of: Data security and compliance requirements Large-scale enterprise architecture
More at aggne