Padmi

Site Reliability Engineer & AI & Data Platforms (LLM & Kubernetes)

ChennaiPosted 1 month ago
Infrastructure And DatabasesMid-levelFull Time; Regular
Apply at POWER COZMO

Opens the source posting on shine.com

Source description

About the role

View original

OverviewCandidate must be willing to travel to Amman, Jordan within 2-3 days upon selection. Experience 5+ Years (Mandatory) Job Summary: We are looking for a hands-on Support Engineer who will act as a Kubernetes maintainer and support engineer for our production AI and data systems. The role requires strong experience in data scraping / crawling systems. Kubernetes operations, Python-based automation, Github workflows and local LLM infrastructure. You will support mission-critical environments involving LLM workloads, data scraping pipelines, and containerized microservices. Key ResponsibilitiesDeploy, manage, and support production-grade Kubernetes clustersTroubleshoot pod, node, networking, and storage issues in live environmentsManage Helm charts, ConfigMaps, Secrets, and Kubernetes manifestsManage infrastructure and application deployments via Git repositoriesLocal LLM & AI Platform SupportTroubleshoot inference performance, memory issues, and model loading failuresSupport AI services exposed via REST APIs or internal microservicesSupport and maintain web scraping and crawling pipelinesDebug data extraction failures, rate limits, and anti-bot challengesEnsure data quality, consistency, and pipeline reliabilityAssist in scheduling and monitoring crawlers (cron jobs)Perform root cause analysis for production incidentsMonitor logs, metrics, and alerts using tools like Prometheus, GrafanaMaintain uptime and SLAs for AI and data platformsParticipate in on-call or escalation support rotationsAutomationMaintain documentation for support procedures and known issuesCollaborate with engineering teams for fixes and optimizationsRequired Skills & Qualifications (Mandatory)5+ years experience as a Data / AI Support EngineerHands-on experience with Docker and containerized workloadsExperience supporting local LLM models or AI inference systemsPractical knowledge of Linux system administrationExperience with data scraping or web crawling systemsBasic scripting skills (Python, Bash)Preferred / Good to HaveFamiliarity with Neo4j, Kafka, or data pipelinesExperience with Ollama, Hugging Face models, or similar LLM runtimesKnowledge of cloud-native monitoring toolsUnderstanding of networking concepts (DNS, ingress, load balancers)What We OfferOpportunity to work on cutting-edge AI and LLM infrastructure over cloud and local AI serverFast-paced startup or scale-up environment OverviewCandidate must be willing to travel to Amman, Jordan within 2-3 days upon selection. Experience 5+ Years (Mandatory) Job Summary: We are looking for a hands-on Support Engineer who will act as a Kubernetes maintainer and support engineer for our production AI and data systems. The role requires strong experience in data scraping / crawling systems. Kubernetes operations, Python-based automation, Github workflows and local LLM infrastructure. You will support mission-critical environments involving LLM workloads, data scraping pipelines, and containerized microservices. Key ResponsibilitiesDeploy, manage, and support production-grade Kubernetes clustersTroubleshoot pod, node, networking, and storage issues in live environmentsManage Helm charts, ConfigMaps, Secrets, and Kubernetes manifestsManage infrastructure and application deployments via Git repositoriesLocal LLM & AI Platform SupportTroubleshoot inference performance, memory issues, and model loading failuresSupport AI services exposed via REST APIs or internal microservicesSupport and maintain web scraping and crawling pipelinesDebug data extraction failures, rate limits, and anti-bot challengesEnsure data quality, consistency, and pipeline reliabilityAssist in scheduling and monitoring crawlers (cron jobs)Perform root cause analysis for production incidentsMonitor logs, metrics, and alerts using tools like Prometheus, GrafanaMaintain uptime and SLAs for AI and data platformsParticipate in on-call or escalation support rotationsAutomationMaintain documentation for support procedures and known issuesCollaborate with engineering teams for fixes and optimizationsRequired Skills & Qualifications (Mandatory)5+ years experience as a Data / AI Support EngineerHands-on experience with Docker and containerized workloadsExperience supporting local LLM models or AI inference systemsPractical knowledge of Linux system administrationExperience with data scraping or web crawling systemsBasic scripting skills (Python, Bash)Preferred / Good to HaveFamiliarity with Neo4j, Kafka, or data pipelinesExperience with Ollama, Hugging Face models, or similar LLM runtimesKnowledge of cloud-native monitoring toolsUnderstanding of networking concepts (DNS, ingress, load balancers)What We OfferOpportunity to work on cutting-edge AI and LLM infrastructure over cloud and local AI serverFast-paced startup or scale-up environment

One address, no account. We’ll tell you when matching roles go live.