Source description
About the role
Site Reliability Engineer (SRE) Engineering Hybrid Remote, Pune, Maharashtra At Relatient, we help healthcare organizations optimize patient access through AI-powered workflows, real-time automation, and flexible access tools. We are trusted by over 50,000 providers to modernize the patient experience. Your Role at Relatient: Were seeking a Site Reliability Engineer (SRE) to join our team to lead production reliability, observability, and operational excellence across our cloud-hosted healthcare platform. Our office is in Amar Tech Park and brings in an amazing culture where we focus on work that makes a difference. This role is responsible for ensuring the availability, performance, scalability, and resiliency of mission-critical customer-facing platforms, including Patient Scheduling, Patient Engagement, Voice, Chat, Messaging, APIs, Reporting, and Healthcare Integrations. The SRE partners closely with Engineering, Infrastructure, Security, Product, Customer Support, and Implementation teams to proactively prevent customer-impacting incidents, improve production stability, automate operational processes, and continuously enhance the reliability of our cloud-native applications and services. This position is part of Relatients 24x7 Site Reliability Engineering organization, requiring participation in planned rotational shifts, including day, evening, overnight, weekend, and holiday coverage to support business-critical production systems. How You'll Make an Impact: Monitor the reliability, availability, performance, and operational health of Relatients customer-facing platforms and services. Respond to P1/P2 production incidents by participating in incident triage, troubleshooting, service restoration, and root cause analysis. Proactively detect, investigate, and resolve production issues before they impact customers through effective monitoring, observability, and operational reviews. Monitor and maintain observability, alerting, dashboards, SLIs, SLOs, and operational KPIs across applications and infrastructure. Monitor and troubleshoot cloud-native applications, APIs, databases, messaging platforms, third-party integrations, and production workloads. Ensure the health and availability of customer-facing services including Patient Scheduling, Voice, Chat, SMS, Email, Reporting, Healthcare Integrations, and EHR platforms. Support production cloud infrastructure, deployments, configurations, platform services, and operational readiness across production environments. Execute deployment validation activities including health checks, smoke testing, dependency validation, and post-deployment verification. Participate in release planning, change management, and production readiness activities to ensure stable production deployments. Collaborate with Engineering, Infrastructure, Security, Product, Customer Support, and third-party vendors to resolve production issues and improve service reliability. Develop and maintain operational runbooks, troubleshooting guides, recovery procedures, and technical documentation. Contribute to automation, monitoring optimization, operational efficiency, and continuous service improvements. Participate in post-incident reviews and support implementation of corrective and preventive actions. Support a culture of operational excellence, proactive ownership, continuous learning, and customer-first reliability. What You Bring: Bachelor's degree in computer science, B.E./ B.Tech, computer engineering, or a related technical discipline. 4+ years supporting enterprise SaaS or cloud-native production environments. Experience supporting production applications and customer-facing platforms. Experience working with cloud infrastructure, monitoring, observability, and incident management. Experience troubleshooting application, API, infrastructure, and database issues. Experience working within a 24x7 production support environment. Experience participating in Sev1/Sev2 incident response and root cause analysis. Experience supporting one or more of the following: Modern programming languages and application frameworks (Java, Spring Boot, PHP, Node.js, Angular, React) Cloud platforms and infrastructure services (AWS preferred) APIs, distributed systems, messaging platforms, and third-party integrations SQL and NoSQL databases, caching, and messaging technologies Observability, monitoring, logging, and distributed tracing platforms CI/CD, source control, containerization, automation, and DevOps practices Linux/Unix operating systems and scripting ITIL Incident, Problem, and Change Management processes Mindsets That Matter: We always look for ways to grow and take pride in what we do. You'll thrive here if you: Act with purpose, focus, and accountability Collaborate across teams and communicate clearly Keep innovating and automate what slows you down Disclaimer : This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
More at RELATIENT INC