Source description
About the role
Job Title: Site Reliability Engineer (SRE) Location: Bangalore Hybrid , India Reports to: Platform Operations Lead (UK) Overview We are looking for a highly motivated and technically skilled Site Reliability Engineer (SRE) to support the reliability, scalability, performance, and operational stability of Care ADHD s technology platforms. This role will work closely with the UK-based Platform Operations Lead and engineering teams across the UK and India to ensure our applications, infrastructure, and cloud environments are secure, resilient, and operating effectively 24/7/365. You will play a key role in supporting cloud infrastructure, deployment automation, observability, incident management, and platform reliability within a modern AWS-based engineering environment. This is an excellent opportunity for someone passionate about cloud operations, DevOps, automation, and reliability engineering within a growing digital healthcare organisation. Key Responsibilities Platform Reliability & Operations Support the operational reliability and availability of production and non-production environments : Monitor platform health, system performance, and operational stability Participate in incident management, troubleshooting, and service recovery activities Help identify and resolve infrastructure, deployment, and application issues Support operational readiness, resilience, and disaster recovery processes Contribute to maintaining high levels of platform uptime and service reliability Cloud Infrastructure & DevOps Support AWS cloud infrastructure and platform services Assist in maintaining and improving CI/CD pipelines and deployment automation Contribute to Infrastructure as Code implementations using Terraform or AWS CDK Work closely with engineering teams to improve deployment reliability and operational processes Support environment provisioning, configuration management, and release activities Help improve scalability, security, and performance across cloud environments Monitoring, Observability & Automation Support monitoring, logging, alerting, and observability solutions Assist with proactive monitoring and operational health checks Help automate operational and infrastructure tasks to reduce manual effort Contribute to platform performance optimisation and reliability improvements Support root cause analysis and post-incident reviews Collaboration & Engineering Support Work closely with: Platform Operations Lead Engineering teams QA teams DevOps and infrastructure teams Support engineering teams with deployment, operational, and infrastructure-related issues Participate in release activities and operational support processes Collaborate across UK and India teams to support platform operations and delivery Technology Environment AWS cloud infrastructure Kubernetes and containerised services Serverless architectures (AWS Lambda, API Gateway) CI/CD pipelines and deployment automation Terraform / AWS CDK Monitoring and observability platforms Node.js / TypeScript applications PostgreSQL and cloud-native databases Microservices and event-driven architectures What We are Looking For Experience 5+ years of experience in Site Reliability Engineering, DevOps, Cloud Operations, or Infrastructure Engineering Experience supporting cloud-based production systems Experience working within Agile engineering teams Experience collaborating with distributed or offshore teams is desirable Technical Skills Experience with: AWS cloud services and infrastructure CI/CD pipelines and deployment tooling Infrastructure as Code (Terraform or AWS CDK preferred) Monitoring and logging platforms Linux systems administration Docker and container technologies Git-based workflows Understanding of: Site Reliability Engineering principles Cloud infrastructure and scalability Incident management and operational support High availability and fault tolerance Networking and infrastructure fundamentals Experience with tools such as: Terraform GitHub Actions / GitLab CI / Jenkins CloudWatch Datadog / Grafana / Prometheus Docker / Kubernetes PagerDuty or similar incident tooling Nice to Have Experience supporting healthcare or regulated environments Experience with Kubernetes administration Familiarity with serverless AWS architectures Experience with scripting or automation (Python, Bash, or TypeScript) Exposure to security and vulnerability management practices Leadership Competencies Problem Solving Strong troubleshooting and analytical skills with the ability to investigate and resolve operational issues effectively. Communication Ability to communicate clearly with technical teams and operational stakeholders. Collaboration Works effectively across engineering, QA, and infrastructure teams in a distributed environment. Ownership Demonstrates accountability for platform reliability, operational support, and continuous improvement. What Success Looks Like Reliable and stable platform operations with reduced downtime Efficient and reliable deployment processes Strong monitoring and operational visibility across environments Fast and effective incident response and resolution Improved operational automation and reduced manual effort Strong collaboration between UK and India engineering and platform teams Why Join Care ADHD This is an opportunity to help operate and scale the technology platforms supporting a growing digital healthcare organisation focused on improving ADHD care and patient outcomes. You ll work within a collaborative engineering environment focused on cloud technology, automation, reliability engineering, and continuous improvement helping ensure our platforms remain secure, scalable, and always available.
More at Care Adhd