Source description
About the role
Congratulations, you have taken the first step towards bagging a career-defining role. Join the team of superheroes that safeguard data wherever it goes. Role: Site Reliability Engineer Experience: 2-5 years Location: Mumbai, India The Mission: Scale our infrastructure to support multi-region and multi-cloud with 99.99% availability. The Stack: Kubernetes, Terraform, AWS, GCP, Jenkins, Ansible, Bash/Python. The Culture: Blameless post-mortems, robust focus on automation (toil reduction), and dedicated deep-work time. Role Overview As a Senior Site Reliability Engineer in Seclore's Cloud Team, you will be responsible for designing, building, and operating highly available and scalable cloud infrastructure supporting Seclore's data-centric security platform. You will lead reliability engineering, performance optimization, monitoring, automation, and incident management across global cloud operations as our platform grows. Here's what you will get to explore - Maintain cloud environments (AWS, Azure, GCP) for Seclore's SaaS offerings with high availability, fault tolerance, and disaster recovery readiness. - Drive SLOs, SLIs, and error-budget practices to improve reliability and reduce downtime and latency issues. - Implement monitoring, logging, and tracing solutions (Prometheus, Grafana, ELK, OpenTelemetry) with robust alerting and runbooks. - Automate infrastructure provisioning (IaC: Terraform, CloudFormation), CI/CD deployments, configuration, and operational workflows. - Lead major incident responses, conduct blameless post-mortems, and drive long-term remediation. - Monitor resource usage, conduct capacity planning, and optimize cloud costs while maintaining performance. - Partner with security and product teams to ensure compliance, IAM best practices, network segmentation, and secure cloud architecture. - Collaborate with engineering, product, QA, and support teams to embed SRE principles across the organization. - Mentor junior team members and contr .
More at Seclore