Source description
About the role
Job Requirements About the Role: We are building large-scale, high-traffic, mission-critical products used globally. As an SRE II, you will play a key role in ensuring our systems are highly available, scalable, secure, and resilient. This is not a traditional support role. We expect SREs to think like product engineers automate everything, eliminate toil, improve system reliability, and proactively solve production challenges before they impact customers. What Youll Own Maintain and improve 99.9%+ uptime for production systems.Define and track SLOs, SLIs, and error budgets.Drive reliability engineering best practices across product teams.Reduce manual operational effort through automation.Improve deployment velocity while maintaining system stability.Lead incident response, postmortems, and preventive action planning. Key Responsibilities Reliability & Performance Ensure scalability and reliability of distributed microservices.Conduct performance tuning and capacity planning.Proactively identify and mitigate system bottlenecks. Infrastructure & Automation Design and manage cloud-native infrastructure (AWS preferred).Implement Infrastructure as Code (Terraform).Build and enhance CI/CD pipelines for safe and fast releases.Automate operational tasks using Python/Bash. Observability & Incident Management Build robust monitoring, logging, and alerting systems.Improve observability using Prometheus, Grafana, ELK, or Datadog.Participate in on-call rotation and drive blameless postmortems.Perform root cause analysis and implement long-term fixes. Security & Compliance Implement production security best practices.Collaborate with Security teams on vulnerability management.Ensure infrastructure compliance standards are met. Required Skills 3+ years in SRE / DevOps / Production Engineering in a product-based company.Strong experience with AWS (EC2, RDS, EKS, S3, IAM, VPC).Hands-on expertise in Kubernetes and container orchestration.Experience building and managing CI/CD pipelines.Strong scripting skills (Python / Bash).Experience with monitoring & logging tools.Solid understanding of networking, load balancing, and distributed systems.Experience handling high-scale production incidents. Good to Have Experience in high-growth SaaS or B2B platforms.Knowledge of Chaos Engineering principles.Experience working with high-traffic, multi-tenant architectures. Understanding of cost optimization in cloud environments. Job Requirements About the Role: We are building large-scale, high-traffic, mission-critical products used globally. As an SRE II, you will play a key role in ensuring our systems are highly available, scalable, secure, and resilient. This is not a traditional support role. We expect SREs to think like product engineers automate everything, eliminate toil, improve system reliability, and proactively solve production challenges before they impact customers. What Youll Own Maintain and improve 99.9%+ uptime for production systems.Define and track SLOs, SLIs, and error budgets.Drive reliability engineering best practices across product teams.Reduce manual operational effort through automation.Improve deployment velocity while maintaining system stability.Lead incident response, postmortems, and preventive action planning. Key Responsibilities Reliability & Performance Ensure scalability and reliability of distributed microservices.Conduct performance tuning and capacity planning.Proactively identify and mitigate system bottlenecks. Infrastructure & Automation Design and manage cloud-native infrastructure (AWS preferred).Implement Infrastructure as Code (Terraform).Build and enhance CI/CD pipelines for safe and fast releases.Automate operational tasks using Python/Bash. Observability & Incident Management Build robust monitoring, logging, and alerting systems.Improve observability using Prometheus, Grafana, ELK, or Datadog.Participate in on-call rotation and drive blameless postmortems.Perform root cause analysis and implement long-term fixes. Security & Compliance Implement production security best practices.Collaborate with Security teams on vulnerability management.Ensure infrastructure compliance standards are met. Required Skills 3+ years in SRE / DevOps / Production Engineering in a product-based company.Strong experience with AWS (EC2, RDS, EKS, S3, IAM, VPC).Hands-on expertise in Kubernetes and container orchestration.Experience building and managing CI/CD pipelines.Strong scripting skills (Python / Bash).Experience with monitoring & logging tools.Solid understanding of networking, load balancing, and distributed systems.Experience handling high-scale production incidents. Good to Have Experience in high-growth SaaS or B2B platforms.Knowledge of Chaos Engineering principles.Experience working with high-traffic, multi-tenant architectures. Understanding of cost optimization in cloud environments.
More at PHENOM PEOPLE INC