Source description
About the role
As a Senior Application Site Reliability Engineer (SRE) with over 8 years of experience, your primary responsibility will be to ensure the reliability, performance, and operational excellence of cloud-native applications. You will be expected to bridge the gap between application development and cloud/platform reliability. Key Responsibilities: - Own application SLOs, SLIs, error budgets, and reliability metrics to maintain high standards of reliability. - Drive incident management, Root Cause Analysis (RCA), and Mean Time To Recovery (MTTR) reduction to enhance operational efficiency. - Enhance application performance, scalability, and resilience to meet the demands of a dynamic environment. - Design and implement observability tools such as logs, metrics, and traces for effective monitoring. - Automate deployments, recovery processes, and operational workflows to streamline operations. - Collaborate with application and platform teams as a reliability advisor to ensure best practices are followed. Qualifications Required: - Minimum of 8 years of experience in Application Engineering, Site Reliability Engineering (SRE), or Production Systems. - Demonstrated expertise in working with distributed, cloud-native applications. - Proficiency in at least one programming language to effectively contribute to development tasks. - Hands-on experience in working with any cloud environment to deploy and manage applications. - Strong communication and consulting skills to engage effectively with stakeholders. - Familiarity with Kubernetes, microservices architecture, and best practices in DevOps/SRE. - Proven experience in driving reliability practices and influencing teams towards operational excellence. As a Senior Application Site Reliability Engineer (SRE) with over 8 years of experience, your primary responsibility will be to ensure the reliability, performance, and operational excellence of cloud-native applications. You will be expected to bridge the gap between application development and cloud/platform reliability. Key Responsibilities: - Own application SLOs, SLIs, error budgets, and reliability metrics to maintain high standards of reliability. - Drive incident management, Root Cause Analysis (RCA), and Mean Time To Recovery (MTTR) reduction to enhance operational efficiency. - Enhance application performance, scalability, and resilience to meet the demands of a dynamic environment. - Design and implement observability tools such as logs, metrics, and traces for effective monitoring. - Automate deployments, recovery processes, and operational workflows to streamline operations. - Collaborate with application and platform teams as a reliability advisor to ensure best practices are followed. Qualifications Required: - Minimum of 8 years of experience in Application Engineering, Site Reliability Engineering (SRE), or Production Systems. - Demonstrated expertise in working with distributed, cloud-native applications. - Proficiency in at least one programming language to effectively contribute to development tasks. - Hands-on experience in working with any cloud environment to deploy and manage applications. - Strong communication and consulting skills to engage effectively with stakeholders. - Familiarity with Kubernetes, microservices architecture, and best practices in DevOps/SRE. - Proven experience in driving reliability practices and influencing teams towards operational excellence.
More at Fifthgen Tech Solution