Source description
About the role
Role Overview: As a Site Reliability Engineer with 7+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering in a production cloud environment, you will be responsible for ensuring the reliability and scalability of our systems. Your role will involve hands-on experience with AWS cloud services, managing Linux-oriented production environments, using Infrastructure-as-Code and GitOps best practices, operating and troubleshooting production Kubernetes environments, and applying AWS Well-Architected Framework principles. Additionally, you will work on cloud security best practices and have experience with PostgreSQL in production. Key Responsibilities: - Lead multi-person technical projects from scoping through delivery - Write automation scripts and tooling in Python, Go, or similar languages - Utilize observability tooling for metrics, logging, and distributed tracing to drive reliability - Implement data retention, backup, and recovery processes across cloud-native systems - Manage CI/CD pipelines, release management, and deployment automation - Understand service mesh, API gateway patterns, and microservices architectures - Use agentic coding assistants for engineering tasks and break down complex infrastructure tasks using AI - Evaluate AI-generated outputs and identify suboptimal or unsafe results - Lead technical discussions, drive alignment across engineering and product, and communicate decisions clearly - Mentor junior and mid-level engineers in technical skills and professional development - Operate independently with minimal supervision and make final technical decisions as DRI - Communicate effectively in English, both written and verbal, to influence cross-functional partners Qualifications Required: - 7+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering in a production cloud environment - 5+ years of hands-on experience with AWS cloud services - 5+ years managing Linux-oriented production environments at scale - 5+ years using Infrastructure-as-Code (Terraform, CDK, CloudFormation) and/or GitOps best practices - 3+ years operating and troubleshooting production Kubernetes environments - 3+ years applying AWS Well-Architected Framework principles - 3+ years in cloud security best practices including IAM, secrets management, network security, and compliance - 3+ years working with PostgreSQL in production: performance tuning, replication, backup, and recovery Please note that the JD does not include any additional details about the company. Role Overview: As a Site Reliability Engineer with 7+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering in a production cloud environment, you will be responsible for ensuring the reliability and scalability of our systems. Your role will involve hands-on experience with AWS cloud services, managing Linux-oriented production environments, using Infrastructure-as-Code and GitOps best practices, operating and troubleshooting production Kubernetes environments, and applying AWS Well-Architected Framework principles. Additionally, you will work on cloud security best practices and have experience with PostgreSQL in production. Key Responsibilities: - Lead multi-person technical projects from scoping through delivery - Write automation scripts and tooling in Python, Go, or similar languages - Utilize observability tooling for metrics, logging, and distributed tracing to drive reliability - Implement data retention, backup, and recovery processes across cloud-native systems - Manage CI/CD pipelines, release management, and deployment automation - Understand service mesh, API gateway patterns, and microservices architectures - Use agentic coding assistants for engineering tasks and break down complex infrastructure tasks using AI - Evaluate AI-generated outputs and identify suboptimal or unsafe results - Lead technical discussions, drive alignment across engineering and product, and communicate decisions clearly - Mentor junior and mid-level engineers in technical skills and professional development - Operate independently with minimal supervision and make final technical decisions as DRI - Communicate effectively in English, both written and verbal, to influence cross-functional partners Qualifications Required: - 7+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering in a production cloud environment - 5+ years of hands-on experience with AWS cloud services - 5+ years managing Linux-oriented production environments at scale - 5+ years using Infrastructure-as-Code (Terraform, CDK, CloudFormation) and/or GitOps best practices - 3+ years operating and troubleshooting production Kubernetes environments - 3+ years applying AWS Well-Architected Framework principles - 3+ years in cloud security best practices including IAM, secrets management, network security, and compliance - 3+ years working with PostgreSQL in production: performance
More at DEVRABBIT IT SOLUTIONS INC