Source description
About the role
What We Are Looking For Required Experience 7+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering in a production cloud setting. 5+ years of hands-on experience with AWS cloud services across compute, networking, storage, and security. 5+ years managing Linux-oriented production environments at scale. 5+ years using Infrastructure-as-Code (Terraform, CDK, CloudFormation) and/or GitOps best practices. 3+ years operating and troubleshooting production Kubernetes environments. 3+ years applying AWS Well-Architected Framework principles across reliability, security, performance, and cost pillars. 3+ years in cloud security best practices including IAM, secrets management, network security, and compliance. 3+ years working with PostgreSQL in production: performance tuning, replication, backup, and recovery. Demonstrated track record of leading multi-person technical projects from scoping through delivery. Technical Skills Solid general programming skills; comfort writing automation scripts and tooling in Python, Go, or similar. Deep knowledge of observability tooling metrics, logging, distributed tracing and how to use them to drive reliability. Solid understanding of data retention, backup, and recovery processes across cloud-native systems. Experience with CI/CD pipelines, release management, and deployment automation. Familiarity with service mesh, API gateway patterns, and microservices architectures. AI Fluency Proficient with agentic coding assistants (e.g., Cursor, Augment, GitHub Copilot) for day-to-day engineering tasks. Able to use AI to break down complex infrastructure tasks, accelerate design documentation, and improve code review quality. Ability to critically evaluate AI-generated outputs and identify when outputs are suboptimal or unsafe. Leadership & Collaboration Proven ability to lead technical discussions, drive alignment across engineering and product, and communicate decisions clearly to stakeholders. Experience mentoring junior and mid-level engineers in both technical skills and career development. Able to operate independently with minimal supervision; comfortable making final technical decisions as DRI. Strong communication skills in English written and verbal with experience influencing cross-functional partners. What We Are Looking For Required Experience 7+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering in a production cloud setting. 5+ years of hands-on experience with AWS cloud services across compute, networking, storage, and security. 5+ years managing Linux-oriented production environments at scale. 5+ years using Infrastructure-as-Code (Terraform, CDK, CloudFormation) and/or GitOps best practices. 3+ years operating and troubleshooting production Kubernetes environments. 3+ years applying AWS Well-Architected Framework principles across reliability, security, performance, and cost pillars. 3+ years in cloud security best practices including IAM, secrets management, network security, and compliance. 3+ years working with PostgreSQL in production: performance tuning, replication, backup, and recovery. Demonstrated track record of leading multi-person technical projects from scoping through delivery. Technical Skills Solid general programming skills; comfort writing automation scripts and tooling in Python, Go, or similar. Deep knowledge of observability tooling metrics, logging, distributed tracing and how to use them to drive reliability. Solid understanding of data retention, backup, and recovery processes across cloud-native systems. Experience with CI/CD pipelines, release management, and deployment automation. Familiarity with service mesh, API gateway patterns, and microservices architectures. AI Fluency Proficient with agentic coding assistants (e.g., Cursor, Augment, GitHub Copilot) for day-to-day engineering tasks. Able to use AI to break down complex infrastructure tasks, accelerate design documentation, and improve code review quality. Ability to critically evaluate AI-generated outputs and identify when outputs are suboptimal or unsafe. Leadership & Collaboration Proven ability to lead technical discussions, drive alignment across engineering and product, and communicate decisions clearly to stakeholders. Experience mentoring junior and mid-level engineers in both technical skills and career development. Able to operate independently with minimal supervision; comfortable making final technical decisions as DRI. Strong communication skills in English written and verbal with experience influencing cross-functional partners.
More at DEVRABBIT IT SOLUTIONS INC