Source description
About the role
Role Overview: As a Site Reliability Engineer (SRE) with expertise in AWS DevOps, you will be responsible for managing production workloads on AWS, ensuring high availability, reliability, and scalability of systems. Your primary focus will be on implementing SRE principles, maintaining infrastructure as code, and building CI/CD pipelines. You will work closely with application teams to enhance observability, security, and cost optimization practices. Key Responsibilities: - Operate production workloads on AWS with 3-5 years of hands-on SRE/DevOps experience - Administer and troubleshoot Linux systems and demonstrate solid networking fundamentals - Implement SRE principles including SLIs/SLOs, error budgets, and incident management - Utilize containerization and orchestration tools such as Docker, EKS/ECS in production environments - Proficient in Infrastructure as Code with Terraform (preferred) or CloudFormation - Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline/CodeBuild - Enhance observability with metrics, logs, traces, alerting, and dashboards using tools like CloudWatch, Prometheus/Grafana, OpenTelemetry, or Datadog/New Relic - Implement AWS security best practices including IAM least privilege, KMS, VPC security groups/NACLs, Config/GuardDuty - Familiarity with cost optimization practices like tagging, rightsizing, Savings Plans/Reserved Instances - Automate tasks using scripting/programming in Python, Go, or Bash - Collaborate effectively with application teams and possess excellent communication skills Qualifications Required: - Bachelors degree in Computer Science/Engineering or equivalent practical experience (preferred) - AWS certification (e.g., Solutions Architect Associate or DevOps Engineer) is a plus - Experience working in a consulting environment is an added advantage Role Overview: As a Site Reliability Engineer (SRE) with expertise in AWS DevOps, you will be responsible for managing production workloads on AWS, ensuring high availability, reliability, and scalability of systems. Your primary focus will be on implementing SRE principles, maintaining infrastructure as code, and building CI/CD pipelines. You will work closely with application teams to enhance observability, security, and cost optimization practices. Key Responsibilities: - Operate production workloads on AWS with 3-5 years of hands-on SRE/DevOps experience - Administer and troubleshoot Linux systems and demonstrate solid networking fundamentals - Implement SRE principles including SLIs/SLOs, error budgets, and incident management - Utilize containerization and orchestration tools such as Docker, EKS/ECS in production environments - Proficient in Infrastructure as Code with Terraform (preferred) or CloudFormation - Build and maintain CI/CD pipelines using GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline/CodeBuild - Enhance observability with metrics, logs, traces, alerting, and dashboards using tools like CloudWatch, Prometheus/Grafana, OpenTelemetry, or Datadog/New Relic - Implement AWS security best practices including IAM least privilege, KMS, VPC security groups/NACLs, Config/GuardDuty - Familiarity with cost optimization practices like tagging, rightsizing, Savings Plans/Reserved Instances - Automate tasks using scripting/programming in Python, Go, or Bash - Collaborate effectively with application teams and possess excellent communication skills Qualifications Required: - Bachelors degree in Computer Science/Engineering or equivalent practical experience (preferred) - AWS certification (e.g., Solutions Architect Associate or DevOps Engineer) is a plus - Experience working in a consulting environment is an added advantage
More at Teamware Solutions