Padmi

SRE AWS

Delhi NCRPosted 2 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at Mobilution It Systems

Opens the source posting on shine.com

Source description

About the role

View original

Role Overview As a Site Reliability Engineer focused on AWS, you will design, automate, and operate highly available, scalable, and secure cloud platforms. You will codify infrastructure using Terraform, run workloads on Kubernetes (EKS), and implement GitOps and CI/CD practices to deliver changes safely. You will establish observability, reliability objectives, and incident response while embedding security, audit, and compliance controls. The role includes mentoring engineers, partnering with stakeholders, and planning delivery to meet business outcomes. Required Qualifications 7+ years in Infrastructure Engineering/SRE/DevOps roles Extensive hands-on experience with AWS (e.g., EC2, EKS, VPC, IAM, RDS, ALB/NLB, CloudWatch, CloudTrail) Deep expertise with Infrastructure as Code using Terraform (modules, state backends, workspaces) Strong experience with Kubernetes (preferably EKS) and related tools (Helm, Kustomize, Argo CD) Solid understanding of Git, branching models (trunk-based, GitFlow), CI/CD pipelines, and deployment strategies (blue/green, canary) Experience implementing security, audit, and compliance best practices (least privilege IAM, encryption, secrets, logging) SRE practices including monitoring/observability, SLIs/SLOs, incident response, and postmortems Software development/scripting experience, preferably Python and Terraform Experience mentoring engineers, team formation, fostering ownership and self-organization Experience with client relationship management and project planning Exposure to machine learning infrastructure (e.g., EKS with GPUs, SageMaker integration) Relevant certifications such as Certified Kubernetes Administrator (CKA), AWS Associate certifications (Developer, Machine Learning Engineer, Data Engineer) Responsibilities Architect, build, and operate AWS infrastructure for high availability, scalability, performance, and resilience Implement and maintain Infrastructure as Code with Terraform, including reusable modules, remote state, and automated pipelines Deploy and operate Kubernetes (EKS) clusters; manage application releases using Helm/Kustomize; implement GitOps with Argo CD Design, implement, and optimize CI/CD pipelines and safe deployment strategies (blue/green, canary, progressive delivery) Embed security, audit, and compliance controls across cloud and Kubernetes environments (IAM, encryption, policies, logging) Establish observability (metrics, logs, traces), define SLIs/SLOs, and lead incident response and postmortems Plan and execute upgrades, capacity/cost optimization, and platform migrations with minimal downtime Mentor and coach engineers; foster self-organization and ownership; collaborate with clients and stakeholders; support project planning and estimations Role Overview As a Site Reliability Engineer focused on AWS, you will design, automate, and operate highly available, scalable, and secure cloud platforms. You will codify infrastructure using Terraform, run workloads on Kubernetes (EKS), and implement GitOps and CI/CD practices to deliver changes safely. You will establish observability, reliability objectives, and incident response while embedding security, audit, and compliance controls. The role includes mentoring engineers, partnering with stakeholders, and planning delivery to meet business outcomes. Required Qualifications 7+ years in Infrastructure Engineering/SRE/DevOps roles Extensive hands-on experience with AWS (e.g., EC2, EKS, VPC, IAM, RDS, ALB/NLB, CloudWatch, CloudTrail) Deep expertise with Infrastructure as Code using Terraform (modules, state backends, workspaces) Strong experience with Kubernetes (preferably EKS) and related tools (Helm, Kustomize, Argo CD) Solid understanding of Git, branching models (trunk-based, GitFlow), CI/CD pipelines, and deployment strategies (blue/green, canary) Experience implementing security, audit, and compliance best practices (least privilege IAM, encryption, secrets, logging) SRE practices including monitoring/observability, SLIs/SLOs, incident response, and postmortems Software development/scripting experience, preferably Python and Terraform Experience mentoring engineers, team formation, fostering ownership and self-organization Experience with client relationship management and project planning Exposure to machine learning infrastructure (e.g., EKS with GPUs, SageMaker integration) Relevant certifications such as Certified Kubernetes Administrator (CKA), AWS Associate certifications (Developer, Machine Learning Engineer, Data Engineer) Responsibilities Architect, build, and operate AWS infrastructure for high availability, scalability, performance, and resilience Implement and maintain Infrastructure as Code with Terraform, including reusable modules, remote state, and automated pipelines Deploy and operate Kubernetes (EKS) clusters; manage application releases using Helm/Kustomize; implement GitOps with Argo CD Design, implement, and optimize CI/CD pipelines and safe deployment st

One address, no account. We’ll tell you when matching roles go live.

More at Mobilution It Systems

Related open roles

View all roles