Padmi

Site Reliability Engineer - Level - 2

Delhi NCRPosted 3 months ago
Software engineeringMid-levelFull Time; Regular
Apply at CorroHealth Infotech Private Limited

Opens the source posting on shine.com

Source description

About the role

View original

You will be responsible for ensuring system reliability, scalability, and performance across cloud platforms, while driving automation, observability, and operational excellence. Your essential duties and responsibilities include: - Designing, deploying, and optimizing workloads on AWS and Azure - Managing Kubernetes clusters and serverless solutions like AWS Lambda - Implementing secure, scalable, and cost-optimized infrastructure - Building automation frameworks and operational tooling using Python - Implementing Infrastructure as Code using Terraform, CloudFormation, or Azure Bicep - Automating deployments, scaling, monitoring, and remediation - Developing and maintaining CI/CD pipelines - Partnering with development teams to improve release velocity and reliability - Implementing end-to-end observability practices: metrics, logging, tracing - Managing logging & monitoring solutions - Establishing alerting & escalation workflows - Participating in 24/7 on-call rotations for critical systems - Diagnosing, resolving, and performing root cause analysis for incidents - Driving post-incident reviews and continuous improvement in reliability - Working closely with Dev, QA, and Security teams to ensure production readiness - Championing SRE best practices across teams - Mentoring junior engineers on cloud-native operations and observability Your required skills and qualifications include: - Experience: 36 years in SRE, DevOps, or Cloud Engineering - Strong experience with AWS and Azure - Hands-on experience with Kubernetes and containerization - Proficiency in Python - Experience with CI/CD tools - Proficiency with Terraform / CloudFormation / Bicep - Experience with observability tools - Familiarity with incident management tools - Solid understanding of networking & security - Strong analytical, troubleshooting, and collaboration skills Preferred qualifications: - Certifications in AWS, Microsoft Azure, or CKA - Experience with multi-cloud deployments - Familiarity with service mesh and advanced observability This job description is intended as a guideline and part of your function. The company has reviewed this description to ensure essential functions and duties are included. Additional functions and requirements may be assigned as deemed appropriate by supervisors. You will be responsible for ensuring system reliability, scalability, and performance across cloud platforms, while driving automation, observability, and operational excellence. Your essential duties and responsibilities include: - Designing, deploying, and optimizing workloads on AWS and Azure - Managing Kubernetes clusters and serverless solutions like AWS Lambda - Implementing secure, scalable, and cost-optimized infrastructure - Building automation frameworks and operational tooling using Python - Implementing Infrastructure as Code using Terraform, CloudFormation, or Azure Bicep - Automating deployments, scaling, monitoring, and remediation - Developing and maintaining CI/CD pipelines - Partnering with development teams to improve release velocity and reliability - Implementing end-to-end observability practices: metrics, logging, tracing - Managing logging & monitoring solutions - Establishing alerting & escalation workflows - Participating in 24/7 on-call rotations for critical systems - Diagnosing, resolving, and performing root cause analysis for incidents - Driving post-incident reviews and continuous improvement in reliability - Working closely with Dev, QA, and Security teams to ensure production readiness - Championing SRE best practices across teams - Mentoring junior engineers on cloud-native operations and observability Your required skills and qualifications include: - Experience: 36 years in SRE, DevOps, or Cloud Engineering - Strong experience with AWS and Azure - Hands-on experience with Kubernetes and containerization - Proficiency in Python - Experience with CI/CD tools - Proficiency with Terraform / CloudFormation / Bicep - Experience with observability tools - Familiarity with incident management tools - Solid understanding of networking & security - Strong analytical, troubleshooting, and collaboration skills Preferred qualifications: - Certifications in AWS, Microsoft Azure, or CKA - Experience with multi-cloud deployments - Familiarity with service mesh and advanced observability This job description is intended as a guideline and part of your function. The company has reviewed this description to ensure essential functions and duties are included. Additional functions and requirements may be assigned as deemed appropriate by supervisors.

One address, no account. We’ll tell you when matching roles go live.

More at CorroHealth Infotech Private Limited

Related open roles

View all roles