Padmi

Sr. Site Reliability Engineer- Azure

IndiaPosted 1 month ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at XenonStack

Opens the source posting on shine.com

Source description

About the role

View original

Gathering Project Requirements from Stakeholders along with Business Analysts and Project ManagersBreak down complex problems and projects into manageable goalsHandle High severity incident and situation.Designing high level Schematics of the infrastructure, tools and process neededPerforming and in depth analysis of the possible risk and countermeasures for themCreate a bridge between development and operations by applying software engineering mindset to system administration topicsConfiguration management platform understanding and experience (Chef/Puppet/Ansible)Release engineering, which involves defining best practices to ensure software releases are consistent and repeatable.Alerting, being on-call, and troubleshooting, along with emergency and incident response and postmortems.Know how best to monitor systems and react when things go wrong, constantly writing and rewriting response playbooks to reduce the time to fix any breakdown which may occurInvolves documenting an incident, understanding all contributing root causes, and implementing future preventive actions.Highly developed skills in managing 24x7 production support comprising of Incident, Problem, Change managementTroubleshooting Support EscalationOn-Call Process OptimizationDocumenting KnowledgeOptimizing SDLC (Software Development Life Cycle)Technical Requirement - Strong understanding of cloud-based architecture and cloud operations. Hands-on experience with AzureExperience in administration/build/management of Linux systemsFoundational understanding of Infrastructure and Platform Technology stacksStrong understanding of Networking concepts and theories, such as different protocols (TCP/IP, UDP, ICMP, etc), MAC addresses, IP packets, DNS, OSI layers, and load balancingWorking knowledge of Infrastructure and Application monitoring platformsUnderstanding of the core DevOps practices (CI/CD pipeline, release management etc)Ability to write code using any one modern programming language (Python, JavaScript, Ruby etc). Additional scripting skills are preferredPrior experience in Cloud management automation tools (Terraform/CloudFormation etc) is preferredExperience with source code management software and API automation is preferred.Deep Understanding of architecture and operations of Container Orchestration tools eg KubernetesDeep understanding of Know Applications ie JAVA, Nodejs, GolangDeep understanding of Databases and SQLStrong understanding of BigData Infrastructure.Understanding of Incident management and Event Register ManagementKnowledge of SDLC methodologies and best practices including Waterfall Process, Agile methodologies, deployment automation, code reviews, and test-driven developmentProfessional Attributes - Excellent communication skillsAttention to detailAnalytical mind and Problem Solving AptitudeStrong Organizational skillsVisual Thinking Gathering Project Requirements from Stakeholders along with Business Analysts and Project ManagersBreak down complex problems and projects into manageable goalsHandle High severity incident and situation.Designing high level Schematics of the infrastructure, tools and process neededPerforming and in depth analysis of the possible risk and countermeasures for themCreate a bridge between development and operations by applying software engineering mindset to system administration topicsConfiguration management platform understanding and experience (Chef/Puppet/Ansible)Release engineering, which involves defining best practices to ensure software releases are consistent and repeatable.Alerting, being on-call, and troubleshooting, along with emergency and incident response and postmortems.Know how best to monitor systems and react when things go wrong, constantly writing and rewriting response playbooks to reduce the time to fix any breakdown which may occurInvolves documenting an incident, understanding all contributing root causes, and implementing future preventive actions.Highly developed skills in managing 24x7 production support comprising of Incident, Problem, Change managementTroubleshooting Support EscalationOn-Call Process OptimizationDocumenting KnowledgeOptimizing SDLC (Software Development Life Cycle)Technical Requirement - Strong understanding of cloud-based architecture and cloud operations. Hands-on experience with AzureExperience in administration/build/management of Linux systemsFoundational understanding of Infrastructure and Platform Technology stacksStrong understanding of Networking concepts and theories, such as different protocols (TCP/IP, UDP, ICMP, etc), MAC addresses, IP packets, DNS, OSI layers, and load balancingWorking knowledge of Infrastructure and Application monitoring platformsUnderstanding of the core DevOps practices (CI/CD pipeline, release management etc)Ability to write code using any one modern programming language (Python, JavaScript, Ruby etc). Additional scripting skills are preferredPrior experience in Cloud management automation tools (Terraform/Clo

One address, no account. We’ll tell you when matching roles go live.

More at XenonStack

Related open roles

View all roles