Padmi

SRE Consultant

BangalorePosted 30 days ago
Software engineeringSenior
Apply at RTown Technologies

Opens the source posting on jobs.pyjamahr.com

Source description

About the role

View original

Role : SRE Consultant Exp : 8-10 years Notice period : 0-15 days Mode of work : WFO Location : Bangalore Mandatory skills : AWS, Microsoft Azure, Iac, Sre, Site Reliability Engineering, Cloud Operations, software development, Golang, Ruby, Ruby Rails, automation, Cloud Infrastructure. SRE Consultant Job Description Overview The Site Reliability Engineer (SRE) Consultant plays a critical role in enhancing the reliability and performance of the organization’s software systems and services. This position is pivotal in bridging the gap between development and operations, ensuring seamless integration and deployment of applications. The SRE Consultant will utilize their expertise in software engineering and systems administration to build scalable and reliable systems. Their focus will be on improving the uptime and overall reliability of services, proactively identifying potential points of failure, and implementing effective solutions. By leveraging automation and monitoring tools, the SRE Consultant will work to optimize the existing infrastructure while cultivating a culture of operational excellence within the organization. This role is vital for driving customer satisfaction through efficient systems and reliable service delivery. Key Responsibilities Design and implement scalable and reliable infrastructure solutions. Develop and maintain tools for deployment, monitoring, and operations. Manage on-call operations to ensure quick recovery from system outages. Collaborate with development teams to ensure reliable feature deployments. Identify and resolve performance bottlenecks in systems and applications. Establish and enhance service-level objectives (SLOs) and indicators (SLIs). Continuously improve monitoring and alerting strategies to minimize downtime. Conduct root cause analysis for incidents and implement preventive measures. Automate manual processes to increase reliability and efficiency. Implement CI/CD pipelines to facilitate rapid and reliable code deployments. Optimize resource utilization and capacity planning to manage loads effectively. Perform regular system assessments to identify vulnerabilities and required improvements. Provide training and mentorship to junior team members on SRE best practices. Document operational procedures and system designs for future reference. Stay updated with the latest industry trends and technologies to recommend improvements. Required Qualifications Bachelor’s degree in Computer Science, Engineering, or a related field. Minimum of 5 years of experience in site reliability engineering or a related role. Proficiency in cloud platforms like AWS, Google Cloud, or Azure. Experience with container orchestration systems such as Kubernetes or Docker. Strong skills in scripting languages like Python, Bash, or Go. In-depth knowledge of networking protocols and services. Experience with monitoring and logging tools such as Prometheus, Grafana, or ELK stack. Familiarity with configuration management tools like Terraform, Ansible, or Chef. Strong understanding of CI/CD processes and tools like Jenkins or GitLab CI. Proven experience in incident management and response best practices. Ability to analyze system performance and perform data-driven optimizations. Excellent problem-solving skills and ability to work under pressure. Strong communication skills for collaboration across teams. Knowledge of Agile methodologies and DevOps principles. Relevant certifications (AWS Certified DevOps Engineer, Google Professional Cloud DevOps Engineer, etc.) are a plus.

One address, no account. We’ll tell you when matching roles go live.

More at RTown Technologies

Related open roles

View all roles
SRE Consultant at RTown Technologies · Padmi