Source description
About the role
Monitor and troubleshoot production and QA systems to identify and resolve performance, scalability, and reliability issues proactively. Participate in the on-call rotation to provide 24/7 critical incident support for supported systems Design, implement, and maintain automated processes and tools to streamline deployment and release processes. Collaborate with cross-functional teams to define , document , and implement operational processes, best practices, and procedures. Implement and maintain system monitoring tools and dashboards to provide real-time insights into system performance and identify potential issues. Work closely with developers to identify and fix bugs and performance bottlenecks in the application code. Ensure that systems and infrastructure comply with security, compliance, and regulatory requirements. Continuously evaluate systems and processes to identify areas for improvement and implement changes as needed. Considerations for top Candidates : Bachelors degree in Computer Science , Information Technology, a related field, or equivalent experience. 13+ years of experience in site reliability engineering, DevOps, QA, SRE or a related field. Strong experience with Java Based solution s Experience with AWS infrastructure and services Experience with IaC solutions like Cloudformation and Terraform Experience with CI/CD solutions - Github , Azure DevOps Strong troubleshooting and critical thinking skills 6+ years of experience and proficiency in one or more programming languages, such as Python (preferred), Javascript (preferred). Solid understanding of networking, load balancing , on prem hosting solutions, and web application architectures. Experience with containerization technologies, such as Docker and Kubernetes. Excellent problem-solving skills and a strong attention to detail. Strong IT and Business communication skills and ability to collaborate effectively with cross-functional teams.
More at Caterpillar