Padmi

Cloud Systems Architect

IndiaPosted 3 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at HIRED

Opens the source posting on shine.com

Source description

About the role

View original

As a Cloud Systems Architect working remotely, you will be responsible for designing, implementing, and maintaining scalable infrastructure using Linux, Kubernetes, and Prometheus. Your role will involve monitoring system health, analyzing performance metrics, and proactively addressing bottlenecks or potential failures to ensure high system availability. Automating operational processes, responding swiftly to incidents, and collaborating closely with development and operations teams are key aspects of this role. Key Responsibilities: - Design, implement, and maintain scalable infrastructure using Linux, Kubernetes, and Prometheus for seamless deployments and high system availability. - Monitor system health, analyze performance metrics, and proactively address bottlenecks or potential failures to increase system reliability. - Automate operational processes to minimize manual intervention, respond swiftly to incidents, conduct root cause analysis, and drive continuous improvements in incident response procedures. - Collaborate closely with development and operations teams to deliver seamless deployments and high system availability, creating comprehensive documentation and clear runbooks for operational excellence. - Respond to incidents, conduct root cause analysis, and drive continuous improvements in incident response procedures to ensure high system availability and minimize downtime. Qualifications Required: - Proven experience in designing, implementing, and maintaining scalable infrastructure using Linux, Kubernetes, and Prometheus. - Strong understanding of automation tools and technologies, with experience in automating operational processes. - Excellent problem-solving skills and the ability to analyze complex system issues and develop effective solutions. - Strong communication and collaboration skills to work closely with cross-functional teams. - Experience in creating comprehensive documentation and clear runbooks for operational excellence. About the Company: This role offers you a unique opportunity to work with a global leader in the AI industry, contributing to the development of cutting-edge AI models and driving innovation. You will have the chance to collaborate with experts on a global scale and shape the future of next-generation AI systems. Equal Opportunity Employer: The company values skills and expertise above all else, welcoming all qualified candidates regardless of background or prior employment history. Applications are evaluated based on demonstrated technical ability and qualifications. Apply Now to be a part of this exciting opportunity! As a Cloud Systems Architect working remotely, you will be responsible for designing, implementing, and maintaining scalable infrastructure using Linux, Kubernetes, and Prometheus. Your role will involve monitoring system health, analyzing performance metrics, and proactively addressing bottlenecks or potential failures to ensure high system availability. Automating operational processes, responding swiftly to incidents, and collaborating closely with development and operations teams are key aspects of this role. Key Responsibilities: - Design, implement, and maintain scalable infrastructure using Linux, Kubernetes, and Prometheus for seamless deployments and high system availability. - Monitor system health, analyze performance metrics, and proactively address bottlenecks or potential failures to increase system reliability. - Automate operational processes to minimize manual intervention, respond swiftly to incidents, conduct root cause analysis, and drive continuous improvements in incident response procedures. - Collaborate closely with development and operations teams to deliver seamless deployments and high system availability, creating comprehensive documentation and clear runbooks for operational excellence. - Respond to incidents, conduct root cause analysis, and drive continuous improvements in incident response procedures to ensure high system availability and minimize downtime. Qualifications Required: - Proven experience in designing, implementing, and maintaining scalable infrastructure using Linux, Kubernetes, and Prometheus. - Strong understanding of automation tools and technologies, with experience in automating operational processes. - Excellent problem-solving skills and the ability to analyze complex system issues and develop effective solutions. - Strong communication and collaboration skills to work closely with cross-functional teams. - Experience in creating comprehensive documentation and clear runbooks for operational excellence. About the Company: This role offers you a unique opportunity to work with a global leader in the AI industry, contributing to the development of cutting-edge AI models and driving innovation. You will have the chance to collaborate with experts on a global scale and shape the future of next-generation AI systems. Equal Opportunity Employer: The company values skills and expertis

One address, no account. We’ll tell you when matching roles go live.

More at HIRED

Related open roles

View all roles