Padmi
Coupang logo
Coupang

e-commerce platform · rocket delivery

Staff Data Center Observability Engineer

BangalorePosted 2 months ago
Software engineeringStaff+Full Time; Regular
Apply at Coupang

Opens the source posting on shine.com

Source description

About the role

View original

As a Staff Data Center Observability and Site Reliability Engineer, you will own the design and operation of scalable observability platforms to ensure the reliability, performance, and availability of data center services. You will apply SRE best practices, automation, and performance optimization to deliver resilient infrastructure. This role partners closely with engineering teams and vendors to drive operational excellence while maintaining security and compliance standards. The core responsibilities for the job include the following: Observability and Monitoring: Design, implement, and maintain observability solutions for data center infrastructure.Develop, deploy, and maintain the operational and reliability components of a large-scale Observability and Telemetry collection platform, emphasizing performance at scale, real-time monitoring, logging, and alerting.Participate in and enhance the entire lifecycle of services, from inception and design to deployment, operation, and refinement.Develop and optimize monitoring systems to ensure high availability and performance.Create and manage dashboards, alerts, and reports to provide visibility into system health and performance. Site Reliability Engineering (SRE): Implement SRE best practices to improve the reliability, scalability, and performance of data center services.Develop and maintain automation scripts for infrastructure provisioning, monitoring, and management.Conduct root cause analysis and post-mortem reviews to prevent recurrence of incidents. Performance Optimization: Analyze and optimize the performance of data center systems and applications.Implement best practices for resource utilization and efficiency. Collaboration:Work closely with other engineering teams to understand and meet their observability and reliability requirements.Collaborate with hardware and software vendors to evaluate and integrate new technologies. Security and Compliance: Ensure that observability and reliability solutions comply with security policies and industry standards.Implement and maintain security measures to protect data and infrastructure. Troubleshooting and Support: Provide support for observability and reliability-related issues, including debugging and resolving hardware and software problems.Develop and maintain documentation for troubleshooting procedures and best practices. Continuous Improvement: Stay updated with the latest advancements in observability and SRE technologies and integrate them into the infrastructure.Continuously improve the reliability, scalability, and performance of data center services. Requirements: Education: Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field.Experience: 8-12 years of progressive software engineering experience, with a heavy emphasis on distributed systems, cloud-native architectures, or platform operations.Programming: Strong proficiency in Go or Python, with a deep understanding of networked systems and performance optimization.Orchestration: Expert-level knowledge of Kubernetes internals (scheduling and controllers) and containerization ecosystems.Traffic Management: Proven experience with load balancing, service mesh, and request routing at scale.Operational Excellence: A strong "ownership" mindset with a track record of maintaining mission-critical, high-availability systems in production. Preferred Qualifications: AI/ML Domain Knowledge: Prior experience building infrastructure specifically for LLM inference or large-scale training clusters.Low-Level Optimization: Familiarity with inference, including mixed precision, kernel tuning, or custom hardware accelerators.Public/Private Cloud: Experience managing hybrid-cloud or multi-AZ deployments across AWS, Azure, or GCP.Compliance: Experience operating in regulated environments with strict security and compliance requirements. As a Staff Data Center Observability and Site Reliability Engineer, you will own the design and operation of scalable observability platforms to ensure the reliability, performance, and availability of data center services. You will apply SRE best practices, automation, and performance optimization to deliver resilient infrastructure. This role partners closely with engineering teams and vendors to drive operational excellence while maintaining security and compliance standards. The core responsibilities for the job include the following: Observability and Monitoring: Design, implement, and maintain observability solutions for data center infrastructure.Develop, deploy, and maintain the operational and reliability components of a large-scale Observability and Telemetry collection platform, emphasizing performance at scale, real-time monitoring, logging, and alerting.Participate in and enhance the entire lifecycle of services, from inception and design to deployment, operation, and refinement.Develop and optimize monitoring systems to ensure high

One address, no account. We’ll tell you when matching roles go live.

More at Coupang

Related open roles

View all roles