Padmi
Level AI logo
Level AI

contact center AI · conversational intelligence

Senior Site Reliability Engineer

IndiaPosted 3 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at Level AI

Opens the source posting on shine.com

Source description

About the role

View original

As a Senior SRE at Level AI, you will play a crucial role at the intersection of backend engineering, infrastructure operations, and FinOps. This position goes beyond the responsibilities of a traditional DevOps engineer and involves a more hands-on approach than a pure architect. Key Responsibilities: - Infrastructure cost efficiency and FinOps: Take ownership of reducing Kubernetes overprovisioning, implement right-sizing programs, and manage cost telemetry for backend teams to make informed decisions. - GPU throughput optimization: Conduct structured experiments on on-premise GPU clusters, collaborate with AI service owners, and provide experimental bandwidth under the leadership of the Engineering team. - Backend enablement: Develop tooling, dashboards, and processes enabling backend teams from various groups to manage their cost and reliability budgets independently, focusing on providing leverage. - Reliability instrumentation: Ensure proper capture of instrumentation across new and offline flows to address cost-at-scale and reliability concerns effectively. - Selective security workstreams: Handle specific security tasks to prevent senior DevOps engineers from being the sole point of execution for security-related platform changes. Qualifications Required: - 4-5 years of hands-on systems experience with a strong judgment capability - Proficiency in backend engineering with production experience in Python, Go/Rust, and the ability to own services end-to-end - Extensive knowledge of Kubernetes at scale, including scheduler behavior, resource management, and cost-aware autoscaling - Experience in managing both cloud and on-premise infrastructure, fluency in GCP, IaC (Terraform), CI/CD, and operation in hybrid setups - Understanding of GPU workloads, including throughput profiling, inference server tuning, and GPU utilization metrics - Expertise in observability and reliability practices, such as metrics, traces, logs, and setting up SLOs - Demonstrated FinOps mindset with a track record of translating infrastructure decisions into measurable cost outcomes - Ability to handle platform-security workstreams independently without constant reliance on the DevOps team Level AI may utilize AI tools in the recruitment process to enhance efficiency but final hiring decisions are made by humans. If you seek further information on data processing, please reach out to us. As a Senior SRE at Level AI, you will play a crucial role at the intersection of backend engineering, infrastructure operations, and FinOps. This position goes beyond the responsibilities of a traditional DevOps engineer and involves a more hands-on approach than a pure architect. Key Responsibilities: - Infrastructure cost efficiency and FinOps: Take ownership of reducing Kubernetes overprovisioning, implement right-sizing programs, and manage cost telemetry for backend teams to make informed decisions. - GPU throughput optimization: Conduct structured experiments on on-premise GPU clusters, collaborate with AI service owners, and provide experimental bandwidth under the leadership of the Engineering team. - Backend enablement: Develop tooling, dashboards, and processes enabling backend teams from various groups to manage their cost and reliability budgets independently, focusing on providing leverage. - Reliability instrumentation: Ensure proper capture of instrumentation across new and offline flows to address cost-at-scale and reliability concerns effectively. - Selective security workstreams: Handle specific security tasks to prevent senior DevOps engineers from being the sole point of execution for security-related platform changes. Qualifications Required: - 4-5 years of hands-on systems experience with a strong judgment capability - Proficiency in backend engineering with production experience in Python, Go/Rust, and the ability to own services end-to-end - Extensive knowledge of Kubernetes at scale, including scheduler behavior, resource management, and cost-aware autoscaling - Experience in managing both cloud and on-premise infrastructure, fluency in GCP, IaC (Terraform), CI/CD, and operation in hybrid setups - Understanding of GPU workloads, including throughput profiling, inference server tuning, and GPU utilization metrics - Expertise in observability and reliability practices, such as metrics, traces, logs, and setting up SLOs - Demonstrated FinOps mindset with a track record of translating infrastructure decisions into measurable cost outcomes - Ability to handle platform-security workstreams independently without constant reliance on the DevOps team Level AI may utilize AI tools in the recruitment process to enhance efficiency but final hiring decisions are made by humans. If you seek further information on data processing, please reach out to us.

One address, no account. We’ll tell you when matching roles go live.

More at Level AI

Related open roles

View all roles
Senior Site Reliability Engineer at Level AI · Padmi