Source description
About the role
As a Senior SRE at Level AI, you will play a crucial role at the intersection of backend engineering, infrastructure operations, and FinOps. This position goes beyond the responsibilities of a traditional DevOps engineer and involves a more hands-on approach than a pure architect. Key Responsibilities: - Infrastructure cost efficiency and FinOps: Take ownership of reducing Kubernetes overprovisioning, implement right-sizing programs, and manage cost telemetry for backend teams to make informed decisions. - GPU throughput optimization: Conduct structured experiments on on-premise GPU clusters, collaborate with AI service owners, and provide experimental bandwidth under the leadership of the Engineering team. - Backend enablement: Develop tooling, dashboards, and processes enabling backend teams from various groups to manage their cost and reliability budgets independently, focusing on providing leverage. - Reliability instrumentation: Ensure proper capture of instrumentation across new and offline flows to address cost-at-scale and reliability concerns effectively. - Selective security workstreams: Handle specific security tasks to prevent senior DevOps engineers from being the sole point of execution for security-related platform changes. Qualifications Required: - 4-5 years of hands-on systems experience with a strong judgment capability - Proficiency in backend engineering with production experience in Python, Go/Rust, and the ability to own services end-to-end - Extensive knowledge of Kubernetes at scale, including scheduler behavior, resource management, and cost-aware autoscaling - Experience in managing both cloud and on-premise infrastructure, fluency in GCP, IaC (Terraform), CI/CD, and operation in hybrid setups - Understanding of GPU workloads, including throughput profiling, inference server tuning, and GPU utilization metrics - Expertise in observability and reliability practices, such as metrics, traces, logs, and setting up SLOs - Demonstrated FinOps mindset with a track record of translating infrastructure decisions into measurable cost outcomes - Ability to handle platform-security workstreams independently without constant reliance on the DevOps team Level AI may utilize AI tools in the recruitment process to enhance efficiency but final hiring decisions are made by humans. If you seek further information on data processing, please reach out to us. As a Senior SRE at Level AI, you will play a crucial role at the intersection of backend engineering, infrastructure operations, and FinOps. This position goes beyond the responsibilities of a traditional DevOps engineer and involves a more hands-on approach than a pure architect. Key Responsibilities: - Infrastructure cost efficiency and FinOps: Take ownership of reducing Kubernetes overprovisioning, implement right-sizing programs, and manage cost telemetry for backend teams to make informed decisions. - GPU throughput optimization: Conduct structured experiments on on-premise GPU clusters, collaborate with AI service owners, and provide experimental bandwidth under the leadership of the Engineering team. - Backend enablement: Develop tooling, dashboards, and processes enabling backend teams from various groups to manage their cost and reliability budgets independently, focusing on providing leverage. - Reliability instrumentation: Ensure proper capture of instrumentation across new and offline flows to address cost-at-scale and reliability concerns effectively. - Selective security workstreams: Handle specific security tasks to prevent senior DevOps engineers from being the sole point of execution for security-related platform changes. Qualifications Required: - 4-5 years of hands-on systems experience with a strong judgment capability - Proficiency in backend engineering with production experience in Python, Go/Rust, and the ability to own services end-to-end - Extensive knowledge of Kubernetes at scale, including scheduler behavior, resource management, and cost-aware autoscaling - Experience in managing both cloud and on-premise infrastructure, fluency in GCP, IaC (Terraform), CI/CD, and operation in hybrid setups - Understanding of GPU workloads, including throughput profiling, inference server tuning, and GPU utilization metrics - Expertise in observability and reliability practices, such as metrics, traces, logs, and setting up SLOs - Demonstrated FinOps mindset with a track record of translating infrastructure decisions into measurable cost outcomes - Ability to handle platform-security workstreams independently without constant reliance on the DevOps team Level AI may utilize AI tools in the recruitment process to enhance efficiency but final hiring decisions are made by humans. If you seek further information on data processing, please reach out to us.
More at Level AI
Related open roles
Senior Site Reliability Engineer (Noida, BLR, India)
Bangalore
Senior Backend Engineer -PE
Bangalore · Delhi NCR · Hybrid
Forward Deployed Engineer - Agents(Remote)
Remote · India
CRM Solution Architect Salesforce
Delhi NCR
Senior Site Reliability Engineer (Noida, BLR, India)
Delhi NCR
Senior Site Reliability Engineer (Noida, BLR, India)
Bangalore · Delhi NCR · Hybrid
