Source description
About the role
As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn’t changed — we’re here to stop breaches, and we’ve redefined modern security with the world’s most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We're proud to work for a mission-driven company leveraging AI to transform the way we work. CrowdStrikers drive their careers through flexibility and autonomy while also being expected to contribute to a culture of responsible AI adoption, experimentation, and innovation. We use an AI-first mindset as a force multiplier to proactively and continuously accelerate execution, build expertise, uncover insights, and solve complex problems. We’re always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you. About the Role: At CrowdStrike, Site Reliability Engineering (SRE) is at the forefront of ensuring the reliability and scalability of our cloud-native security platform. In this role, you'll manage a team of talented engineers, providing technical leadership on key projects and empowering them to excel in their roles. As an SRE Manager, you will lead a team of SRE engineers ensuring the reliability, scalability, and performance of CrowdStrike's cloud-native security platform. You'll provide technical leadership and mentorship, owning both reliability engineering and software delivery pipelines - driving engineering velocity while maintaining zero tolerance for downtime in security-critical infrastructure. What you will Do Define and enforce SLOs, SLIs, and error budgets across distributed systems processing millions of events per second Drive system reliability by blending software engineering principles with AI-driven automation, moving from reactive firefighting to proactive, automated operations Lead major incident response and facilitate blameless postmortems, driving systemic reliability improvements Own capacity planning, traffic management, and load shedding strategies for high-throughput distributed systems Own the end-to-end software delivery pipeline strategy — designing, building, and maintaining scalable, reliable pipelines using Jenkins, GitLab CI, and Bitbucket Pipelines Build and maintain observability frameworks including metrics, distributed tracing, and log aggregation across the full stack Champion chaos engineering and resilience validation practices for security-critical systems Lead and grow a high-performing SRE team, mentoring engineers and fostering a culture of continuous learning and operational excellence Partner with cross-functional engineering teams to embed reliability practices early in the software development lifecycle What You'll Need Experience & Leadership Proven track record of building, growing, and retaining high-performing SRE/DevOps engineering teams in a fast-paced, high-growth environment 10+ years of software engineering experience with significant focus on reliability engineering, platform infrastructure, and production operations at scale 3+ years of hands-on management experience overseeing SRE/DevOps engineering teams, including incident command and reliability ownership Bachelor's degree in Computer Science or related field, or equivalent work experience Reliability Engineering Deep understanding of SRE principles including SLOs, SLAs, SLIs, and error budgeting strategies applied to large-scale distributed systems Proven experience owning reliability for high-throughput distributed systems processing millions of events per second, including capacity planning, traffic management, and load shedding strategies Strong incident management facilitating blameless postmortems, and driving system reliability improvements Demonstrated ability to build, operationalize, and maintain highly scalable, security-critical microservices-based distributed systems with zero tolerance for data loss or downtime. Advanced observability experience including Prometheus, Grafana, distributed tracing (Jaeger/OpenTelemetry), and large-scale log aggregation (ELK/Splunk) with a focus on building custom SLO dashboards and reliability scorecards. Experience owning disaster recovery strategies including backup automation, failover testing, and business continuity planning for stateful distributed systems Platform and Delivery Engineering Proficiency in Python and/or Golang for automation, tooling, and platform services Hands-on experience designing and managing scalable software delivery pipelines using Jenkins, GitLab CI, Bitbucket Pipelines, or equivalent Strong proficiency in Infrastructure as Code (IaC) - Terraform, Ansible, Pulumi, or equivalent Familiarity with GitOps workflows using ArgoCD or Flux for managing infrastructure deployments at scale Cloud and Big Data Exposure Proficiency in at least one cloud environment (AWS, Azure, GCP) with emphasis on multi-region architecture, cloud-native reliability patterns, and security-first cloud design Strong experience with Kubernetes at scale - managing large cluster fleets, workload orchestration, and container lifecycle management Familiarity with distributed data systems including relational databases (PostgreSQL), NoSQL (Cassandra), OLAP (Pinot), Indexing(OpenSearch) and real-time streaming platforms (Kafka, Flink) Exposure to Big Data and analytics technologies like Spark,Storm. #LI-AP1 Benefits of Working at CrowdStrike: Market leader in compensation and equity awards Comprehensive physical and mental wellness programs Competitive vacation and holidays for recharge Paid parental and adoption leaves Professional development opportunities for all employees regardless of level or role Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections Vibrant office culture with world class amenities Great Place to Work Certified™ across the globe CrowdStrike is proud to be an equal opportunity employer. We are committed to fostering a culture of belonging where everyone is valued for who they are and empowered to succeed. We support veterans and individuals with disabilities through our affirmative action program. CrowdStrike is committed to providing equal employment opportunity for all employees and applicants for employment. The Company does not discriminate in employment opportunities or practices on the basis of race, color, creed, ethnicity, religion, sex (including pregnancy or pregnancy-related medical conditions), sexual orientation, gender identity, marital or family status, veteran status, age, national origin, ancestry, physical disability (including HIV and AIDS), mental disability, medical condition, genetic information, membership or activity in a local human rights commission, status with regard to public assistance, or any other characteristic protected by law. We base all employment decisions--including recruitment, selection, training, compensation, benefits, discipline, promotions, transfers, lay-offs, return from lay-off, terminations and social/recreational programs--on valid job requirements. If you need assistance accessing or reviewing the information on this website or need help submitting an application for employment or requesting an accommodation, please contact us at recruiting@crowdstrike.com for further assistance. Find out more about your rights as an applicant. CrowdStrike participates in the E-Verify program. Notice of E-Verify Participation Right to Work CrowdStrike, Inc. is committed to fair and equitable compensation practices. Placement within the pay range is dependent on a variety of factors including, but not limited to, relevant work experience, skills, certifications, job level, supervisory status, and location. The base salary range for this position for all U.S. candidates is $140,000 - $215,000 per year, with eligibility for bonuses, equity grants and a comprehensive benefits package that includes health insurance, 401k and paid time off. For detailed information about the U.S. benefits package, please click here. CrowdStrike was founded in 2011 to fix a fundamental problem: The sophisticated attacks that were forcing the world’s leading businesses into the headlines could not be solved with existing malware-based defenses. Founder George Kurtz realized that a brand new approach was needed — one that combines the most advanced endpoint protection with expert intelligence to pinpoint the adversaries perpetrating the attacks, not just the malware. There’s much more to the story of how Falcon has redefined endpoint protection but there’s only one thing to remember about CrowdStrike: We stop breaches.
More at CrowdStrike
Related open roles
DevOps Engineer III
Tel Aviv
Sr. SRE Engineer II - EPICS, NG-SIEM (Hybrid, Bucharest)
Romania
Engineer III - TechOps CICD SRE (Reliability Focused)
Remote · India
Engineer II - SRE
Mumbai
Sr. Engineer, iAuto (Remote)
Remote · United Kingdom
Senior Linux Systems Engineer - Object Storage (2PM -11PM IST)
Bangalore