Padmi

Booking Holdings is Hiring Fresher Site Reliability Engineer I (SRE I)

BangalorePosted 2 months ago
Software engineeringNew gradFull Time; Regular
Apply at Booking Holdings

Opens the source posting on shine.com

Source description

About the role

View original

Booking Holdings (NASDAQ: BKNG) is a global leader in online travel and digital commerce, operating well-known brands such as Booking.com, KAYAK, Priceline, Agoda, Rentalcars.com, and OpenTable. The company operates in over 220 countries and is focused on making travel and commerce simple, reliable, and accessible for everyone. The Bangalore Center of Excellence supports global engineering teams by building scalable systems, improving reliability, and delivering high-performance infrastructure solutions across multiple business units. --- # Job Summary Booking Holdings is hiring a Site Reliability Engineer I (SRE I) to work on building and maintaining highly reliable, scalable, and performant systems. This role focuses on software engineering for infrastructure reliability, including system automation, cloud operations, monitoring, incident management, CI/CD pipelines, and performance optimization. The ideal candidate is passionate about DevOps, SRE, Cloud Computing, Distributed Systems, Kubernetes, AWS/GCP/Azure, Infrastructure Automation, Observability, and High-Availability Systems. --- # Key Responsibilities ## Software Engineering & Development - Build and maintain software applications supporting infrastructure and platform services. - Write clean, reusable, and maintainable code. - Refactor existing systems using modern design patterns. - Follow engineering best practices and testing standards. - Ensure application security, reliability, and data integrity. --- ## System Design & Architecture - Support evaluation of system architecture and design decisions. - Assist in designing scalable and cost-efficient solutions. - Understand system dependencies and infrastructure impact. - Participate in prototyping and technical solution design. --- ## End-to-End System Ownership - Monitor system health, performance, and availability. - Define and track service-level metrics (SLAs/SLOs). - Maintain production systems and ensure reliability. - Create and maintain runbooks and operational documentation. - Reduce operational risk and improve system stability. --- ## Incident Management - Respond to production incidents and minimize customer impact. - Perform root cause analysis (RCA). - Contribute to postmortem documentation. - Improve system reliability through long-term fixes. --- ## Automation & Reliability Engineering - Build automation to reduce manual operational work (toil). - Improve CI/CD pipelines and deployment workflows. - Write scripts and small services for automation. - Improve system scalability and efficiency. --- ## Observability & Monitoring - Monitor infrastructure, applications, and business metrics. - Improve logging, metrics, and alerting systems. - Support capacity planning and performance tuning. - Work with development teams to improve observability. --- ## Continuous Improvement - Identify system bottlenecks and performance issues. - Improve engineering processes and operational workflows. - Support implementation of best practices across teams. - Drive efficiency through automation and tooling. --- ## Communication & Collaboration - Work closely with product and engineering teams. - Communicate technical issues clearly. - Participate in cross-functional discussions. - Contribute ideas for system improvements. --- ## Architectural Support - Provide input on system design decisions. - Evaluate trade-offs between scalability, cost, and performance. - Ensure solutions align with enterprise architecture standards. --- # Required Qualifications - Bachelors degree in Computer Science, Engineering, or related field. - Strong foundation in software engineering principles. - Understanding of distributed systems and cloud computing. - Experience (academic or professional) with DevOps/SRE concepts. - Strong problem-solving and analytical thinking skills. - Good communication and teamwork abilities. --- # Technical Skills Candidates should have exposure to: - Linux Administration - Cloud Platforms (AWS / Azure / GCP) - CI/CD Pipelines - Git / Version Control Systems - Docker - Kubernetes - Scripting (Python / Bash / Go) - Infrastructure as Code (Terraform / Ansible) - Monitoring Tools (Prometheus, Grafana, ELK Stack) - Networking Basics (HTTP, DNS, TCP/IP) - Microservices Architecture - Distributed Systems Fundamentals - System Design Basics - Incident Management Tools --- # Preferred Skills - Site Reliability Engineering (SRE) experience - DevOps automation experience - Cloud-native architecture exposure - Performance tuning & optimization - Security best practices - High availability system design - Large-scale system operations - Observability engineering --- # Professional Competencies - Infrastructure Reliability Engineering - Cloud Operations - Automation Engineering - Platform Engineering - Incident Response - Root Cause Analysis - Performance Engineering - Scalability Engin

One address, no account. We’ll tell you when matching roles go live.

More at Booking Holdings

Related open roles

View all roles