Padmi

Service Reliability Engineer

BangalorePosted 1 month ago
Software engineeringSeniorFull Time; Regular
Apply at WirelessCar

Opens the source posting on shine.com

Source description

About the role

View original

WirelessCars Journey WirelessCar drives the future of mobility by connecting vehicles and developing cutting-edge digital services. We help leading brandslike Volkswagen, Volvo, and Jaguar Land Roverinnovate, enhance mobility, and accelerate their digital transformation. Join us and shape the next era of automotive technology! We are looking for a Service Reliability Engineer (SRE) to join our Operations team. As an SRE, you'll be responsible for ensuring the reliability, availability, and performance of our connected vehicle services running in production. You'll combine software engineering principles with operational excellence to troubleshoot complex production issues, improve system resilience, and drive continuous improvements across our cloud-native platform. Working closely with developers, architects, and other stakeholders, you'll play a key role in delivering highly reliable services that support millions of connected vehicles. We offer you High-tech company that offers an exciting working environment Involvement with Electric Vehicle services and products that drive sustainability. Tech fund where you pick your choice of tech tools. Flat organizational culture founded on trust and autonomy; at WirelessCar, you are not just a number. Hybrid Work-life balance What You'll Do As a Service Reliability Engineer, you will: Monitor and maintain production services to ensure high availability, reliability, and performance. Lead and participate in incident response, including major incidents, with a focus on minimizing Mean Time to Recovery (MTTR). Coordinate stakeholders during incidents and communicate status updates clearly and effectively. Perform root cause analyses (RCA) and facilitate post-incident reviews (PIR) to drive long-term improvements. Identify recurring issues and implement preventive measures through problem management. Support change management activities while ensuring operational stability and service continuity. Troubleshoot complex issues across cloud infrastructure, container platforms, distributed systems, databases, APIs, and backend applications. Collaborate closely with DevOps teams to improve observability, automation, and operational readiness. Continuously improve monitoring, alerting, operational processes, and system reliability. Ensure compliance with security requirements, operational standards, and agreed service level agreements (SLAs). Participate in a 24/7 on-call rotation according to team schedules. What You'll Bring Strong Linux administration and troubleshooting skills. Experience with Kubernetes, Docker, and Helm. Solid understanding of AWS fundamentals. Hands-on experience with monitoring and observability platforms such as Dynatrace, Grafana, Prometheus, Splunk, or OpenTelemetry (OTEL). Experience troubleshooting Kafka and distributed systems. Database troubleshooting experience with PostgreSQL, Oracle, and/or MongoDB. REST API troubleshooting. Experience with CI/CD tools such as Jenkins or Concourse. Experience troubleshooting Java backend applications and JVM-based services. Nice to Have MQTT experience. AWS EKS. Infrastructure automation using Ansible or scripting. Experience working with connected vehicle platforms or automotive backend systems. Operational Experience We're looking for someone with experience in, Incident Management. Major Incident Management. Acting as Incident Commander during critical incidents. Problem Management. Root Cause Analysis (RCA). Post Incident Reviews (PIR). Change Management. SLA-driven operations. ITIL-based service management. Experience with the following is considered an advantage: Service transition and knowledge transfer. SAFe or Agile delivery. Vendor coordination within multi-supplier environments. Who You Are You are someone who, Takes ownership and drives issues through to resolution. Has a structured, analytical approach to troubleshooting. Remains calm and focused during high-pressure production incidents. Communicates clearly with both technical and non-technical stakeholders. Can coordinate multiple teams during critical situations. Makes sound decisions, even when information is incomplete. Is passionate about continuous improvement and operational excellence. Enjoys coaching and mentoring colleagues. Demonstrates technical leadership and strong stakeholder management skills. Thrives in collaborative, cross-functional teams. Qualifications Bachelor's degree in Computer Science, Computer Engineering, Information Technology, or a related field, or equivalent professional experience. 58 years of experience in Site Reliability Engineering, DevOps, Cloud Operations, .

One address, no account. We’ll tell you when matching roles go live.

More at WirelessCar

Related open roles

View all roles
Service Reliability Engineer at WirelessCar · Padmi