Source description
About the role
Job Summary Key Duties & Responsibilities 1Collaborate with cross-functional teams to define and implement observability and reliability standards and best practices.2Design, deploy, and maintain the ELK stack for log aggregation, monitoring, and analysis.3Develop and maintain alerts and monitoring systems, ensuring early detection of issues and rapid incident response.4Create, customize, and maintain dashboards in Kibana for different stakeholders.5Collaborate with software development teams to identify performance bottlenecks and recommend solutions.6Automate manual tasks and workflows to streamline observability and reliability processes.78Generate and deliver detailed reports on system performance and reliability metrics.9Stay up to date with industry trends and best practices in observability and reliability engineering.Qualifications/Skills/Abilities Minimum Requirements Formal EducationBachelors degree in computer science, Information Technology, or a related field (or equivalent experience).Experience (type & duration)5+ years of experience in Site Reliability Engineering, Obervability & reliability, DevOpsSkills- Proficiency in configuring and maintaining the ELK stack (Elasticsearch, Logstash, Kibana) is mandatory. Strong scripting and automation skills, with expertise in Python, Bash, or similar languages. Experience in Data structures using Elasticsearch Indices. Experience in writing Data Ingestion Pipelines using Logstash. Experience with infrastructure as code (IaC) and configuration management tools (e.g., Ansible, Terraform). Handson and experience with cloud platforms ( AWS preferred) and containerization technologies (e.g., Docker, Kubernetes). Valuable to have Telecom domain expertise but not mandatory Solid problem-solving skills and the ability to troubleshoot complex issues in a production workplace. Excellent communication and collaboration skills. Accreditation/certifications/licensesRelevant certifications (e.g., Elastic C .