Source description
About the role
As an Observability Subject Matter Expert & Engineer at Concierto, you will play a crucial role in defining and building the observability backbone of our multi-tenant AIOps platform. Your responsibilities will include owning the alert ingestion pipeline, metrics architecture, log management strategy, and distributed tracing framework. Here are the key responsibilities associated with this role: - Define standards for metrics collection, structured logging, distributed tracing, and alerting for cloud and on-premises workloads within Concierto. - Engineer microservice features in Node.js and/or Python to power alert ingestion, deduplication, correlation, and root-cause analysis. - Integrate and extend open-source observability stacks like Prometheus, Grafana, OpenTelemetry, Loki, Tempo, and Alertmanager within Kubernetes environments. - Design and implement alert correlation rules, noise-reduction algorithms, and intelligent routing logic to reduce MTTR for platform customers. - Define SLIs, SLOs, and error budgets for the platform and build dashboards to surface reliability signals to engineering and customer stakeholders. - Collaborate with AI/Bedrock engineering to feed observability telemetry into AIOps models for anomaly detection and predictive alerting. - Evaluate and recommend commercial or open-source tooling additions and lead POC implementations. - Act as the go-to SME for monitoring gaps, alerting fidelity, or observability integration questions. - Conduct knowledge-transfer sessions, write technical RFCs, and author platform observability documentation. In terms of required skills and qualifications, you should have expertise in Observability & Monitoring Domain, Product Engineering, and Cloud & Platform. Some of the specific requirements include: Observability & Monitoring Domain: - Expert-level knowledge of metrics, logs, and traces. - Hands-on experience with tools like Prometheus, Grafana, OpenTelemetry, Loki, Tempo, Jaeger, etc. - Proficiency in alert management tools like Alertmanager, PagerDuty, etc. Product Engineering: - Hands-on backend engineering experience in Node.js or Python for building production microservices. - Experience with event-driven pipelines using message brokers like Apache Kafka, AWS SQS/SNS, etc. - Proficiency with relational databases, API design, containerization, and Kubernetes workload management. Cloud & Platform: - Hands-on experience deploying and operating observability stacks on AWS and/or Azure. - Understanding of multi-tenant data isolation requirements in observability platforms. Additionally, it would be nice to have experience with AI/ML-enhanced observability, eBPF-based observability tools, commercial APM tools, SRE practices, and MSP or multi-tenant SaaS product environments. As an Observability SME & Engineer at Concierto, you will be expected to independently own observability features and be recognized internally as a domain expert. As an Observability Subject Matter Expert & Engineer at Concierto, you will play a crucial role in defining and building the observability backbone of our multi-tenant AIOps platform. Your responsibilities will include owning the alert ingestion pipeline, metrics architecture, log management strategy, and distributed tracing framework. Here are the key responsibilities associated with this role: - Define standards for metrics collection, structured logging, distributed tracing, and alerting for cloud and on-premises workloads within Concierto. - Engineer microservice features in Node.js and/or Python to power alert ingestion, deduplication, correlation, and root-cause analysis. - Integrate and extend open-source observability stacks like Prometheus, Grafana, OpenTelemetry, Loki, Tempo, and Alertmanager within Kubernetes environments. - Design and implement alert correlation rules, noise-reduction algorithms, and intelligent routing logic to reduce MTTR for platform customers. - Define SLIs, SLOs, and error budgets for the platform and build dashboards to surface reliability signals to engineering and customer stakeholders. - Collaborate with AI/Bedrock engineering to feed observability telemetry into AIOps models for anomaly detection and predictive alerting. - Evaluate and recommend commercial or open-source tooling additions and lead POC implementations. - Act as the go-to SME for monitoring gaps, alerting fidelity, or observability integration questions. - Conduct knowledge-transfer sessions, write technical RFCs, and author platform observability documentation. In terms of required skills and qualifications, you should have expertise in Observability & Monitoring Domain, Product Engineering, and Cloud & Platform. Some of the specific requirements include: Observability & Monitoring Domain: - Expert-level knowledge of metrics, logs, and traces. - Hands-on experience with tools like Prometheus, Grafana, OpenTelemetry, Loki, Tempo, Jaeger, etc. - Proficiency in alert management tools like Alertmanager, Page
More at Trianz
Related open roles
Linux OS Subject Matter Expert (Telangana)
Hyderabad
Technology Innovator & Prototyping Engineer
Bangalore
Technology Innovator & Prototyping Engineer
Bangalore
Salesforce Revenue Cloud Architect
Location not specified
Ai Research Bengaluru (India)
Bangalore
Technology Innovation & Prototyping Engineer
Bangalore