Padmi

Site Reliability Architect

MumbaiPosted 6 months ago
Software engineeringSenior
Apply at LOTUSFLARE INC

Opens the source posting on naukri.com

Source description

About the role

View original

Operational Empathy & Developer Enablement Partner closely with feature and product engineers to embed observability into the development lifecycle Translate complex logs and telemetry into clear, actionable Grafana dashboards that help teams understand system behavior and blast radius Full-Stack Observability & Forensics Lead distributed tracing initiatives by correlating frontend exceptions (Sentry) with backend logs and traces (VictoriaLogs/OpenSearch), enabling a seamless end-to-end (north-to-south) view of system health Telemetry Gap Identification & Instrumentation Proactively identify blind spots in logging, metrics, and traces Implement custom instrumentation across service layers to capture high-cardinality data while maintaining system performance Incident Response Automation Design and build Python-based automation tools to reduce Time to Truth during incidents by automating log aggregation, telemetry correlation, and diagnostic reporting Service Level Ownership Advocate for meaningful reliability metrics by defining and refining SLIs and SLOs that truly reflect user experience and satisfaction, balancing rapid feature delivery with long-term system stability Job Requirements Technical Skills & Experience Expert-level experience with OpenSearch and VictoriaLogs, including indexing strategies and optimization for high-volume log ingestion and querying Strong hands-on expertise with Grafana, building intuitive dashboards that clearly communicate system behavior and incident patterns Python: Advanced proficiency for building automation scripts, diagnostic tooling, and observability glue code TypeScript & Lua: Working familiarity required You should be comfortable reading and understanding service codebases to trace request flows end-to-end (deep expertise not required on day one) Experience with Sentry, including performance monitoring and profiling capabilities to proactively identify regressions and bottlenecks Additional Expectations Strong analytical and problem-solving mindset with a passion for system reliability and visibility Ability to collaborate effectively with application engineers, platform teams, and incident responders Clear communication skills to translate complex system data into actionable insights for diverse stakeholders

One address, no account. We’ll tell you when matching roles go live.

More at LOTUSFLARE INC

Related open roles

View all roles