Padmi

HCL Software Hiring For ClickHouse Expert + Application Performance Mo

Bangalore · Delhi NCRPosted 2 months ago
Infrastructure And DatabasesSeniorFull Time, Permanent
Apply at HCLTech

Opens the source posting on naukri.com

Source description

About the role

View original

Job Title: ClickHouse Expert + Application Performance Monitoring (APM) Specialist Experience: 8 to 10+ yrs Location: Noida, Bangalore Job Summary: ClickHouse Expert + APM Specialist with 7+ years of experience to architect, manage, and optimize high-performance ClickHouse deployments and enterprise-grade Application Performance Monitoring (APM) solutions. You will be the primary subject matter expert driving observability strategy, enabling real-time analytics at scale, and ensuring platform reliability across distributed systems. Key Responsibilities: ClickHouse Architecture & Administration: Design, deploy, and manage ClickHouse clusters; define schema design, partitioning strategies, sharding, and replication for high availability and performance. Query Optimization: Analyze and tune complex analytical queries, materialized views, and aggregations to achieve sub-second response times on large datasets. APM Implementation: Implement and own APM tooling (e.g., Datadog, Dynatrace, New Relic, Grafana + Tempo) across microservices and infrastructure to provide end-to-end visibility. Observability Strategy: Define and drive the observability roadmap covering metrics, logs, and traces; integrate with Prometheus, OpenTelemetry, and alerting pipelines. Performance Baselining & SLOs: Establish performance baselines, define SLIs/SLOs, build dashboards, and drive proactive capacity planning. Incident Analysis: Lead root cause analysis for performance degradations; correlate APM signals with infrastructure metrics and ClickHouse query telemetry. Data Ingestion Pipelines: Design and maintain high-throughput data ingestion pipelines (Kafka to ClickHouse via Vector, Fluent Bit, etc.) ensuring low-latency delivery. Collaboration: Work closely with product engineering, SRE, and DevOps teams to embed observability practices in CI/CD pipelines and release processes. Documentation & Standards: Maintain runbooks, architecture diagrams, and best-practice guidelines for ClickHouse operations and APM tooling. Required Skills & Qualifications: Experience: 7+ years in data engineering, backend, or SRE roles with at least 3 years of hands-on ClickHouse production experience. ClickHouse Expertise: Deep knowledge of ClickHouse internals: MergeTree engine families, TTLs, materialized views, dictionaries, distributed tables, and cluster management. APM Tools: Proficiency with one or more APM platforms: Datadog, Dynatrace, New Relic, AppDynamics, or open-source stacks (Grafana, Prometheus, Jaeger/Tempo). Observability Stack: Hands-on experience with OpenTelemetry, distributed tracing, structured logging, and metrics instrumentation across polyglot environments. Query & Schema Design: Strong SQL skills with expertise in ClickHouse-specific optimizations (ORDER BY keys, PREWHERE, skip indexes, sampling). Data Pipelines: Experience with streaming ingestion via Apache Kafka, Kafka Connect, Vector, Fluent Bit, or similar tools. Infrastructure: Familiarity with Kubernetes, Docker, and cloud platforms (AWS/GCP/Azure) for running stateful workloads at scale. Scripting: Proficiency in Python or Go for automation, tooling, and pipeline development. Alerting & On-Call: Experience configuring alerting pipelines (PagerDuty, OpsGenie) and participating in on-call rotations. Communication: Strong verbal and written communication skills; ability to present observability insights to technical and non-technical stakeholders. Nice to Have: Experience with TimescaleDB, InfluxDB, or other time-series databases; exposure to eBPF-based monitoring tools.

One address, no account. We’ll tell you when matching roles go live.

More at HCLTech

Related open roles

View all roles