Source description
About the role
What you bring in: 6+ years of hands-on experience in data engineering and large-scale distributed systems Proven expertise in building and maintaining complex ETL/ELT pipelines Deep knowledge of orchestration frameworks (e g , Airflow), and workflow optimization Strong cloud infrastructure experience (GCP preferred; AWS or Azure also relevant) Expert-level programming in Python or Scala; solid understanding of Spark internals Experience with CI/CD tools (e g , Jenkins, GitHub Actions) and infrastructure as code Familiarity with managing self-hosted tools like Spark or Airflow on Kubernetes Experience managing data warehouse in BigQuery or Redshift Strong communication skills and a proactive, problem-solving mindset The impact you will create: Design , build and optimize the data ingestion pipeline to reliably deliver billions of events daily in defined SLA Lead initiatives to improve scalability , performance and reliability Provide support for all product teams in building and optimizing their complex pipelines Identify and address pain points in the existing data platform; propose and implement high-leverage improvements Develop new tools and frameworks to streamline the data platform workflows Drive adoption of best practices in data and software engineering (testing, CI/CD, version control, monitoring) Work in close collaboration with data scientists and data analysts to help support their work in production Support production ML workflows and real-time streaming use cases Mentor other engineers and contribute to a culture of technical excellence and knowledge sharing It would be great if you also have: Working with messaging systems like Kafka , Redpanda
More at Truecaller