Source description
About the role
Requirements Should have a total of 3-5 years of experience in Data Engineering. Should have hands-on experience in distributed computing. Should have working experience in Data Architecture design. Strong in programming languages like Python and Java. Should be aware of storage and compute options, and when to choose what. Must have one cloud hands-on experience. Proficient in either Apache Spark, Apache Beam or Apache Flink. Advanced SQL knowledge is mandatory. Must be aware of streaming concepts like Windowing, Late arrival, Triggers, etc. Should have exposure to GCP tools to develop an end-to-end data pipeline for various scenarios (including ingesting data from traditional databases as well as integration of API based data sources). Should have a business mindset to understand data and how it will be used for BI and Analytics purposes. Should have working experience on CI/CD pipelines, Deployment methodologies, and Infrastructure as Code. Should have advance understanding of different file formats, e. g., AVRO, Parquet, ORC, CSV and when to choose what. Good to have Hands-on experience on Kubernetes. Good to have a vector-based database like Qdrant. Should have a good understanding of Cluster Optimisation/ Pipeline Optimisation strategies. This job was posted by Rahul Kiran from Falabella.
More at Falabella India