Source description
About the role
● Should have experience in Data and Analytics and overseen end-to-end implementation of data pipelines on cloud-based data platforms. ● Strong programming skills in Python, Pyspark and some combination Java, Scala (good to have) ● Deep understanding of modern data processing technology stacks: Spark, HBase, Hive and other Hadoop ecosystem technologies. Development using Scala. ● Experience writing SQL, Structuring data, and data storage practices. ● Experience in Pyspark for Data Processing and transformation. ● Experience building stream-processing applications (Spark streaming, Apache-Flink, Kafka, etc.) ● Maintaining and developing CI/CD pipelines based on Gitlab. ● You have been involved in assembling large, complex structured and unstructured datasets that meet functional/non-functional business requirements. ● Experience of working with cloud data platform and services. ● Conduct code reviews, maintain code quality, and ensure best practices are followed. ● Debug and upgrade existing systems. ● Nice to have some knowledge in Devops.
More at Logic Planet