Source description
About the role
We are looking for a Lead Data Engineer to join our Content Tech Big Data Engineering Team in India. This is an amazing opportunity to work on Real Wor l d Data using big data technologies. W e would love to speak with you if you have skills in Python, Spark and have experience on building big data platforms. About You - experience, education, skills, and accomplishments Bachelor s Degree or equivalent in computer science, software engineering, or a related field At least 5 y ears of r elevant e xperience . Good experience working with Python, Py S park , AWS, AWS Glue, EMR and Delta Lake. Good knowledge of ETL, including the ability to read and write efficient, robust code, follow or implement best practices and coding standards, design/implement common ETL strategies (CDC, SCD, etc.), and create reusable/maintainable jobs. Solid background in database systems (such as Postgres, Oracle, Snowflake /Databricks ) along with strong knowledge of PL/SQL and SQL. Experience in handling large volume of data and building data pipelines. Possess good knowledge of Agile/other SDLC methodologies. Exposure to a Data warehouse / BI project in a Healthcare Domain . Strong oral and written communication skills . It would be great if you also had . . . Familiarity with s would be added advantage. Experience in building big data platforms. Understanding on healthcare data. What will you be doing in this role As a member of Data Engineering Team, you ll Step into a key role on an expanding data engineering team to build our data platforms, data pipelines, and data transformation capabilities. Define and implement our data platform strategy on Cloud, have a meaningful impact on our customers, and working in our high energy, innovative, fast-paced Agile culture. Drive rapid prototyping and development with Product and Technical teams in building and scaling high-value medical data capabilities. Interface with other technology teams to extract, transform, and load data from a wide variety of data sources using Apache suite (airflow, spark), SQL, Python, ETL, and AWS big data technologies. Creation and support of batch and real-time data pipelines and ongoing data monitoring and validation built on AWS/Snowflake/Apache technologies for medical data from many different sources. Conduct functional and non-functional testing, writing test scenarios and test scripts. Evaluate existing applications to update and add new features to meet business requirements.
More at Clarivate