Padmi

Principal Data Engineer (Real World Data)

BangalorePosted 2 months ago
Infrastructure And DatabasesStaff+Full Time; Regular
Apply at Compile

Opens the source posting on shine.com

Source description

About the role

View original

You will be responsible for designing, building, and maintaining data pipelines that handle Real-world data at Compile. You will be handling both inbound and outbound data deliveries at Compile for datasets including Claims, Remittances, EHR, SDOH, etc. You will Work on building and maintaining data pipelines (specifically RWD).Build, enhance and maintain existing pipelines in pyspark, python and help build analytical insights and datasets.Scheduling and maintaining pipeline jobs for RWD.Develop, test, and implement data solutions based on the design.Design and implement quality checks on existing and new data pipelines.Ensure adherence to security and compliance that is required for the products.Maintain relationships with various data vendors and track changes and issues across vendors and deliveries.You have Hands-on experience with ETL process (min of 5 years).Excellent communication skills and ability to work with multiple vendors.High proficiency with Spark, SQL.Proficiency in Data modeling, validation, quality check, and data engineering concepts.Experience in working with big-data processing technologies using - databricks, dbt, S3, Delta lake, Deequ, Griffin, Snowflake, BigQuery.Familiarity with version control technologies, and CI/CD systems.Understanding of scheduling tools like Airflow/Prefect.Min of 3 years of experience managing data warehouses.Familiarity with healthcare datasets is a plus.Compile embraces diversity and equal opportunity in a serious way. We are committed to building a team of people from many backgrounds, perspectives, and skills. We know the more inclusive we are, the better our work will be. You will be responsible for designing, building, and maintaining data pipelines that handle Real-world data at Compile. You will be handling both inbound and outbound data deliveries at Compile for datasets including Claims, Remittances, EHR, SDOH, etc. You will Work on building and maintaining data pipelines (specifically RWD).Build, enhance and maintain existing pipelines in pyspark, python and help build analytical insights and datasets.Scheduling and maintaining pipeline jobs for RWD.Develop, test, and implement data solutions based on the design.Design and implement quality checks on existing and new data pipelines.Ensure adherence to security and compliance that is required for the products.Maintain relationships with various data vendors and track changes and issues across vendors and deliveries.You have Hands-on experience with ETL process (min of 5 years).Excellent communication skills and ability to work with multiple vendors.High proficiency with Spark, SQL.Proficiency in Data modeling, validation, quality check, and data engineering concepts.Experience in working with big-data processing technologies using - databricks, dbt, S3, Delta lake, Deequ, Griffin, Snowflake, BigQuery.Familiarity with version control technologies, and CI/CD systems.Understanding of scheduling tools like Airflow/Prefect.Min of 3 years of experience managing data warehouses.Familiarity with healthcare datasets is a plus.Compile embraces diversity and equal opportunity in a serious way. We are committed to building a team of people from many backgrounds, perspectives, and skills. We know the more inclusive we are, the better our work will be.

One address, no account. We’ll tell you when matching roles go live.

Principal Data Engineer (Real World Data) at Compile · Padmi