Source description
About the role
As a skilled and detail-oriented Data Engineer with expertise in PySpark and Databricks, your role will involve designing, building, and optimizing scalable data pipelines. You should have hands-on experience in data warehousing, ETL processes, and modern data lake architectures, with the ability to translate business requirements into robust technical solutions. - Design, develop, and optimize large-scale data pipelines using PySpark and Databricks - Build and maintain ETL workflows for data extraction, transformation, and loading across multiple sources - Develop scalable data solutions for Data Warehousing and Data Lake environments - Translate business requirements into technical specifications and low-level ETL design documents - Ensure data quality, integrity, and performance optimization across pipelines - Collaborate with cross-functional teams including analysts, architects, and business stakeholders - Troubleshoot and resolve data-related issues in production environments Required Skills & Experience: - Strong experience in Data Warehousing, ETL, Data Integration, and Data Lake projects - Hands-on expertise in Databricks for developing and deploying ETL solutions - Proficiency in PySpark for large-scale data processing - Strong understanding of data management ecosystems and architecture - Experience with at least one database system: Redshift, Snowflake, SQL Server, Oracle, Teradata, or Azure SQL - Solid SQL skills and experience in performance tuning - Ability to independently design and deliver complex data solutions Good to Have: - Experience in Pharma / Life Sciences domain - Exposure to cloud platforms (AWS/Azure) Educational Qualification: - B.E./B.Tech / MCA / BCA / Computer Science or related field - Minimum 60% throughout academics As a skilled and detail-oriented Data Engineer with expertise in PySpark and Databricks, your role will involve designing, building, and optimizing scalable data pipelines. You should have hands-on experience in data warehousing, ETL processes, and modern data lake architectures, with the ability to translate business requirements into robust technical solutions. - Design, develop, and optimize large-scale data pipelines using PySpark and Databricks - Build and maintain ETL workflows for data extraction, transformation, and loading across multiple sources - Develop scalable data solutions for Data Warehousing and Data Lake environments - Translate business requirements into technical specifications and low-level ETL design documents - Ensure data quality, integrity, and performance optimization across pipelines - Collaborate with cross-functional teams including analysts, architects, and business stakeholders - Troubleshoot and resolve data-related issues in production environments Required Skills & Experience: - Strong experience in Data Warehousing, ETL, Data Integration, and Data Lake projects - Hands-on expertise in Databricks for developing and deploying ETL solutions - Proficiency in PySpark for large-scale data processing - Strong understanding of data management ecosystems and architecture - Experience with at least one database system: Redshift, Snowflake, SQL Server, Oracle, Teradata, or Azure SQL - Solid SQL skills and experience in performance tuning - Ability to independently design and deliver complex data solutions Good to Have: - Experience in Pharma / Life Sciences domain - Exposure to cloud platforms (AWS/Azure) Educational Qualification: - B.E./B.Tech / MCA / BCA / Computer Science or related field - Minimum 60% throughout academics
More at Axtria