Padmi

Big Data Engineer - PySpark/Hadoop

BangalorePosted 2 months ago
Software engineeringMid-levelFull Time; Regular
Apply at Virtuoso Staffing Solutions (P) Ltd

Opens the source posting on shine.com

Source description

About the role

View original

Description : Roles & Responsibility : - To build a Big Data Platform for APAC region - To develop data pipelines to load different kind of data (both structured and unstructured data) to HADOOP system - To migrate existing datasets from Data warehouse to Data Lake. - Build views in hive based on the business needs. - Establish connection between various source systems from and to hive. - Prepare and maintain the inventory, mappings, documentation, assets and data dictionary for the platform. - Conduct training / support end user and technical team on how to use hue and establish connection to data lake for automation. - To work closely with our BI Team and users to understand their needs and provide necessary data for their reporting and analysis. - Perform QandA, unit testing for all use cases by preparing proper test cases. - Provided L2 support for the workflows built by Data Management team - To build a Big Data Platform for APAC region - To develop data pipelines to load different kind of data (both structured and unstructured data) to HADOOP system - To migrate existing datasets from Data warehouse to Data Lake. - Build views in hive based on the business needs. - Establish connection between various source systems from and to hive. - Prepare and maintain the inventory, mappings, documentation, assets and data dictionary for the platform. - Conduct training / support end user and technical team on how to use hue and establish connection to data lake for automation. - To work closely with our BI Team and users to understand their needs and provide necessary data for their reporting and analysis. - Perform QandA, unit testing for all use cases by preparing proper test cases. - Provided L2 support for the workflows built by Data Management team Requirements : - Experience in ETL / Data Engineering / Data Warehouse - Solid experience in Python, PySpark and SQL or ETL tools - Experience in handling huge volume of data is a must - Solid experience in structured and unstructured data in traditional and Big Data environments Oracle / SQL server / MongoDB / HIVE / PostgreSQL. - Experience in Adapting governance principles and coding best practices is a must. - Experience in Setting up connecting to application interaction, Hue Client is an advantage. - Experience with Cloud deployment and CI/CD pipeline building is a plus - Experience with indexing engines like Indexima is a plus - Experience in Data Cataloging would be an advantage - Experience in Data Governance Principle, Ranger Policy is a plus. Description : Roles & Responsibility : - To build a Big Data Platform for APAC region - To develop data pipelines to load different kind of data (both structured and unstructured data) to HADOOP system - To migrate existing datasets from Data warehouse to Data Lake. - Build views in hive based on the business needs. - Establish connection between various source systems from and to hive. - Prepare and maintain the inventory, mappings, documentation, assets and data dictionary for the platform. - Conduct training / support end user and technical team on how to use hue and establish connection to data lake for automation. - To work closely with our BI Team and users to understand their needs and provide necessary data for their reporting and analysis. - Perform QandA, unit testing for all use cases by preparing proper test cases. - Provided L2 support for the workflows built by Data Management team - To build a Big Data Platform for APAC region - To develop data pipelines to load different kind of data (both structured and unstructured data) to HADOOP system - To migrate existing datasets from Data warehouse to Data Lake. - Build views in hive based on the business needs. - Establish connection between various source systems from and to hive. - Prepare and maintain the inventory, mappings, documentation, assets and data dictionary for the platform. - Conduct training / support end user and technical team on how to use hue and establish connection to data lake for automation. - To work closely with our BI Team and users to understand their needs and provide necessary data for their reporting and analysis. - Perform QandA, unit testing for all use cases by preparing proper test cases. - Provided L2 support for the workflows built by Data Management team Requirements : - Experience in ETL / Data Engineering / Data Warehouse - Solid experience in Python, PySpark and SQL or ETL tools - Experience in handling huge volume of data is a must - Solid experience in structured and unstructured data in traditional and Big Data environments Oracle / SQL server / MongoDB / HIVE / PostgreSQL. - Experience in Adapting governance principles and coding best practices is a must. - Experience in Setting up connecting to application interaction, Hue Client is an adv

One address, no account. We’ll tell you when matching roles go live.

More at Virtuoso Staffing Solutions (P) Ltd

Related open roles

View all roles
Big Data Engineer - PySpark/Hadoop at Virtuoso Staffing Solutions (P) Ltd · Padmi