Source description
About the role
Key Responsibilities: Analyze and modernize legacy ETL processes and translate them into scalable ETL/ELT solutions . Design, develop, and optimize PySpark and Scala-based Spark jobs . Interpret and support Sqoop jobs , including JDBC configurations, incremental load logic, and syntax. Develop and manage AWS Glue jobs , crawlers, and workflows. Orchestrate Glue jobs using triggers, workflows, and scheduling mechanisms . Implement complex data transformation logic including joins, filters, aggregations, cleansing, and formatting. Apply Protegrity integration for data masking and tokenization of sensitive data during ETL processing. Optimize Glue job performance using partitioning strategies, parallel execution, and resource tuning . Handle schema evolution and manage changes in source data structures. Manage and maintain Glue Data Catalog , including metadata, table definitions, and schema versioning. Perform unit testing and integration testing for data pipelines. Lead data migration initiatives , including S3 to S3 migrations using AWS DataSync . Conduct source system analysis , including schema, data types, dependencies, and relationships. Design and document system architecture for individual components and the overall data pipeline. Produce and maintain technical documentation and manage development lifecycle. Support deployed code in production environments and troubleshoot issues as needed.
More at Artifint Technologies