Source description
About the role
About the role: This role focuses on building and maintaining scalable data pipelines and platforms using AWS, Python/SQL, and big data tools. It involves managing large datasets, optimizing ETL/ELT workflows, ensuring strong security and governance, and supporting real-time and batch data integration. The position requires hands-on experience with AWS services, DevOps practices, and Agile collaboration, along with strong problem solving skills and the ability to translate business needs into reliable data solutions. Responsibilities: Data Pipeline Development Design and build ETL/ELT pipelines to ingest, transform, and load data. Use AWS services like Glue, Lambda and Step Functions. Schedule and monitor workflows using Step Functions or AWS Glue Workflows. Create CI/CD Pipelines based on CloudFormation, Terraform, or CDK. Data Storage and Management Manage structured and unstructured data using services like: Amazon S3 Amazon Glue Catalogue Amazon DynamoDB Amazon RDS or Aurora Amazon Redshift Ensure data partitioning, indexing, and lifecycle management. Data Integration Integrate data from multiple sources (APIs, on-prem databases, third-party tools). Use AWS Glue, Kinesis, or Kafka for real-time or batch data streaming. Performance Optimization Tune ETL jobs and query performance on Athena or Databricks. Optimize storage formats (e.g., Parquet, ORC) for cost and speed. Security & Compliance Implement data encryption, IAM roles, and access policies. Ensure compliance with data governance and privacy policies (GDPR). Use AWS Lake Formation and IAM for access control and auditing. Monitoring & Maintenance Monitor pipeline health and performance using CloudWatch, CloudTrail, and custom dashboards. Set up alerts and logging for failures and anomalies. Working with Data Scientists, Analysts, and BI teams to understand data needs. Participate in Agile processes, sprint planning, and retrospectives. Requirements: Typically, 3+ years of professional data engineering or software engineering experience with a strong data focus. Proven delivery of complex, production-grade data pipelines and platforms across multiple components and domains. Experience translating ambiguous business and analytical requirements into scalable, governed data solutions. Experience working in Agile, cross-functional teams with product, analytics, and data science partners. Bachelor s Degree (Engineering/Computer Science preferred but not required); or equivalent experience required. AWS Services: S3, Glue, Redshift, Athena, EMR, Lambda, Cloud Formation, Kinesis, DynamoDB, SQS, SNS, Lake Formation Languages: Python, SQL, Scala Big Data Tools: Apache Spark, DataBricks, Hadoop, Kafka DevOps: Git, CI/CD, CloudFormation or Terraform Data Security: Strong experience ETL Pipelines: Strong experience AWS infrastructure: Familiarity with AWS VPC, EC2 Instances, Network policies, and Cloud Watch Experience operating very large data warehouses or data lakes. Strong Team Player. Security and governance-minded approach aligned to Secure SDLC and data protection practices. Ability to interface competently with other technical personnel or team members to finalize requirements. Ability to write and review portions of detailed specifications for the development of complex system components. Understanding of automation/other tools to increase efficiency. Strong problem solving and research skills. Ability to troubleshoot and resolve process inefficiencies (i.e. bugs) in the repository. Strong attention to detail. Strong oral and written communications skills.
More at RELX