Source description
About the role
We are looking for an experienced Senior Data Engineer to design, build, and optimize scalable, high-performance data platforms using AWS cloud services and Python. The ideal candidate will play a key role in architecting end-to-end data pipelines, driving automation, ensuring data quality, and enabling analytics and AI workloads across the organization. This role requires deep technical expertise in AWS data services, modern data architecture, and a passion for delivering reliable, high-quality data solutions at scale. Key Responsibilities Architect and implement scalable, fault-tolerant data pipelines using AWS Glue, Lambda, EMR, Step Functions, and Redshift Build and optimize data lakes and data warehouses on Amazon S3, Redshift, and Athena Develop Python-based ETL/ELT frameworks and reusable data transformation modules Integrate multiple data sources (RDBMS, APIs, Kafka/Kinesis, SaaS systems) into unified data models Lead efforts in data modeling, schema design, and partitioning strategies for performance and cost optimization Drive data quality, observability, and lineage using AWS Data Catalog, Glue Data Quality, or third-party tools Define and enforce data governance, security, and compliance best practices (IAM policies, encryption, access control) Collaborate with cross-functional teams (Data Science, Analytics, Product, DevOps) to support analytical and ML workloads Implement CI/CD pipelines for data workflows using AWS CodePipeline, GitHub Actions, or Cloud Build Provide technical leadership, code reviews, and mentoring to junior engineers Monitor data infrastructure performance, troubleshoot issues, and lead capacity planning Required Skills & Qualifications Bachelors or Masters degree in Computer Science, Information Systems, or related field 4–9 years of hands-on experience in data engineering or data platform development Expert-level proficiency in Python (pandas, PySpark, boto3, SQLAlchemy) Advanced experience with AWS Data Services, including: AWS Glue, Lambda, EMR, Step Functions, DynamoDB, EDW Redshift, Athena, S3, Kinesis, Amazon Quicksight. IAM, CloudWatch, CloudFormation / Terraform (for infrastructure automation) Strong experience in SQL, data modeling, and performance tuning Proven ability to design and deploy data lakes, data warehouses, and streaming solutions Solid understanding of ETL best practices, partitioning, error handling, and data validation Hands-on experience in version control (Git) and CI/CD for data pipelines Knowledge of containerization (Docker/Kubernetes) and DevOps concepts Excellent analytical, debugging, and communication skills
More at People Prime Worldwide