Padmi

Lead Data Engineer (Python, PySpark, AWS) 100% Remote

IndiaPosted 1 month ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at Apetan Consulting

Opens the source posting on shine.com

Source description

About the role

View original

Lead Data Engineer (Python, PySpark, AWS) 100% Remote Job Summary We are looking for an experienced Lead Data Engineer to design, develop, and optimize scalable data platforms and pipelines. The ideal candidate will have strong expertise in Python, PySpark, and AWS cloud services, along with experience leading technical initiatives and mentoring engineering teams. You will collaborate with cross-functional stakeholders to build reliable, high-performance data solutions that support analytics, reporting, and machine learning use cases. Key Responsibilities Design, build, and maintain scalable ETL/ELT data pipelines. Develop and optimize data processing solutions using Python and PySpark. Build and manage cloud-based data platforms on AWS. Design and implement data lakes and data warehouse solutions. Optimize data workflows for performance, scalability, and reliability. Ensure data quality, governance, security, and compliance best practices. Collaborate with data scientists, analysts, architects, and business stakeholders. Lead code reviews, establish engineering best practices, and mentor junior team members. Monitor production data pipelines and troubleshoot performance issues. Participate in solution architecture, effort estimation, and technical planning. Required Skills 710+ years of experience in Data Engineering. Strong programming experience in Python. Hands-on expertise with PySpark and Apache Spark. Experience with AWS services such as S3, Glue, EMR, Lambda, Redshift, Athena, IAM, and CloudWatch. Strong SQL skills and experience with relational and NoSQL databases. Experience building batch and large-scale data processing pipelines. Knowledge of data modeling, data warehousing, and ETL best practices. Familiarity with version control (Git) and CI/CD pipelines. Excellent problem-solving, communication, and leadership skills. Preferred Skills Experience with Apache Airflow or similar workflow orchestration tools. Knowledge of Delta Lake, Apache Iceberg, or Apache Hudi. Experience with Kafka or other streaming technologies. Exposure to Docker, Kubernetes, and Infrastructure as Code (Terraform/CloudFormation). Understanding of DevOps and DataOps practices. .

One address, no account. We’ll tell you when matching roles go live.

More at Apetan Consulting

Related open roles

View all roles