Source description
About the role
Job Title: Data Engineer Level: Senior Consultant Experience: 3-6 Years Location: Mumbai and Ahmedabad
Role Summary
The Data Engineer role focuses on the design, development, and maintenance of scalable, reliable, and high-performance data platforms. The role requires strong hands-on expertise in Python-based microservices, distributed data processing frameworks, and cloud-based big data ecosystems, with a primary emphasis on building robust applications to process ETL pipelines and production-grade data systems.
Key Responsibilities Design, develop, and maintain scalable Python-based microservices for data ingestion and processing Build and optimize microservices to support end-to-end ETL pipelines for large-scale batch and streaming data Develop and manage distributed data processing workfl ows using Apache Spark and PySpark Create, schedule, and monitor workfl ows using Apache Airfl ow Design and deploy data pipelines on AWS EMR or similar big data platforms using Python Implement real-time data ingestion and streaming pipelines Ensure data quality, reliability, performance, and scalability of data systems Apply best practices for code quality, testing, CI/CD, and production monitoring Troubleshoot production issues and perform root-cause analysis
Required Skills & Experience 6–8 years of experience in Data Engineering or backend data platform roles Strong profi ciency in Python with experience building production-grade microservices Hands-on expertise with Apache Spark and PySpark Experience working with AWS EMR or equivalent cloud-based big data services. (Must have knowledge on AWS eco system) Strong experience with Apache Airfl ow for workfl ow orchestration Deep understanding of ETL pipeline design and data modeling Hands-on experience with Kafka for inter-domain communication and real-time or streaming data pipelines Strong experience with PostgreSQL or similar RDBMS Experience working in Agile/Scrum environments
Good to Have Experience with Scala, Docker and Kubernetes Exposure to cloud platforms such as AWS, Azure, or GCP Experience with data lake architectures and schema evolution Familiarity with monitoring and logging tools
Desired Competencies Strong ownership and accountability Excellent problem-solving and analytical skills Good communication and collaboration skills Ability to work independently in a fast-paced environment
More at EY