Source description
About the role
Lead Data Engineer Experience: 10+ Years Location: Bangalore (Hybrid 23 Days/Week) Shift: 12:00 PM – 10:00 PM IST Job Description We are looking for an experienced Lead Data Engineer to design, develop, and deliver scalable, high-performance data platforms using modern cloud technologies. The ideal candidate should have strong hands-on expertise in Databricks, Apache Spark, PySpark, Python, SQL, ETL/ELT, and cloud platforms (AWS/Azure/GCP) , along with experience in building enterprise data lakes and lakehouse architectures. This role involves working with cross-functional teams to build robust data pipelines, optimize large-scale data processing, and enable advanced analytics and GenAI use cases. Key Responsibilities Design and develop scalable batch and streaming data pipelines. Build and maintain ETL/ELT workflows using Databricks and Apache Spark. Develop and optimize Data Lake/Lakehouse architectures using Delta Lake and Medallion (Bronze/Silver/Gold) architecture. Design data models for analytics, reporting, machine learning, and AI workloads. Build high-quality datasets to support GenAI/LLM applications. Implement workflow orchestration using Apache Airflow or Databricks Workflows. Optimize Spark jobs for performance, scalability, and cost efficiency. Lead cloud data migration initiatives from legacy platforms to modern cloud-native architectures. Implement CI/CD, monitoring, logging, and data quality frameworks. Collaborate with architects, data scientists, and business stakeholders to deliver scalable solutions. Mentor junior engineers and drive engineering best practices. Required Skills 10+ years of Data Engineering experience. Strong hands-on experience with Databricks Lakehouse Platform . Expertise in Apache Spark, PySpark, Spark SQL . Advanced proficiency in Python and SQL . Strong experience building ETL/ELT pipelines . Experience with Delta Lake, Delta Live Tables (DLT), Unity Catalog , and Medallion Architecture. Hands-on experience with AWS, Azure, or GCP . Strong knowledge of Data Lakes, Data Warehousing, and Data Modeling. Experience with Apache Airflow or Databricks Workflows . Strong understanding of distributed data processing and Spark performance tuning. Experience with Git, CI/CD, Agile, and DevOps practices. Excellent communication and stakeholder management skills. Preferred Skills Experience with LLMs, Generative AI, RAG, LangChain, LlamaIndex, Embeddings, and Vector Databases . Experience with Kafka or Spark Structured Streaming. Knowledge of Terraform , Jenkins, or Azure DevOps. Cloud or Databricks certifications are an added advantage. Notice Period: Immediate
More at eTeam
Related open roles
SAP Datasphere / SAP Business Data Cloud (BDC) Lead
Dallas–Fort Worth
Data Governance Architect/Analyst
Dallas–Fort Worth
GCP Cloud Infrastructure Engineer– Terraform & SaaS Platforms
Remote · United States
Senior Kafka Devops Administrator
United States
DevOps Engineer
Dallas–Fort Worth
Senior Data Engineer / Data Architect (AWS, Python, Telecom)
Los Angeles