Source description
About the role
Job Summary We are looking for a skilled Data Engineer with 35 years of experience in building scalable data pipelines and integration solutions. The ideal candidate should have strong hands-on expertise in Databricks (must-have), Python, and API-based data integration. Key Responsibilities Design, develop, and maintain scalable ETL/ELT pipelines using Databricks and Python. Build and optimize data workflows using Apache Spark within the Databricks environment. Integrate data from internal systems and third-party platforms through REST APIs and other integration mechanisms. Develop reusable data ingestion frameworks and automation scripts. Perform data transformation, cleansing, validation, and enrichment for analytics and reporting. Work with structured and semi-structured data sources such as JSON, CSV, APIs, databases, and cloud storage. Optimize Spark jobs for performance, scalability, and cost efficiency. Collaborate with Data Analysts, Data Scientists, and application teams to deliver reliable datasets. Implement monitoring, logging, and error-handling mechanisms for production pipelines. Participate in code reviews, testing, and deployment activities following engineering best practices. Required Skills & Qualifications Bachelors degree in Computer Science, Information Technology, Engineering, or a related field. 35 years of experience in Data Engineering or Big Data development. Strong hands-on experience with Databricks (mandatory). Proficiency in Python for data processing and automation. Experience with REST APIs, API authentication, and data integration. Strong understanding of Apache Spark and distributed data processing. Experience writing complex SQL queries and working with relational databases. Knowledge of data lake/lakehouse concepts and Delta Lake. Experience with cloud platforms such as Azure, AWS, or GCP. Familiarity with Git, CI/CD pipelines, and Agile development practices. Preferred Skills Experience with Azure Databricks and Azure Data Factory. Knowledge of orchestration tools such as Airflow or Databricks Workflows. Exposure to streaming technologies such as Kafka. Understanding of data governance, security, and access control. Key Technical Stack Databricks (Must Have) Python Apache Spark / PySpark REST APIs SQL Delta Lake Azure / AWS / GCP Git Airflow / Databricks Workflows Compensation: 1,000,000.00 - 1,400,000.00 per year Benefits: Work from home Application Question(s): What is your Current CTC What is your Expected CTC How soon can you join if selected Work Location: Remote Job Summary We are looking for a skilled Data Engineer with 35 years of experience in building scalable data pipelines and integration solutions. The ideal candidate should have strong hands-on expertise in Databricks (must-have), Python, and API-based data integration. Key Responsibilities Design, develop, and maintain scalable ETL/ELT pipelines using Databricks and Python. Build and optimize data workflows using Apache Spark within the Databricks environment. Integrate data from internal systems and third-party platforms through REST APIs and other integration mechanisms. Develop reusable data ingestion frameworks and automation scripts. Perform data transformation, cleansing, validation, and enrichment for analytics and reporting. Work with structured and semi-structured data sources such as JSON, CSV, APIs, databases, and cloud storage. Optimize Spark jobs for performance, scalability, and cost efficiency. Collaborate with Data Analysts, Data Scientists, and application teams to deliver reliable datasets. Implement monitoring, logging, and error-handling mechanisms for production pipelines. Participate in code reviews, testing, and deployment activities following engineering best practices. Required Skills & Qualifications Bachelors degree in Computer Science, Information Technology, Engineering, or a related field. 35 years of experience in Data Engineering or Big Data development. Strong hands-on experience with Databricks (mandatory). Proficiency in Python for data processing and automation. Experience with REST APIs, API authentication, and data integration. Strong understanding of Apache Spark and distributed data processing. Experience writing complex SQL queries and working with relational databases. Knowledge of data lake/lakehouse concepts and Delta Lake. Experience with cloud platforms such as Azure, AWS, or GCP. Familiarity with Git, CI/CD pipelines, and Agile development practices. Preferred Skills Experience with Azure Databricks and Azure Data Factory. Knowledge of orchestration tools such as Airflow or Databricks Workflows. Exposure to streaming technologies such as Kafka. Understanding of data governance, security, and access control. Key Technical Stack Databricks (Must Have) Python Apache Spark / PySpark REST APIs SQL Delta Lake Azure / AWS / GCP Git Airfl
More at Acronotics