Source description
About the role
Azure Data Engg Number of Openings 5 No of years experience 3- 5 yrs Detailed job description - Skill Set: Provided below Mandatory Skills(ONLY 2 or 3) PySpark & Azure Work Location Hyderabad, Trivandrum WFO/WFH/Hybrid WFO Hybrid WFO Joining time ( Notice period) Immediate Is there any working in shifts from standard Daylight (to avoid confusions post onboarding) YES/ NO No Detailed Job Description : Data Pipeline Development: Design, build, and optimize scalable and resilient ETL/ELT pipelines using Azure Databricks (PySpark, Scala, SQL) for batch and streaming data processing. Ingest data from diverse sources (e.g., relational databases, APIs, cloud storage, streaming platforms) into Azure Data Lake Storage (ADLS) and other Azure data services. Implement data transformations, aggregations, and quality checks within Databricks notebooks and jobs. Work with various data formats like Delta Lake, Parquet, Avro, JSON, and CSV. Azure Cloud & Data Services: Leverage a wide range of Azure data services, including Azure Data Factory (ADF) for orchestration, Azure Synapse Analytics, Azure SQL Database, Azure Event Hubs, Azure Stream Analytics, and Azure Cosmos DB. Design and manage data storage solutions in ADLS (Gen1/Gen2) and ensure optimal data partitioning and indexing. Monitor and optimize the performance and cost-efficiency of Azure data solutions. DevOps & CI/CD: Implement and maintain CI/CD pipelines using Jenkins to automate the build, test, and deployment of data engineering solutions to Azure Databricks workspaces. Manage source code effectively using Git (e.g., GitHub, Azure DevOps Repos, GitLab), including branching strategies, pull requests, and code reviews. Ensure code quality, testability, and adherence to best practices through automated testing and linting. Collaborate with DevOps teams to streamline deployment processes and infrastructure as code (IaC) where applicable. Data Governance & Security: Implement data security, privacy, and compliance best practices within Azure and Databricks environments. Work with Azure Active Directory (AAD) for access management and secure data access. Ensure data quality, integrity, and reliability across all data pipelines. Collaboration & Communication: Collaborate closely with data scientists, data analysts, business stakeholders, and other engineering teams to understand data requirements and deliver robust solutions. Provide technical leadership, mentorship, and support to junior data engineers. Document data architectures, pipelines, and technical specifications clearly and concisely. Participate in agile development methodologies (Scrum/Kanban). Troubleshooting & Optimization: Monitor production data pipelines for performance, reliability, and data quality issues. Perform root cause analysis and implement solutions for production incidents. Continuously identify opportunities for performance tuning and optimization of data processes. Required Qualifications: Bachelors or Masters degree in Computer Science, Engineering, Information Technology, or a related quantitative field. 3 - 5 years of experience in data engineering, with a strong focus on building and maintaining large-scale data solutions. Proficiency in Azure: Extensive hands-on experience with key Azure data services, including Azure Databricks, Azure Data Factory, Azure Data Lake Storage, and Azure SQL Database. Expertise in Databricks: Strong experience with Apache Spark (PySpark, Scala) for data processing and transformation within Databricks. Familiarity with Delta Lake is essential. Version Control: Expert-level proficiency with Git for source code management. CI/CD: Hands-on experience designing, implementing, and maintaining CI/CD pipelines using Jenkins. Programming Languages: Strong programming skills in Python is a must. Experience with Scala or Java is a plus. SQL: Advanced SQL scripting and query optimization skills. Data Warehousing: Solid understanding of data warehousing concepts, data modeling (dimensional modeling, Kimball/Inmon), and ETL/ELT principles. Problem-Solving: Excellent analytical and problem-solving skills with a keen eye for detail. Communication: Strong verbal and written communication skills, with the ability to articulate complex technical concepts to both technical and non-technical audiences. Preferred Qualifications (Nice to Have): Azure Data Engineer Associate (DP-203) or Databricks Certified Data Engineer certifications. Experience with other big data technologies like Kafka, Hadoop, or Snowflake. Familiarity with other CI/CD tools (e.g., Azure DevOps Pipelines, GitHub Actions). Experience with Infrastructure as Code (e.g., Terraform). Knowledge of monitoring tools like Azure Monitor, Splunk, or similar. Experience working in an Agile/Scrum environment. Location- Hyderabad/Trivandrum Yrs of Exp-3+Yrs
More at Diverse Lynx