Padmi

Data Engineer / Architect

ChennaiPosted 3 months ago
Infrastructure And DatabasesSeniorFull Time; Regular
Apply at Bahwan CyberTek

Opens the source posting on shine.com

Source description

About the role

View original

As a candidate for this role, your primary responsibility will be to demonstrate strong proficiency in the Databricks platform, including Delta Lake, Spark SQL, PySpark, Unity Catalog, MLflow, and Databricks Workflows. You should have deep expertise in data modeling, encompassing dimensional, data vault, and medallion/lakehouse architectures. Your experience in building ETL/ELT pipelines using Databricks, Apache Spark, or similar data engineering tools will be crucial. Proficiency in SQL and Python for data transformation, pipeline orchestration, and automation is essential. Your understanding of data governance principles such as data cataloging, lineage, quality monitoring, access control, and metadata management is required. Additionally, familiarity with cloud data platforms like Azure Data Lake Storage, Azure Synapse, AWS S3/Glue, or similar is expected. You should also have knowledge of AI/ML data requirements, including feature engineering, RAG data preparation, embedding storage, and LLM training/fine-tuning data pipelines. Experience in integrating data from enterprise systems such as ServiceNow, Workday, Active Directory, CMDB, and Jira will be beneficial. Knowledge of data privacy and compliance standards (GDPR, LGPD) and security best practices for data platforms is necessary. You should be comfortable with CI/CD pipelines for data, including Databricks Asset Bundles, Terraform, GitHub Actions, and Azure DevOps. Strong skills in documentation, data storytelling, and cross-functional communication are essential for this role. Qualifications required: - Bachelors or Masters degree in Computer Science, Data Engineering, Information Systems, or related field. (Note: No additional details about the company were provided in the job description.) As a candidate for this role, your primary responsibility will be to demonstrate strong proficiency in the Databricks platform, including Delta Lake, Spark SQL, PySpark, Unity Catalog, MLflow, and Databricks Workflows. You should have deep expertise in data modeling, encompassing dimensional, data vault, and medallion/lakehouse architectures. Your experience in building ETL/ELT pipelines using Databricks, Apache Spark, or similar data engineering tools will be crucial. Proficiency in SQL and Python for data transformation, pipeline orchestration, and automation is essential. Your understanding of data governance principles such as data cataloging, lineage, quality monitoring, access control, and metadata management is required. Additionally, familiarity with cloud data platforms like Azure Data Lake Storage, Azure Synapse, AWS S3/Glue, or similar is expected. You should also have knowledge of AI/ML data requirements, including feature engineering, RAG data preparation, embedding storage, and LLM training/fine-tuning data pipelines. Experience in integrating data from enterprise systems such as ServiceNow, Workday, Active Directory, CMDB, and Jira will be beneficial. Knowledge of data privacy and compliance standards (GDPR, LGPD) and security best practices for data platforms is necessary. You should be comfortable with CI/CD pipelines for data, including Databricks Asset Bundles, Terraform, GitHub Actions, and Azure DevOps. Strong skills in documentation, data storytelling, and cross-functional communication are essential for this role. Qualifications required: - Bachelors or Masters degree in Computer Science, Data Engineering, Information Systems, or related field. (Note: No additional details about the company were provided in the job description.)

One address, no account. We’ll tell you when matching roles go live.

More at Bahwan CyberTek

Related open roles

View all roles