Source description
About the role
Role Overview: You will be responsible for designing, building, and maintaining Databricks data pipelines (ETL/ELT) for ingestion, transformation, and orchestration using Spark/Delta Lake/Databricks Workflows. Additionally, you must have practical experience in Databricks and MLflow, including model development, experiment tracking, model management, and deployment in a production environment. Your role will involve operationalizing machine learning models by building inference pipelines that invoke models authored by data scientists (batch or real-time) to ensure consistency between training and inference environments. Furthermore, you will collaborate closely with data scientists to productionize models, manage model deployment lifecycles, and optimize inference performance and cost. Key Responsibilities: - Design, build, and maintain Databricks data pipelines (ETL/ELT) using Spark/Delta Lake/Databricks Workflows. - Have practical experience in Databricks and MLflow for model development, experiment tracking, model management, and deployment in production. - Operationalize machine learning models by building inference pipelines to ensure consistency between training and inference environments. - Ensure data reliability, quality, and observability through validation, monitoring, alerting, and automated recovery mechanisms. - Collaborate with data scientists to productionize models, manage model deployment lifecycles, and optimize inference performance and cost. - Implement best-practice DevOps/MLOps processes such as CI/CD for pipelines, model versioning, environment promotion, and infrastructure-as-code. - Optimize performance and cost across compute clusters, jobs, and storage layers. - Manage the enterprise data catalog, including schema design, table ownership, lineage, governance, and documentation using Unity Catalog. - Experience with Databricks infrastructure, building BI dashboards, visualization, coding agents, and best practices. Qualifications Required: - Databricks platform experience. - Proficiency in Python for data processing and ETL pipelines. - Knowledge of Unity Catalog and AWS data services (S3, IAM, VPC, potentially Glue/Lambda). - Familiarity with data lake/lakehouse architecture patterns. - Experience with RESTful API design and development, authentication/authorization patterns, query optimization, performance tuning, PySpark optimization, ML/AI pipeline, and Databricks AI/BI would be advantageous. Role Overview: You will be responsible for designing, building, and maintaining Databricks data pipelines (ETL/ELT) for ingestion, transformation, and orchestration using Spark/Delta Lake/Databricks Workflows. Additionally, you must have practical experience in Databricks and MLflow, including model development, experiment tracking, model management, and deployment in a production environment. Your role will involve operationalizing machine learning models by building inference pipelines that invoke models authored by data scientists (batch or real-time) to ensure consistency between training and inference environments. Furthermore, you will collaborate closely with data scientists to productionize models, manage model deployment lifecycles, and optimize inference performance and cost. Key Responsibilities: - Design, build, and maintain Databricks data pipelines (ETL/ELT) using Spark/Delta Lake/Databricks Workflows. - Have practical experience in Databricks and MLflow for model development, experiment tracking, model management, and deployment in production. - Operationalize machine learning models by building inference pipelines to ensure consistency between training and inference environments. - Ensure data reliability, quality, and observability through validation, monitoring, alerting, and automated recovery mechanisms. - Collaborate with data scientists to productionize models, manage model deployment lifecycles, and optimize inference performance and cost. - Implement best-practice DevOps/MLOps processes such as CI/CD for pipelines, model versioning, environment promotion, and infrastructure-as-code. - Optimize performance and cost across compute clusters, jobs, and storage layers. - Manage the enterprise data catalog, including schema design, table ownership, lineage, governance, and documentation using Unity Catalog. - Experience with Databricks infrastructure, building BI dashboards, visualization, coding agents, and best practices. Qualifications Required: - Databricks platform experience. - Proficiency in Python for data processing and ETL pipelines. - Knowledge of Unity Catalog and AWS data services (S3, IAM, VPC, potentially Glue/Lambda). - Familiarity with data lake/lakehouse architecture patterns. - Experience with RESTful API design and development, authentication/authorization patterns, query optimization, performance tuning, PySpark optimization, ML/AI pipeline, and Databricks AI/BI would be advantageous.
More at /codvo