Source description
About the role
Role & responsibilities - Design, build, and maintain Databricks workspaces, clusters, and compute pools across development, testing, and production environments. - Configure and manage Unity Catalog for data governance, fine-grained access control, permissions, metadata management, and data lineage. - Optimize Databricks cluster configurations, including instance types, auto -scaling, spot/preemptible nodes, and compute pools to improve performance and reduce costs. - Implement workspace best practices, including folder structures, access controls, secret management using Databricks Secrets, Azure Key Vault, or AWS Secrets Manager. - Create, schedule, and manage Databricks Jobs, Workflows, and multi-task job orchestration with dependency management. - Design and implement Delta Lake tables using partitioning, Z-Ordering, OPTIMIZE, VACUUM, and file compaction techniques. - Build and maintain Medallion Architecture (Bronze, Silver, and Gold layers) for scalable and governed data lakehouse solutions. - Develop Delta Live Tables (DLT) pipelines with built-in data quality expectations for reliable ETL/ELT processing. - Manage schema evolution, table versioning, Time Travel, and Change Data Feed (CDF) to support incremental data processing. - Design and implement lakehouse architectures integrating Delta Lake with cloud storage and external systems such as Azure Data Lake Storage (ADLS), Kafka, Event Hubs, and Kinesis. - Develop scalable batch and real-time data pipelines using PySpark, Spark SQL, Structured Streaming, and Delta Lake. - Build streaming ingestion pipelines from Kafka, Azure Event Hubs, and other streaming platforms into Delta tables. - Optimize PySpark applications using broadcast joins, Adaptive Query Execution (AQE), dynamic partition pruning, caching, and Photon Engine. - Develop reusable transformation frameworks, utility libraries, and pipeline templates to improve engineering productivity and standardization. - Implement robust error handling, retry mechanisms, logging, monitoring, and dead -letter queue (DLQ) patterns for production-grade pipelines. - Set up and manage MLflow experiment tracking, model registry, and model lifecycle management. - Support machine learning workloads by enabling scalable model training, inference, and GPU-based compute environments. - Develop feature engineering pipelines using Databricks Feature Store to create reusable and versioned machine learning features. - Enable Generative AI solutions, including Retrieval-Augmented Generation (RAG), vector search, LLM fine-tuning, and Mosaic AI capabilities. - Implement MLOps best practices, including model versioning, model deployment, A/B testing, and Databricks Model Serving. - Integrate Databricks with Azure Data Lake Storage (ADLS) and other cloud -native services. - Develop and maintain CI/CD pipelines using Azure DevOps, GitHub Actions, or GitLab CI for Databricks notebooks, jobs, and workflows. - Automate Databricks infrastructure deployment using Databricks Asset Bundles (DABs), Terraform, and Infrastructure-as-Code (IaC) practices. - Build and manage data ingestion frameworks using Auto Loader, COPY INTO, and third -party integration tools such as Fivetran, dbt, and Airbyte. - Monitor pipeline execution, cluster utilization, system performance, and cloud costs using Databricks system tables and cloud monitoring tools. - Implement row-level security, column-level masking, energetic views, and governance policies using Unity Catalog. - Enforce data quality through Delta Live Tables expectations and Great Expectations frameworks. - Perform query optimization, execution plan analysis, caching strategies, and performance tuning to improve workload efficiency. - Maintain enterprise data cataloging, metadata management, and end -to-end data lineage. - Prepare technical documentation, architecture diagrams, operational runbooks, and standard operating procedures for Databricks platform and data engineering solutions. .
More at WOW Softech