Source description
About the role
OverviewLine of Service Advisory Industry/Sector Not Applicable Specialism Operations Management Level Senior Associate Job Description & Summary At PwC, our people in software and product innovation focus on developing cutting-edge software solutions and driving product innovation to meet the evolving needs of clients. These individuals combine technical experience with creative thinking to deliver innovative software products and solutions. In emerging technology at PwC, you will focus on exploring and implementing cutting-edge technologies to drive innovation and transformation for clients. You will work in areas such as artificial intelligence, blockchain, and the internet of things (IoT). Why PwC: At PwC, you will be part of a vibrant community of solvers that leads with trust and creates distinctive outcomes for our clients and communities. This purpose-led and values-driven work, powered by technology in an environment that drives innovation, will enable you to make a tangible impact in the real world. We reward your contributions, support your wellbeing, and offer inclusive benefits, flexibility programmes and mentorship that will help you thrive in work and life. Together, we grow, learn, care, collaborate, and create a future of infinite experiences for each other. Learn more about us. At PwC, we believe in providing equal employment opportunities, without any discrimination on the grounds of gender, ethnic background, age, disability, marital status, sexual orientation, pregnancy, gender identity or expression, religion or other beliefs, perceived differences and status protected by law. We strive to create an environment where each one of our people can bring their true selves and contribute to their personal growth and the firms growth. To enable this, we have zero tolerance for any discrimination and harassment based on the above considerations. ResponsibilitiesDesign, build, and maintain robust ETL/ELT pipelines on Azure using Apache Spark (PySpark/Scala) on Databricks/HDInsightOrchestrate complex data workflows with Azure Data Factory (pipelines, triggers, integration runtimes) and Databricks Jobs/WorkflowsDevelop and manage data lakes on Azure Blob Storage/ADLS Gen2, including naming, partitioning, lifecycle policies, and schema evolutionImplement curated layers (raw, staged, curated), leveraging columnar formats (Parquet/ORC/Avro) and table formats (Delta Lake); manage metastore/Unity CatalogOptimize Spark jobs and cluster configurations for performance and cost (autoscaling, spot VMs, Photon, adaptive query execution, caching, partition tuning)Operationalize jobs with monitoring, logging, and alerting via Azure Monitor, Log Analytics, and Databricks metrics; build runbooks and dashboardsImplement data quality, testing, and observability for pipelines (unit/integration tests, Great Expectations, SLAs, lineage)Collaborate with Analytics, Data Science, and Product to deliver modeled, trustworthy datasets for BI, ML, and applicationsEnforce security and governance best practices (AAD RBAC, Managed Identities, ACLs, Key Vault, Private Endpoints, VNet integration, encryption with CMK/SSE)Contribute to infrastructure as code and CI/CD (Azure DevOps or GitHub Actions) including Databricks objects deploymentParticipate in on-call rotations, incident response, and postmortems; drive continuous improvement and documentationMandatory skill sets6+ years as a Data Engineer (or similar) with a strong focus on Azure data servicesExpert-level experience with Apache Spark (PySpark and/or Scala) and distributed data processingHands-on experience with Databricks and Azure HDInsight for large-scale batch processingProficient in Azure Data Factory for orchestration (pipelines, data flows, triggers)Strong Python skills and solid SQL (window functions, performance tuning, optimization)Practical experience with Azure Blob Storage/ADLS Gen2, Azure Key Vault, Azure Monitor/Log AnalyticsUnderstanding of Hadoop ecosystem fundamentals (HDFS, YARN, Hive/Metastore)Strong grasp of data modeling, file formats (Parquet/ORC/Avro), partitioning, and performance best practicesExperience building production-grade pipelines with testing, monitoring, and alertingVersion control with Git and collaborative development practicesExcellent communication and cross-functional collaboration skillsUnderstanding of Hadoop ecosystem fundamentals (HDFS, YARN, Hive/Metastore)Strong grasp of data modeling, file formats (Parquet/ORC/Avro), partitioning, and performance best practicesExperience building production-grade pipelines with testing, monitoring, and alertingVersion control with Git and collaborative development practicesExcellent communication and ability to work cross-functionallyPreferred skill setsStream processing and event-driven architectures (Kafka, Azure Event Hubs, Azure Functions)Lakehouse technologies (Delta Lake, Unity Catalog) and query engines (Synapse Serverless, Databricks SQL)Governance and lineage tools OverviewLi
More at PwC South Africa
Related open roles
Senior Associate Technical Business Analyst Digital Integration Advisory Gurgaon
Delhi NCR
Associate Infrastructure Automation
Delhi NCR
IN-Associate SAP SD Advisory
Delhi NCR
Senior Associate - SAP ABAP - AI in Delivery Advisory
Bangalore
Senior Associate Azure Data Engineer
India
Associate GenAI and Agentic AI Engineer
Delhi NCR