Source description
About the role
Role Overview: As the Lead, Big Data Analytics & Engineering at Mastercard, you will hold a senior technical leadership position focusing on data engineering, data unification, and large-scale analytics enablement across enterprise data assets. Your role will involve integrating diverse internal and external data sources to build a single, trusted, and scalable view of data. By partnering closely with various teams, you will design and deliver high-impact data solutions that generate measurable business value, particularly in areas such as Value Quantification, Cyber Intelligence, and Analytics led solutions. Key Responsibilities: - Data Engineering & Platform Leadership: - Lead the ingestion, transformation, aggregation, and processing of large-scale datasets to enable advanced analytics and downstream consumption. - Design, build, and maintain robust, scalable data pipelines across Hadoop and enterprise data platforms, ensuring high standards of data quality, reliability, performance, and availability. - Drive data unification initiatives by integrating multiple structured and semi-structured data sources into a cohesive, governed analytical foundation. - Advanced Analytics Enablement: - Manipulate and analyze high volume, high velocity, and high-dimensional datasets using modern big data frameworks. - Analyze large volumes of transactional and product data to produce insights and actionable recommendations that support business growth and value realization. - Apply metrics, measurement frameworks, and benchmarking techniques to evaluate solution effectiveness and drive continuous improvement. - Cross Functional Collaboration: - Partner with Product Managers, Data Science, Platform Strategy, and Technology teams to understand analytical and data requirements and translate them into scalable engineering solutions. - Act as a technical bridge between business, analytical, and engineering teams, articulating architecture decisions, trade-offs, and implementation approaches. - Ensure alignment across stakeholders to tie data solutions directly to business and customer outcomes. - Innovation & Value Creation: - Identify innovation opportunities and deliver proofs of concept, prototypes, and pilot solutions aligned with near-term and future business needs. - Integrate new and emerging data assets that enhance existing platforms, products, and services, strengthening overall value propositions. - Gather and synthesize feedback from clients, product, engineering, and sales teams to inform new solutions and product enhancements. - Technical Leadership & Mentorship: - Provide technical leadership, guidance, and mentorship to data engineers and analysts, setting standards for engineering quality, scalability, performance, and maintainability. - Promote best practices in data modeling, pipeline design, performance optimization, and data governance. - Influence engineering standards, architectural consistency, and long-term platform sustainability. Qualification Required: - Technical Skills & Experience: - Strong proficiency in Python, including Pandas, NumPy, PySpark, with hands-on experience using Impala. - Proven experience working on Hadoop-based platforms, performing large-scale data extraction, transformation, and processing. - Strong SQL skills and experience working with both relational and distributed data stores. - Experience with enterprise data platforms and business intelligence ecosystems. - Hands-on experience with ETL / ELT and data integration tools such as Apache Airflow, Apache NiFi, Azure Data Factory, Pentaho, or Talend. - Experience in data modeling, querying, data mining, and reporting over large volumes of granular data. - Exposure to machine learning concepts and analytical techniques used in advanced data solutions. - Experience with Graph Databases is a plus. - 8+ years of experience in data engineering, big data analytics, or enterprise data platforms, including 2+ years in a lead or technical leadership role. - Experience working with cloud-based data platforms (Azure, AWS, or GCP), including data lakes, distributed compute, and storage services. - Experience implementing CI/CD pipelines and DevOps practices for data engineering workflows. - GenAI / LLM Skills (Preferred): - Experience enabling GenAI/AI products through scalable, reliable data ingestion and transformation pipelines (batch and streaming). - Exposure to unstructured and semi-structured data processing (documents/logs/text) and building curated datasets for downstream consumption. - Strong understanding of data governance, privacy, and security requirements when using enterprise data with AI (PII handling, access control, auditability). - Familiarity with operationalizing AI data workflows (monitoring, data quality checks, reproducibility, and cost-aware scaling in cloud environments). - Analytical & Business Acumen: - Strong
More at Mastercard