Source description
About the role
As a Data Platform Engineer at our company, you will be responsible for taking end-to-end ownership of large-scale data pipelines and distributed systems that power our next-generation Graph Analytics platform. Role Overview: You will be in charge of designing and building end-to-end data pipelines, including both batch and streaming processes. Your responsibilities will include owning large-scale Apache Spark workloads, implementing data ingestion, transformation, and serving layers, managing schema evolution, data contracts, and ensuring pipeline reliability. Key Responsibilities: - Design and build end-to-end data pipelines (batch + streaming) - Own large-scale Apache Spark workloads and distributed data processing - Implement data ingestion, transformation, and serving layers - Manage schema evolution, data contracts, and pipeline reliability - Work on systems handling high-volume graph datasets, optimizing for latency, throughput, and fault tolerance - Design scalable architectures using Kafka, Spark, Flink, or Beam - Deploy and operate systems on GCP or AWS (GKE, Dataproc, Cloud Run, etc.) - Build and maintain CI/CD pipelines for data and microservices - Implement data quality checks, monitoring, and alerting - Ensure data integrity across pipelines and services - Work with GraphQL, REST, or gRPC APIs for data access layers - Ensure seamless integration between data systems and application layers Qualifications Required: - At least 4 years of experience in Data Engineering, Platform Engineering, or Distributed Systems - Strong hands-on experience with Apache Spark, distributed data processing, Cloud platforms (GCP or AWS), and streaming systems (Kafka, Flink, Beam) - Solid programming skills in Python, Java, Scala, or Node.js - Experience building and owning production data pipelines end-to-end - Understanding of microservices architecture, data modeling, and large-scale system design - Ability to debug and optimize systems in real production environments In addition to your core responsibilities, you will have the opportunity to work on real-scale projects, solve distributed systems, graph, and real-time problems, and influence architecture in an early-stage environment. You will operate close to production impact, not isolated dev work, and have a direct impact on the core platform architecture. Join us for a high-ownership, low bureaucracy environment where you will work on cutting-edge graph and AI-driven data systems, gain exposure to complex, real-world data problems, and experience fast growth with direct impact on core platform architecture. As a Data Platform Engineer at our company, you will be responsible for taking end-to-end ownership of large-scale data pipelines and distributed systems that power our next-generation Graph Analytics platform. Role Overview: You will be in charge of designing and building end-to-end data pipelines, including both batch and streaming processes. Your responsibilities will include owning large-scale Apache Spark workloads, implementing data ingestion, transformation, and serving layers, managing schema evolution, data contracts, and ensuring pipeline reliability. Key Responsibilities: - Design and build end-to-end data pipelines (batch + streaming) - Own large-scale Apache Spark workloads and distributed data processing - Implement data ingestion, transformation, and serving layers - Manage schema evolution, data contracts, and pipeline reliability - Work on systems handling high-volume graph datasets, optimizing for latency, throughput, and fault tolerance - Design scalable architectures using Kafka, Spark, Flink, or Beam - Deploy and operate systems on GCP or AWS (GKE, Dataproc, Cloud Run, etc.) - Build and maintain CI/CD pipelines for data and microservices - Implement data quality checks, monitoring, and alerting - Ensure data integrity across pipelines and services - Work with GraphQL, REST, or gRPC APIs for data access layers - Ensure seamless integration between data systems and application layers Qualifications Required: - At least 4 years of experience in Data Engineering, Platform Engineering, or Distributed Systems - Strong hands-on experience with Apache Spark, distributed data processing, Cloud platforms (GCP or AWS), and streaming systems (Kafka, Flink, Beam) - Solid programming skills in Python, Java, Scala, or Node.js - Experience building and owning production data pipelines end-to-end - Understanding of microservices architecture, data modeling, and large-scale system design - Ability to debug and optimize systems in real production environments In addition to your core responsibilities, you will have the opportunity to work on real-scale projects, solve distributed systems, graph, and real-time problems, and influence architecture in an early-stage environment. You will operate close to production impact, not isolated dev work, and have a direct impact on the core platform architecture. Join us for a high-
More at ContexQ