Source description
About the role
As an ideal candidate for the role, you will be responsible for the following key responsibilities: - Designing, developing, and maintaining scalable data pipelines using PySpark - Building and managing batch and real-time data processing systems - Developing and integrating Kafka-based streaming solutions - Optimizing Spark jobs for performance, cost, and scalability - Working with cloud-native services to deploy and manage data solutions - Ensuring data quality, reliability, and security across platforms - Collaborating with data scientists, analysts, and application teams - Participating in code reviews, design discussions, and production support In order to excel in this role, you must possess the following must-have skills: - Strong hands-on experience with PySpark / Apache Spark - Solid understanding of distributed data processing concepts - Experience with Apache Kafka (producers, consumers, topics, partitions) - Hands-on experience with any one cloud platform: AWS (S3, EMR, Glue, EC2, IAM) or Azure (ADLS, Synapse, Databricks) or GCP (GCS, Dataproc, BigQuery) - Proficiency in Python - Strong experience with SQL and data modeling - Experience working with large-scale datasets - Familiarity with Linux/Unix environments - Understanding of ETL/ELT frameworks - Experience with CI/CD pipelines for data applications Additionally, possessing the following good-to-have skills would be advantageous: - Experience with Spark Structured Streaming - Knowledge of Kafka Connect and Kafka Streams - Exposure to Databricks - Experience with NoSQL databases (Cassandra, MongoDB, HBase) - Familiarity with workflow orchestration tools (Airflow, Oozie) - Knowledge of containerization (Docker, Kubernetes) - Experience with data lake architectures - Understanding of security, governance, and compliance in cloud environments - Exposure to Scala or Java is a plus - Prior experience in Agile/Scrum environments Candidates ready to join immediately can share their details via email for quick processing at nitin.patil@ust.com. Act fast for immediate attention! As an ideal candidate for the role, you will be responsible for the following key responsibilities: - Designing, developing, and maintaining scalable data pipelines using PySpark - Building and managing batch and real-time data processing systems - Developing and integrating Kafka-based streaming solutions - Optimizing Spark jobs for performance, cost, and scalability - Working with cloud-native services to deploy and manage data solutions - Ensuring data quality, reliability, and security across platforms - Collaborating with data scientists, analysts, and application teams - Participating in code reviews, design discussions, and production support In order to excel in this role, you must possess the following must-have skills: - Strong hands-on experience with PySpark / Apache Spark - Solid understanding of distributed data processing concepts - Experience with Apache Kafka (producers, consumers, topics, partitions) - Hands-on experience with any one cloud platform: AWS (S3, EMR, Glue, EC2, IAM) or Azure (ADLS, Synapse, Databricks) or GCP (GCS, Dataproc, BigQuery) - Proficiency in Python - Strong experience with SQL and data modeling - Experience working with large-scale datasets - Familiarity with Linux/Unix environments - Understanding of ETL/ELT frameworks - Experience with CI/CD pipelines for data applications Additionally, possessing the following good-to-have skills would be advantageous: - Experience with Spark Structured Streaming - Knowledge of Kafka Connect and Kafka Streams - Exposure to Databricks - Experience with NoSQL databases (Cassandra, MongoDB, HBase) - Familiarity with workflow orchestration tools (Airflow, Oozie) - Knowledge of containerization (Docker, Kubernetes) - Experience with data lake architectures - Understanding of security, governance, and compliance in cloud environments - Exposure to Scala or Java is a plus - Prior experience in Agile/Scrum environments Candidates ready to join immediately can share their details via email for quick processing at nitin.patil@ust.com. Act fast for immediate attention!
More at UST
Related open roles
Lead Ii Data Engineering Python, Aws, Sql Kerala
India
Lead Ii Devops Engineering Chennai
Chennai
Nodejs Developer(Javascript, Nodejs with Express JS, SQL, ReactJS)
Hyderabad
Lead I / Lead II - Fullstack Dev (React, Node JS, Typescript, JavaScript)
India
Developer II - Software Engineering
Chennai
Lead II - Software Engineering
Chennai