Padmi

Data Engineer

BangalorePosted 3 months ago
Software engineeringMid-levelFull Time; Regular
Apply at Snapmint

Opens the source posting on shine.com

Source description

About the role

View original

You will be responsible for designing, building, and managing real-time data pipelines using tools such as Apache Kafka, Apache Flink, and Apache Spark Streaming. Your role will involve optimizing data pipelines for performance, scalability, and fault-tolerance, as well as performing real-time transformations, aggregations, and joins on streaming data. Collaboration with data scientists to onboard new features and ensuring they are discoverable, documented, and versioned will be a key aspect of your responsibilities. Additionally, you will optimize feature retrieval latency for real-time inference use cases and maintain strong data governance including lineage, auditing, schema evolution, and quality checks using tools like dbt and Open Lineage. Qualifications required for this role include a Bachelor's degree in Engineering from a premier institute (IIT/NIT/BIT) and 3-5 years of experience in an Indian startup/tech company. Strong programming skills in Python, Java, or Scala, along with proficiency in SQL, are essential. You should have a solid understanding of data modeling, data warehousing concepts, and the differences between OLTP and OLAP workloads. Experience with ingesting and processing various data formats, including semi-structured, unstructured, and document-based data from sources like NoSQL databases, APIs, and event tracking platforms is necessary. Hands-on experience with Change Data Capture tools, designing scalable data lakes, and working with real-time streaming technologies are also required. Proficiency with data pipeline orchestration tools and exposure to event-driven microservices architecture will be beneficial. Strong written and verbal communication skills are essential for this role. It is good to have familiarity with cloud data warehouse systems like BigQuery or Snowflake, experience with real-time analytical databases like ClickHouse, and familiarity with designing, building, and maintaining feature store infrastructure to support machine learning use cases. This position is based in Bangalore, and the working days are 5 days a week. You will be responsible for designing, building, and managing real-time data pipelines using tools such as Apache Kafka, Apache Flink, and Apache Spark Streaming. Your role will involve optimizing data pipelines for performance, scalability, and fault-tolerance, as well as performing real-time transformations, aggregations, and joins on streaming data. Collaboration with data scientists to onboard new features and ensuring they are discoverable, documented, and versioned will be a key aspect of your responsibilities. Additionally, you will optimize feature retrieval latency for real-time inference use cases and maintain strong data governance including lineage, auditing, schema evolution, and quality checks using tools like dbt and Open Lineage. Qualifications required for this role include a Bachelor's degree in Engineering from a premier institute (IIT/NIT/BIT) and 3-5 years of experience in an Indian startup/tech company. Strong programming skills in Python, Java, or Scala, along with proficiency in SQL, are essential. You should have a solid understanding of data modeling, data warehousing concepts, and the differences between OLTP and OLAP workloads. Experience with ingesting and processing various data formats, including semi-structured, unstructured, and document-based data from sources like NoSQL databases, APIs, and event tracking platforms is necessary. Hands-on experience with Change Data Capture tools, designing scalable data lakes, and working with real-time streaming technologies are also required. Proficiency with data pipeline orchestration tools and exposure to event-driven microservices architecture will be beneficial. Strong written and verbal communication skills are essential for this role. It is good to have familiarity with cloud data warehouse systems like BigQuery or Snowflake, experience with real-time analytical databases like ClickHouse, and familiarity with designing, building, and maintaining feature store infrastructure to support machine learning use cases. This position is based in Bangalore, and the working days are 5 days a week.

One address, no account. We’ll tell you when matching roles go live.

More at Snapmint

Related open roles

View all roles