Padmi

Data Engineer For Python

MumbaiPosted 3 months ago
Software engineeringMid-levelFull Time; Regular
Apply at A2Tech Consultants

Opens the source posting on shine.com

Source description

About the role

View original

As a candidate for this position, you should have the following qualifications and experience: Role Overview: You should possess strong Python coding skills, along with a solid understanding of Object-Oriented Programming (OOP). Additionally, you must have practical experience in working on Big Data product architecture. Your familiarity with SQL-based databases such as MySQL, PostgreSQL, and NoSQL-based databases like Cassandra and Elasticsearch is crucial. Hands-on experience with frameworks like Spark RDD, DataFrame, and Dataset is necessary. Proficiency in developing ETL for data products is also expected from you. Moreover, you should have expertise in performance optimization, optimal resource utilization, parallelism, and tuning of Spark jobs. Your knowledge of file formats including CSV, JSON, XML, PARQUET, ORC, and AVRO is essential. It would be advantageous if you have worked with Analytical Databases like Druid, MongoDB, or Apache Hive. Experience in handling real-time data feeds and familiarity with Apache Kafka or similar tools is a plus. Key Responsibilities: - Strong Python coding skills and OOP skills - Work on Big Data product Architecture - Experience with SQL-based databases (e.g., MySQL, PostgreSQL) and NoSQL-based databases (e.g., Cassandra, Elasticsearch) - Hands-on experience with frameworks like Spark RDD, DataFrame, Dataset - Development of ETL for data product - Working knowledge on performance optimization, optimal resource utilization, parallelism, and tuning of Spark jobs - Familiarity with file formats: CSV, JSON, XML, PARQUET, ORC, AVRO - Good to have experience with Analytical Databases like Druid, MongoDB, Apache Hive - Ability to handle real-time data feeds (knowledge of Apache Kafka or similar tool is a bonus) Qualifications Required: - Strong Python coding skills - Experience with SQL-based and NoSQL-based databases - Hands-on experience with Spark and PySpark - Knowledge of parallel programming Please note that Scala is optional but would be beneficial for this role. As a candidate for this position, you should have the following qualifications and experience: Role Overview: You should possess strong Python coding skills, along with a solid understanding of Object-Oriented Programming (OOP). Additionally, you must have practical experience in working on Big Data product architecture. Your familiarity with SQL-based databases such as MySQL, PostgreSQL, and NoSQL-based databases like Cassandra and Elasticsearch is crucial. Hands-on experience with frameworks like Spark RDD, DataFrame, and Dataset is necessary. Proficiency in developing ETL for data products is also expected from you. Moreover, you should have expertise in performance optimization, optimal resource utilization, parallelism, and tuning of Spark jobs. Your knowledge of file formats including CSV, JSON, XML, PARQUET, ORC, and AVRO is essential. It would be advantageous if you have worked with Analytical Databases like Druid, MongoDB, or Apache Hive. Experience in handling real-time data feeds and familiarity with Apache Kafka or similar tools is a plus. Key Responsibilities: - Strong Python coding skills and OOP skills - Work on Big Data product Architecture - Experience with SQL-based databases (e.g., MySQL, PostgreSQL) and NoSQL-based databases (e.g., Cassandra, Elasticsearch) - Hands-on experience with frameworks like Spark RDD, DataFrame, Dataset - Development of ETL for data product - Working knowledge on performance optimization, optimal resource utilization, parallelism, and tuning of Spark jobs - Familiarity with file formats: CSV, JSON, XML, PARQUET, ORC, AVRO - Good to have experience with Analytical Databases like Druid, MongoDB, Apache Hive - Ability to handle real-time data feeds (knowledge of Apache Kafka or similar tool is a bonus) Qualifications Required: - Strong Python coding skills - Experience with SQL-based and NoSQL-based databases - Hands-on experience with Spark and PySpark - Knowledge of parallel programming Please note that Scala is optional but would be beneficial for this role.

One address, no account. We’ll tell you when matching roles go live.