Padmi

Python and SQL Developer

IndiaPosted 3 months ago
Software engineeringSeniorFull Time; Regular
Apply at WarpDrive Tech Works LLP

Opens the source posting on shine.com

Source description

About the role

View original

As a Data Engineer, your key responsibility will be to design and develop scalable data pipelines to process and transform large datasets on local systems. You will be expected to: - Build and manage ETL pipelines to transform large volumes of data efficiently. - Optimize data processing using generators, iterators, and chunked file processing to minimize memory usage. - Implement parallel and asynchronous processing using multi-threading, multi-processing, and asyncio. - Write and optimize complex SQL queries using window functions, CTEs, and indexing. - Manage large-scale data operations including bulk inserts, updates, and partitioning. - Perform memory and performance tuning using profiling tools like cProfile and memory_profiler. Qualifications required for this role include: - 5+ years of experience in data engineering with Python. - Advanced proficiency in Python libraries including Pandas, NumPy, and SQLAlchemy. - Expertise in SQL query optimization, execution plan analysis, and partitioning. - Experience handling large files in CSV, JSON, and Parquet formats. - Proficiency in memory management techniques such as streaming and compression (gzip, zlib). - Ability to use parallel processing tools like Joblib and Dask. - Strong knowledge of data structures optimization for speed and memory efficiency. - Experience with ETL processes and data validation mechanisms. In addition to the above key responsibilities and qualifications, the following preferred skills would be advantageous: - Knowledge of PySpark. - Experience with AWS S3. As a Data Engineer, your key responsibility will be to design and develop scalable data pipelines to process and transform large datasets on local systems. You will be expected to: - Build and manage ETL pipelines to transform large volumes of data efficiently. - Optimize data processing using generators, iterators, and chunked file processing to minimize memory usage. - Implement parallel and asynchronous processing using multi-threading, multi-processing, and asyncio. - Write and optimize complex SQL queries using window functions, CTEs, and indexing. - Manage large-scale data operations including bulk inserts, updates, and partitioning. - Perform memory and performance tuning using profiling tools like cProfile and memory_profiler. Qualifications required for this role include: - 5+ years of experience in data engineering with Python. - Advanced proficiency in Python libraries including Pandas, NumPy, and SQLAlchemy. - Expertise in SQL query optimization, execution plan analysis, and partitioning. - Experience handling large files in CSV, JSON, and Parquet formats. - Proficiency in memory management techniques such as streaming and compression (gzip, zlib). - Ability to use parallel processing tools like Joblib and Dask. - Strong knowledge of data structures optimization for speed and memory efficiency. - Experience with ETL processes and data validation mechanisms. In addition to the above key responsibilities and qualifications, the following preferred skills would be advantageous: - Knowledge of PySpark. - Experience with AWS S3.

One address, no account. We’ll tell you when matching roles go live.