Padmi

Bioinformatics Data Engineer ( Python / Shell Scripting / Docker )

MumbaiPosted 1 month ago
Software engineeringMid-levelFull Time; Regular
Apply at Sanskruti Solutions

Opens the source posting on shine.com

Source description

About the role

View original

(Important Notes: (1) It is a Desk Job (Onsite Job), operating for 6 days / week working pattern 5 days working from an office environment and Saturday currently allowed with remote working and (2) Immediate Joiners Preferred) Job Summary: We are looking for a highly skilled Data Engineer with strong programming and workflow automation skills to support and enhance our production bioinformatics pipelines. You do not need a biology background but you must be comfortable working with biological data formats, large datasets, and cloud-based workflows. In this role, you will develop, maintain, and optimize our data processing pipelines used for NGS (Next-Generation Sequencing) analysis, clinical workflows, and research innovation. You will work closely with bioinformaticians, software engineers, and data scientists to build scalable, reliable, and efficient systems. Responsibilities: Pipeline & Data Engineering: Develop and maintain scalable data pipelines for genomic and clinical datasets. Build workflow automation using Python, Shell, Docker, and workflow managers (Nextflow, Snakemake, Airflow, etc.). Optimize existing pipelines for performance, resource usage, and reliability. Handle large biological datasets (FASTQ, BAM, VCF, CSV/TSV, metadata). Software Engineering: Write clean, modular, production-level code in Python and Shell. Implement CI/CD processes for pipeline deployment. Maintain code repositories (Git) and ensure high-quality documentation. Cloud & Infrastructure: Work with AWS/GCP/Azure services for scalable pipeline execution. Manage container-based deployments using Docker. Monitor job performance, logs, and system behavior. Data Management: Maintain data integrity, versioning, and audit trails. Develop automated QC checks and validation workflows. Support data ingestion, transformation, and ETL processes. Collaboration: Work with bioinformaticians to translate analysis logic into scalable workflows. Collaborate with clinical and operations teams to ensure pipeline readiness. Troubleshoot pipeline failures and optimize workflows in production. Required Skills: Core Technical Skills: Strong proficiency in Python Hands-on experience with Shell scripting (bash) Strong understanding of Docker / containerization Knowledge of Git, CI/CD, and software development best practices Experience with workflow orchestration: Nextflow, Snakemake, Airflow, Cromwell, Prefect (any one) Data Engineering Skills: Experience with large datasets, ETL pipelines, log processing Strong understanding of file formats, data transformation, and automation Familiarity with Linux and HPC or distributed systems Bioinformatics Data Handling (Training can be provided): Understanding of at least one biological data type: FASTQ, BAM/CRAM, VCF, BED, GTF Ability to process unstructured or semi-structured scientific data Good to Have Skills (Will Prefer If Available): Experience with R Exposure to ML/AI workflows Experience with cloud (AWS Batch, Lambda, S3, EC2) Knowledge of Next Generation Sequencing (NGS) pipelines Familiarity with database systems (SQL/NoSQL) .

One address, no account. We’ll tell you when matching roles go live.

More at Sanskruti Solutions

Related open roles

View all roles