Padmi

Bioinformatics Data Engineer ( Python / Shell Scripting / Docker ) (Maharashtra)

IndiaPosted 1 month ago
Software engineeringJuniorFull Time; Regular
Apply at Sanskruti Solutions

Opens the source posting on shine.com

Source description

About the role

View original
  • (Important Notes: (1) It is a Desk Job (Onsite Job), operating for 6 days / week working pattern - 5 days working from an office environment and Saturday currently allowed with remote working and (2) Immediate Joiners Preferred) Job Summary: - We are looking for a highly skilled Data Engineer with strong programming and workflow automation skills to support and enhance our production bioinformatics pipelines. - You do not need a biology background - but you must be comfortable working with biological data formats, large datasets, and cloud-based workflows. - In this role, you will develop, maintain, and optimize our data processing pipelines used for NGS (Next-Generation Sequencing) analysis, clinical workflows, and research innovation. - You will work closely with bioinformaticians, software engineers, and data scientists to build scalable, reliable, and efficient systems. Responsibilities: - Pipeline & Data Engineering: - Develop and maintain scalable data pipelines for genomic and clinical datasets. - Build workflow automation using Python, Shell, Docker, and workflow managers (Nextflow, Snakemake, Airflow, etc.). - Optimize existing pipelines for performance, resource usage, and reliability. - Handle large biological datasets (FASTQ, BAM, VCF, CSV/TSV, metadata). - Software Engineering: - Write clean, modular, production-level code in Python and Shell. - Implement CI/CD processes for pipeline deployment. - Maintain code repositories (Git) and ensure high-quality documentation. - Cloud & Infrastructure: - Work with AWS/GCP/Azure services for scalable pipeline execution. - Manage container-based deployments using Docker. - Monitor job performance, logs, and system behavior. - Data Management: - Maintain data integrity, versioning, and audit trails. - Develop automated QC checks and validation workflows. - Support data ingestion, transformation, and ETL processes. - Collaboration: - Work with bioinformaticians to translate analysis logic into scalable workflows. - Collaborate with clinical and operations teams to ensure pipeline readiness. - Troubleshoot pipeline failures and optimize workflows in production. Required Skills: Core Technical Skills: - Strong proficiency in Python - Hands-on experience with Shell scripting (bash) - Strong understanding of Docker / containerization - Knowledge of Git, CI/CD, and software development best practices - Experience with workflow orchestration: Nextflow, Snakemake, Airflow, Cromwell, Prefect (any one) Data Engineering Skills: - Experience with large datasets, ETL pipelines, log processing - Strong understanding of file formats, data transformation, and automation - Familiarity with Linux and HPC or distributed systems Bioinformatics Data Handling (Training can be provided): - Understanding of at least one biological data type: FASTQ, BAM/CRAM, VCF, BED, GTF - Ability to process unstructured or semi-structured scientific data Positive to Have Skills (Will Prefer If Available): - Experience with R - Exposure to ML/AI workflows - Experience with cloud (AWS Batch, Lambda, S3, EC2) - Knowledge of Next Generation Sequencing (NGS) pipelines - Familiarity with database systems (SQL/NoSQL) .

One address, no account. We’ll tell you when matching roles go live.

More at Sanskruti Solutions

Related open roles

View all roles