Padmi

Lead Data Engineer

IndiaPosted 5 months ago
Infrastructure And DatabasesSeniorFull Time
Apply at Providence

Opens the source posting on foundit.in

Source description

About the role

View original

About the Role We are looking for a handson Data Engineer to design, build, and operate robust data pipelines and platforms on Snowflake with Azure. You will use strong SQL , Python/PySpark , ADF pipelines , and modern datamodeling practices to ingest data from diverse data sources, and enable AI/ML usecases via VectorDB indexing and embeddings . The role emphasizes reliability, performance, costefficiency, and secure data operations in line with our enterprise platforms and standards. Key Responsibilities Design & build data pipelines on Snowflake and Azure (ADF, PySpark) to ingest data from REST APIs, files, and databases into curated zones. Model data optimized for analytics, reporting, and downstream applications. Develop embeddings & VectorDB indices to power semantic search/retrieval (e.g., generating embeddings and indexing into enterpriseapproved vector stores integrate with pipeline orchestration). Own performance & cost optimization in Snowflake (SQL tuning, partitioning, caching, clustering, compute sizing). Implement CI/CD and DevOps practices (Git branching, automated deploys for ADF/Snowflake). Harden reliability (monitoring, alerting, retry logic, SLA tracking) and security/compliance (RBAC, secrets management, data governance, data lineage). Collaborate with stakeholders (product, analytics, and platform teams) to translate requirements into technical design and deliver incremental value. MustHave Qualifications 4-7 years total experience in data engineering in large scale enterprise systems Snowflake : Min 3 years of experience in Snowflake with exposure to warehouse configuration, schema design, performance tuning stored procedures/tasks loading strategies. Exposure to Snowflake Cortex AI. SQL/Python/PySpark : Design and implement scalable data processing solutions using SQL, Python, and distributed compute frameworks, including unit/integration tests. Azure & ADF : ADLS Gen2, ADF pipelines/activities, triggers, parameterization monitoring & troubleshooting. Data modeling : Apply data modeling techniques, including medallion architecture (Bronze/Silver/Gold). API ingestion : designing resilient ingestion of REST/JSON, pagination, auth, ratelimit handling. VectorDB & embeddings : Experience generating embeddings and building vector indices for retrievalaugmented scenariosExposure to building knowledge graphs and Gremlin or Cypher graph query languages on CosmosDB/Neo4j Version control & CI/CD : Git, pull requests, automated deployment pipelines.Maintain a results-oriented mindset with strong analytical and problem-solving skills. GoodtoHave Experience in Healthcare IndustryPrior experience working on data migration projects.

One address, no account. We’ll tell you when matching roles go live.

More at Providence

Related open roles

View all roles