Padmi

ML Engineer / PySpark

IndiaPosted 3 months ago
Software engineeringMid-levelFull Time; Regular
Apply at Gloify

Opens the source posting on shine.com

Source description

About the role

View original

Role Overview: You will be joining the Test & Learn Platform team as an ML Engineer. Your primary responsibility will involve building and scaling the experimentation and causal inference services. This includes developing statistical engines, API integrations, and cloud pipelines to empower business teams globally in making data-driven decisions. Key Responsibilities: - Develop and maintain statistical/ML modules (DID, Synthetic Control, A/B Testing, Multi-Treatment Effects) in Python - Build and extend Quick API services and integrate them with our web application via SDK wrappers - Design and optimize large-scale data pipelines using PySpark, Delta Lake, and Azure Data Lake - Profile and resolve OOM issues in PySpark jobs - optimize memory allocation, partitioning, broadcast joins, caching strategies, and Spark configurations - Deploy and manage workloads on Databricks, including job clusters, notebooks, and Delta Lake tables - Containerize and deploy services using Docker, Kubernetes, and CI/CD pipelines - Ensure code quality and security via Sonar Cloud, Snyk, and PyTest - Collaborate with data scientists and product teams to translate research into production-ready modules Qualifications Required: - 3+ years of production experience in Python (3.9+) - Strong experience with PySpark & Spark Internals, including understanding Spark memory model, executor tuning, shuffle optimization, and diagnosing/resolving OOM errors - Hands-on experience with Databricks, job orchestration, cluster configuration, notebook workflows, and Delta Lake optimization - Proficiency in Causal Inference & Experimentation methodologies such as DID, synthetic control, A/B testing, hypothesis testing, and panel data methods - Familiarity with Statistics/ML Libraries like statsmodels, scikit-learn, scipy, pandas, numpy - Experience in API Development, building RESTful services with FastAPI or similar frameworks - Knowledge of Cloud platforms like Azure, specifically Azure Storage, Azure ML, and Data Lake - Proficiency in Docker & Kubernetes for containerization and orchestration of ML workloads - Experience in writing robust unit/integration tests with pytest Additional Company Details: Good-to-Have: - Experience with Celery/Redis for async task orchestration - Familiarity with Polars, PyArrow, or SQL Alchemy - Background in econometrics or experimental design - Proficiency in Spark UI profiling and performance benchmarking - Knowledge of CI/CD tooling such as Sonar Cloud, Snyk, GitHub Actions Role Overview: You will be joining the Test & Learn Platform team as an ML Engineer. Your primary responsibility will involve building and scaling the experimentation and causal inference services. This includes developing statistical engines, API integrations, and cloud pipelines to empower business teams globally in making data-driven decisions. Key Responsibilities: - Develop and maintain statistical/ML modules (DID, Synthetic Control, A/B Testing, Multi-Treatment Effects) in Python - Build and extend Quick API services and integrate them with our web application via SDK wrappers - Design and optimize large-scale data pipelines using PySpark, Delta Lake, and Azure Data Lake - Profile and resolve OOM issues in PySpark jobs - optimize memory allocation, partitioning, broadcast joins, caching strategies, and Spark configurations - Deploy and manage workloads on Databricks, including job clusters, notebooks, and Delta Lake tables - Containerize and deploy services using Docker, Kubernetes, and CI/CD pipelines - Ensure code quality and security via Sonar Cloud, Snyk, and PyTest - Collaborate with data scientists and product teams to translate research into production-ready modules Qualifications Required: - 3+ years of production experience in Python (3.9+) - Strong experience with PySpark & Spark Internals, including understanding Spark memory model, executor tuning, shuffle optimization, and diagnosing/resolving OOM errors - Hands-on experience with Databricks, job orchestration, cluster configuration, notebook workflows, and Delta Lake optimization - Proficiency in Causal Inference & Experimentation methodologies such as DID, synthetic control, A/B testing, hypothesis testing, and panel data methods - Familiarity with Statistics/ML Libraries like statsmodels, scikit-learn, scipy, pandas, numpy - Experience in API Development, building RESTful services with FastAPI or similar frameworks - Knowledge of Cloud platforms like Azure, specifically Azure Storage, Azure ML, and Data Lake - Proficiency in Docker & Kubernetes for containerization and orchestration of ML workloads - Experience in writing robust unit/integration tests with pytest Additional Company Details: Good-to-Have: - Experience with Celery/Redis for async task orchestration - Familiarity with Polars, PyArrow, or SQL Alchemy - Background in econometrics or experimental design - Proficiency in Spark UI profiling and performance benchmarking - Knowledge of CI/CD tooling such as

One address, no account. We’ll tell you when matching roles go live.

More at Gloify

Related open roles

View all roles