Source description
About the role
We are looking for a highly skilled Senior Data Scientist with strong expertise in PySpark, Machine Learning, and Big Data technologies to design, build, and deploy production-grade machine learning solutions. The ideal candidate will have hands-on experience working with cloud platforms and MLOps practices in real-world environments. Key Responsibilities Design, develop, and deploy scalable machine learning models for production useWork extensively with PySpark and Apache Spark for large-scale data processingBuild and optimize end-to-end ML pipelines on cloud platformsApply advanced statistical, machine learning, and deep learning techniquesDevelop and maintain production-grade ML systemsCollaborate with data engineers, software engineers, and business stakeholdersImplement MLOps/DevOps practices including CI/CD pipelinesMonitor, evaluate, and improve model performance in productionRequired Skills Strong proficiency in Python & PySparkSolid understanding of Machine Learning & StatisticsHands-on experience with Apache Spark, Hadoop, KafkaExperience with Big Data platforms: AWS / Google Cloud / Azure / Oracle CloudExpertise in ML techniques: Tree-based models (XGBoost, Random Forest, LightGBM)Deep Learning (CNNs, RNNs, GNNs, Transformers)Reinforcement Learning and Unsupervised LearningExperience in building production-ready ML modelsKnowledge of MLOps / DevOps tools: Docker, Kubernetes, CI/CDFamiliarity with TensorFlow / PyTorch / Scikit-learnQualifications Bachelors or Masters degree in a STEM discipline4 to 6 years of experience as a Data Scientist or ML EngineerProven track record of delivering ML solutions in productionPreferred Skills (Good to Have) Experience with MLflow, KubeflowCloud certifications (AWS/GCP/Azure)Experience working in Agile/Scrum teams We are looking for a highly skilled Senior Data Scientist with strong expertise in PySpark, Machine Learning, and Big Data technologies to design, build, and deploy production-grade machine learning solutions. The ideal candidate will have hands-on experience working with cloud platforms and MLOps practices in real-world environments. Key Responsibilities Design, develop, and deploy scalable machine learning models for production useWork extensively with PySpark and Apache Spark for large-scale data processingBuild and optimize end-to-end ML pipelines on cloud platformsApply advanced statistical, machine learning, and deep learning techniquesDevelop and maintain production-grade ML systemsCollaborate with data engineers, software engineers, and business stakeholdersImplement MLOps/DevOps practices including CI/CD pipelinesMonitor, evaluate, and improve model performance in productionRequired Skills Strong proficiency in Python & PySparkSolid understanding of Machine Learning & StatisticsHands-on experience with Apache Spark, Hadoop, KafkaExperience with Big Data platforms: AWS / Google Cloud / Azure / Oracle CloudExpertise in ML techniques: Tree-based models (XGBoost, Random Forest, LightGBM)Deep Learning (CNNs, RNNs, GNNs, Transformers)Reinforcement Learning and Unsupervised LearningExperience in building production-ready ML modelsKnowledge of MLOps / DevOps tools: Docker, Kubernetes, CI/CDFamiliarity with TensorFlow / PyTorch / Scikit-learnQualifications Bachelors or Masters degree in a STEM discipline4 to 6 years of experience as a Data Scientist or ML EngineerProven track record of delivering ML solutions in productionPreferred Skills (Good to Have) Experience with MLflow, KubeflowCloud certifications (AWS/GCP/Azure)Experience working in Agile/Scrum teams
More at People Prime Worldwide