Source description
About the role
As a Senior Data Scientist specializing in ML & Semantic AI, your primary responsibility will be to build and enhance machine learning models at scale using Python, Azure Cloud, and NLP technologies. You will work on embedding optimization, semantic matching, LDA, and RAG architectures, dense and sparse retrieval pipelines, and migration of cloud-native data pipelines to Azure Databricks. Key Responsibilities: - Design and execute end-to-end machine learning pipelines including data extraction, preprocessing, feature engineering, model development, tuning, and deployment. - Develop machine learning pipelines using Azure Synapse, Databricks, and Snowflake. - Build and deploy classification, regression, and clustering models. - Develop and deploy proof-of-concept solutions for client use cases. - Implement semantic matching and similarity search using cosine similarity, dot-product scoring, and bi-encoder/cross-encoder architectures (e.g., SBERT, sentence-transformers). - Build embedding models by fine-tuning pre-trained models and optimizing embedding storage in vector databases such as Chroma DB, FAISS, and Azure AI Search. Qualifications Required: - Proficiency in Python, Azure Databricks, Azure ML, Azure Synapse, Azure Blob Storage, Scikit-learn, NumPy, Pandas, Hugging Face, sentence-transformers, FAISS, Chroma DB, Azure AI Search, LangChain, TensorFlow, PyTorch, Statsmodels, Azure OpenAI. - Experience in developing forecasting models for marketing, demand prediction, and trend analysis. - Knowledge of NLP-based forecasting techniques using sentiment and external data. - Ability to apply semantic similarity for audience intelligence, including zero-shot and few-shot classification techniques. - Strong understanding of data extraction, preprocessing, feature engineering, model development, tuning, and deployment. Location: DGS India - Mumbai - Thane Ashar IT Park Brand: Merkle Time Type: Full time Contract Type: Permanent As a Senior Data Scientist specializing in ML & Semantic AI, your primary responsibility will be to build and enhance machine learning models at scale using Python, Azure Cloud, and NLP technologies. You will work on embedding optimization, semantic matching, LDA, and RAG architectures, dense and sparse retrieval pipelines, and migration of cloud-native data pipelines to Azure Databricks. Key Responsibilities: - Design and execute end-to-end machine learning pipelines including data extraction, preprocessing, feature engineering, model development, tuning, and deployment. - Develop machine learning pipelines using Azure Synapse, Databricks, and Snowflake. - Build and deploy classification, regression, and clustering models. - Develop and deploy proof-of-concept solutions for client use cases. - Implement semantic matching and similarity search using cosine similarity, dot-product scoring, and bi-encoder/cross-encoder architectures (e.g., SBERT, sentence-transformers). - Build embedding models by fine-tuning pre-trained models and optimizing embedding storage in vector databases such as Chroma DB, FAISS, and Azure AI Search. Qualifications Required: - Proficiency in Python, Azure Databricks, Azure ML, Azure Synapse, Azure Blob Storage, Scikit-learn, NumPy, Pandas, Hugging Face, sentence-transformers, FAISS, Chroma DB, Azure AI Search, LangChain, TensorFlow, PyTorch, Statsmodels, Azure OpenAI. - Experience in developing forecasting models for marketing, demand prediction, and trend analysis. - Knowledge of NLP-based forecasting techniques using sentiment and external data. - Ability to apply semantic similarity for audience intelligence, including zero-shot and few-shot classification techniques. - Strong understanding of data extraction, preprocessing, feature engineering, model development, tuning, and deployment. Location: DGS India - Mumbai - Thane Ashar IT Park Brand: Merkle Time Type: Full time Contract Type: Permanent
More at Dentsu