Padmi

Data Science Lead - Generative AI

HyderabadPosted 3 months ago
Data Science And StatisticsSeniorFull Time; Regular
Apply at Jasper Colin

Opens the source posting on shine.com

Source description

About the role

View original

As a Data Science Lead, your role involves designing and deploying Generative AI systems using Large Language Models (LLMs) and embeddings. You will be responsible for optimizing chunking strategies, fine-tuning transformer-based models, implementing and managing vector databases, and building machine learning models for various applications. Collaboration with engineering teams, monitoring performance metrics, staying updated with the latest research, mentoring junior data scientists, and contributing to technical leadership are also key aspects of your responsibilities. Key Responsibilities: - Lead the design and deployment of Retrieval-Augmented Generation (RAG) systems using Large Language Models (LLMs) and embeddings - Develop and optimize chunking strategies for improved retrieval and context relevance - Fine-tune transformer-based models (e.g., GPT, LLaMA, Mistral) using techniques like LoRA, PEFT, or full-model training - Implement and manage vector databases (e.g., FAISS, Pinecone, Weaviate) for semantic search - Build and evaluate machine learning models for classification, regression, clustering, and recommendation systems - Collaborate with engineering teams to integrate ML and Generative AI pipelines into production - Define and monitor performance metrics for model accuracy, latency, and scalability - Stay updated with the latest research in Generative AI, Large Language Models (LLMs), and MLOps - Mentor junior data scientists and contribute to technical leadership Required Skills & Tools: - Programming: Python, SQL - ML Libraries: scikit-learn, XGBoost, LightGBM, TensorFlow, PyTorch - GenAI Frameworks: Hugging Face Transformers, LangChain, OpenAI API - Vector Databases: FAISS, Pinecone, Weaviate, Qdrant - LLM Fine-Tuning: LoRA, PEFT, RLHF, prompt tuning - Data Engineering: Pandas, NumPy, Spark - Deployment: FastAPI, Docker, Kubernetes, CI/CD - Cloud Platforms: AWS, GCP, Azure - Strong problem-solving and communication skills - Bachelors or Masters degree in Computer Science, Data Science, AI, or related field - 5 - 7 years of experience in data science, NLP, or machine learning - Proven track record of building and deploying ML and Generative AI applications Preferred Attributes: - Experience with open-source Large Language Models (LLMs) such as Mistral, LLaMA, Falcon - Familiarity with semantic search, prompt engineering, and retrieval optimization - Contributions to AI research or open-source projects - Certification in Generative AI, NLP, or Machine Learning (e.g., DeepLearning.AI, Hugging Face) As a Data Science Lead, your role involves designing and deploying Generative AI systems using Large Language Models (LLMs) and embeddings. You will be responsible for optimizing chunking strategies, fine-tuning transformer-based models, implementing and managing vector databases, and building machine learning models for various applications. Collaboration with engineering teams, monitoring performance metrics, staying updated with the latest research, mentoring junior data scientists, and contributing to technical leadership are also key aspects of your responsibilities. Key Responsibilities: - Lead the design and deployment of Retrieval-Augmented Generation (RAG) systems using Large Language Models (LLMs) and embeddings - Develop and optimize chunking strategies for improved retrieval and context relevance - Fine-tune transformer-based models (e.g., GPT, LLaMA, Mistral) using techniques like LoRA, PEFT, or full-model training - Implement and manage vector databases (e.g., FAISS, Pinecone, Weaviate) for semantic search - Build and evaluate machine learning models for classification, regression, clustering, and recommendation systems - Collaborate with engineering teams to integrate ML and Generative AI pipelines into production - Define and monitor performance metrics for model accuracy, latency, and scalability - Stay updated with the latest research in Generative AI, Large Language Models (LLMs), and MLOps - Mentor junior data scientists and contribute to technical leadership Required Skills & Tools: - Programming: Python, SQL - ML Libraries: scikit-learn, XGBoost, LightGBM, TensorFlow, PyTorch - GenAI Frameworks: Hugging Face Transformers, LangChain, OpenAI API - Vector Databases: FAISS, Pinecone, Weaviate, Qdrant - LLM Fine-Tuning: LoRA, PEFT, RLHF, prompt tuning - Data Engineering: Pandas, NumPy, Spark - Deployment: FastAPI, Docker, Kubernetes, CI/CD - Cloud Platforms: AWS, GCP, Azure - Strong problem-solving and communication skills - Bachelors or Masters degree in Computer Science, Data Science, AI, or related field - 5 - 7 years of experience in data science, NLP, or machine learning - Proven track record of building and deploying ML and Generative AI applications Preferred Attributes: - Experience with open-source Large Language Models (LLMs) such as Mistral, LLaMA, Falcon - Familiarity with semantic search, prompt engineering, and retrieval optimization - Contributions to AI resear

One address, no account. We’ll tell you when matching roles go live.

More at Jasper Colin

Related open roles

View all roles