Source description
About the role
Job Title: Data Scientist LLM, NLP, ML & Anomaly Detection Location: Pune Experience Level: 4-8 years We are looking for a Data Scientist with expertise in Large Language Models (LLMs), Natural Language Processing (NLP), Machine Learning, and Anomaly Detection to join our AI & Data Science team. Key Responsibilities Design, build, and deploy end-to-end AI/ML solutions.Develop NLP applications including:Named Entity Recognition (NER)Text Classification & Sentiment AnalysisInformation ExtractionDocument Summarization & Question AnsweringSemantic Search & Document SimilarityBuild and fine-tune models using Hugging Face Transformers (BERT, RoBERTa, T5, etc.).Process unstructured data from PDFs, emails, OCR documents, and web content.Develop LLM-based applications using OpenAI, Azure OpenAI, LangChain, and LlamaIndex.Design anomaly detection solutions for fraud detection, operational monitoring, and document/text analytics.Implement Statistical, ML-based, Deep Learning, and Time-Series anomaly detection techniques.Build scalable real-time and batch ML pipelines.Perform feature engineering, model training, hyperparameter tuning, and model evaluation.Deploy models using FastAPI/Flask and monitor production performance.Collaborate with engineering, product, and business teams to deliver AI solutions.Create data visualizations and communicate insights to stakeholders.Required Skills 48 years of experience in Data Science or Machine Learning.Strong Python skills (Pandas, NumPy, Scikit-learn, SciPy).Experience with LangChain, LlamaIndex, OpenAI, or Azure OpenAI.Hands-on experience with Hugging Face, spaCy, and NLTK.Expertise in ML and Deep Learning using PyTorch or TensorFlow.Experience with anomaly detection techniques such as Isolation Forest, LOF, One-Class SVM, Autoencoders, and LSTM.Knowledge of Time-Series Analysis (ARIMA, Prophet, LSTM).Proficiency in SQL and Vector Databases (FAISS, Pinecone, Chroma, Weaviate).Experience with AWS, Azure, or GCP.Knowledge of Git, Docker, CI/CD, and REST API development.Experience with Pydantic for structured outputs and data validation.Preferred Skills OCR and document processing tools (Tesseract, Google Vision API, PyMuPDF).MLOps tools such as MLflow, Kubeflow, or Weights & Biases.Kafka, Spark Streaming, Kubernetes, Apache Spark, or Dask.Experience in Financial Services, Fraud Detection, Risk Management, or Compliance.Open-source contributions or AI/ML research publications.Education Bachelor's or Master's degree in Computer Science, Data Science, Statistics, Mathematics, or a related quantitative field. Job Title: Data Scientist LLM, NLP, ML & Anomaly Detection Location: Pune Experience Level: 4-8 years We are looking for a Data Scientist with expertise in Large Language Models (LLMs), Natural Language Processing (NLP), Machine Learning, and Anomaly Detection to join our AI & Data Science team. Key Responsibilities Design, build, and deploy end-to-end AI/ML solutions.Develop NLP applications including:Named Entity Recognition (NER)Text Classification & Sentiment AnalysisInformation ExtractionDocument Summarization & Question AnsweringSemantic Search & Document SimilarityBuild and fine-tune models using Hugging Face Transformers (BERT, RoBERTa, T5, etc.).Process unstructured data from PDFs, emails, OCR documents, and web content.Develop LLM-based applications using OpenAI, Azure OpenAI, LangChain, and LlamaIndex.Design anomaly detection solutions for fraud detection, operational monitoring, and document/text analytics.Implement Statistical, ML-based, Deep Learning, and Time-Series anomaly detection techniques.Build scalable real-time and batch ML pipelines.Perform feature engineering, model training, hyperparameter tuning, and model evaluation.Deploy models using FastAPI/Flask and monitor production performance.Collaborate with engineering, product, and business teams to deliver AI solutions.Create data visualizations and communicate insights to stakeholders.Required Skills 48 years of experience in Data Science or Machine Learning.Strong Python skills (Pandas, NumPy, Scikit-learn, SciPy).Experience with LangChain, LlamaIndex, OpenAI, or Azure OpenAI.Hands-on experience with Hugging Face, spaCy, and NLTK.Expertise in ML and Deep Learning using PyTorch or TensorFlow.Experience with anomaly detection techniques such as Isolation Forest, LOF, One-Class SVM, Autoencoders, and LSTM.Knowledge of Time-Series Analysis (ARIMA, Prophet, LSTM).Proficiency in SQL and Vector Databases (FAISS, Pinecone, Chroma, Weaviate).Experience with AWS, Azure, or GCP.Knowledge of Git, Docker, CI/CD, and REST API development.Experience with Pydantic for structured outputs and data validation.Preferred Skills OCR and document processing tools (Tesseract, Google Vision API, PyMuPDF).MLOps tools such as MLflow, Kubeflow, or Weights & Biases.Kafka, Spark Streaming, Kubernetes, Apache Spark, or Dask.Experience in Financial Services, Fraud Detection, Risk Management, or Compli
More at STRATACENT INC