Source description
About the role
As an experienced candidate with 46 years of experience in ML/NLP, particularly in document-heavy domains such as finance, legal, and policy, you will be responsible for the following key areas: - Data Ingestion and Preprocessing: - Build and maintain data pipelines to ingest unstructured data from PDFs, gazettes, HTML circulars, etc. - Process data extraction, parsing, and normalization. - NLP & LLM Modeling: - Fine-tune or prompt-tune LLMs for summarization, classification, and change detection in regulations. - Develop embeddings for semantic similarity. - Knowledge Graph Engineering: - Design entity relationships (regulation, control, policy). - Implement retrieval over Neo4j or similar graph DBs. - Information Retrieval (RAG): - Build RAG pipelines for natural language querying of regulations. - Annotation and Validation: - Annotate training data by collaborating with Subject Matter Experts (SMEs). - Validate model outputs. - MLOps: - Build CI/CD for model retraining, versioning, and evaluation (precision, recall, BLEU, etc.). - API and Integration: - Expose ML models as REST APIs (FastAPI) for integration with product frontend. Additionally, the required skills for this role include proficiency in: - Languages: Python, SQL - AI/ML/NLP: Hugging face transformers, OpenAI API, Spacy, Scikit-Learn, LangChain, RAG, LLM prompt-tuning, LLM fine-tuning - Vector Search: Pinecone, Weaviate, FAISS - Data Engineering: Airflow, Kafka, OCR (Tesseract, pdfminer) - MLOps: MLflow, Docker If you believe you possess the necessary skills and experience for this role, we encourage you to apply now. As an experienced candidate with 46 years of experience in ML/NLP, particularly in document-heavy domains such as finance, legal, and policy, you will be responsible for the following key areas: - Data Ingestion and Preprocessing: - Build and maintain data pipelines to ingest unstructured data from PDFs, gazettes, HTML circulars, etc. - Process data extraction, parsing, and normalization. - NLP & LLM Modeling: - Fine-tune or prompt-tune LLMs for summarization, classification, and change detection in regulations. - Develop embeddings for semantic similarity. - Knowledge Graph Engineering: - Design entity relationships (regulation, control, policy). - Implement retrieval over Neo4j or similar graph DBs. - Information Retrieval (RAG): - Build RAG pipelines for natural language querying of regulations. - Annotation and Validation: - Annotate training data by collaborating with Subject Matter Experts (SMEs). - Validate model outputs. - MLOps: - Build CI/CD for model retraining, versioning, and evaluation (precision, recall, BLEU, etc.). - API and Integration: - Expose ML models as REST APIs (FastAPI) for integration with product frontend. Additionally, the required skills for this role include proficiency in: - Languages: Python, SQL - AI/ML/NLP: Hugging face transformers, OpenAI API, Spacy, Scikit-Learn, LangChain, RAG, LLM prompt-tuning, LLM fine-tuning - Vector Search: Pinecone, Weaviate, FAISS - Data Engineering: Airflow, Kafka, OCR (Tesseract, pdfminer) - MLOps: MLflow, Docker If you believe you possess the necessary skills and experience for this role, we encourage you to apply now.