Padmi

AI/ML & Data Engineer

ChennaiPosted 2 months ago
Software engineeringMid-levelFull Time; Regular
Apply at Congruent Info-Tech

Opens the source posting on shine.com

Source description

About the role

View original

As an experienced candidate with 46 years of experience in ML/NLP, particularly in document-heavy domains such as finance, legal, and policy, you will be responsible for the following key areas: - Data Ingestion and Preprocessing: - Build and maintain data pipelines to ingest unstructured data from PDFs, gazettes, HTML circulars, etc. - Process data extraction, parsing, and normalization. - NLP & LLM Modeling: - Fine-tune or prompt-tune LLMs for summarization, classification, and change detection in regulations. - Develop embeddings for semantic similarity. - Knowledge Graph Engineering: - Design entity relationships (regulation, control, policy). - Implement retrieval over Neo4j or similar graph DBs. - Information Retrieval (RAG): - Build RAG pipelines for natural language querying of regulations. - Annotation and Validation: - Annotate training data by collaborating with Subject Matter Experts (SMEs). - Validate model outputs. - MLOps: - Build CI/CD for model retraining, versioning, and evaluation (precision, recall, BLEU, etc.). - API and Integration: - Expose ML models as REST APIs (FastAPI) for integration with product frontend. Additionally, the required skills for this role include proficiency in: - Languages: Python, SQL - AI/ML/NLP: Hugging face transformers, OpenAI API, Spacy, Scikit-Learn, LangChain, RAG, LLM prompt-tuning, LLM fine-tuning - Vector Search: Pinecone, Weaviate, FAISS - Data Engineering: Airflow, Kafka, OCR (Tesseract, pdfminer) - MLOps: MLflow, Docker If you believe you possess the necessary skills and experience for this role, we encourage you to apply now. As an experienced candidate with 46 years of experience in ML/NLP, particularly in document-heavy domains such as finance, legal, and policy, you will be responsible for the following key areas: - Data Ingestion and Preprocessing: - Build and maintain data pipelines to ingest unstructured data from PDFs, gazettes, HTML circulars, etc. - Process data extraction, parsing, and normalization. - NLP & LLM Modeling: - Fine-tune or prompt-tune LLMs for summarization, classification, and change detection in regulations. - Develop embeddings for semantic similarity. - Knowledge Graph Engineering: - Design entity relationships (regulation, control, policy). - Implement retrieval over Neo4j or similar graph DBs. - Information Retrieval (RAG): - Build RAG pipelines for natural language querying of regulations. - Annotation and Validation: - Annotate training data by collaborating with Subject Matter Experts (SMEs). - Validate model outputs. - MLOps: - Build CI/CD for model retraining, versioning, and evaluation (precision, recall, BLEU, etc.). - API and Integration: - Expose ML models as REST APIs (FastAPI) for integration with product frontend. Additionally, the required skills for this role include proficiency in: - Languages: Python, SQL - AI/ML/NLP: Hugging face transformers, OpenAI API, Spacy, Scikit-Learn, LangChain, RAG, LLM prompt-tuning, LLM fine-tuning - Vector Search: Pinecone, Weaviate, FAISS - Data Engineering: Airflow, Kafka, OCR (Tesseract, pdfminer) - MLOps: MLflow, Docker If you believe you possess the necessary skills and experience for this role, we encourage you to apply now.

One address, no account. We’ll tell you when matching roles go live.