Padmi

NLP Engineer

Delhi NCRPosted 2 months ago
Software engineeringSeniorFull Time; Regular
Apply at CareerNet Technologies Pvt Ltd

Opens the source posting on shine.com

Source description

About the role

View original

As an experienced NLP engineer, you will be responsible for designing and building end-to-end text extraction pipelines for a variety of document types, including policy, regulatory, fintech, and healthcare documents. Your role will involve extracting key entities, structuring policy clauses and obligations, fine-tuning BERT and RoBERTa models for NER and text classification tasks, and leveraging LLM APIs using prompt engineering and structured output extraction. Additionally, you will build scalable Python pipelines for processing high volumes of PDF, DOCX, and HTML files, define and enforce JSON schemas, and evaluate model performance to ensure extraction quality. Key Responsibilities: - Design and build end-to-end text extraction pipelines for complex document types - Extract key entities and structure policy clauses and obligations - Fine-tune BERT and RoBERTa for NER, text classification, and relation extraction tasks - Leverage LLM APIs using prompt engineering, tool/function calling, and structured output extraction - Build scalable Python pipelines for high-volume processing of PDF, DOCX, and HTML files - Define and enforce JSON schemas to ensure outputs are compatible with knowledge graph ingestion - Evaluate model performance and implement feedback loops to improve extraction quality Qualifications Required: - 5+ years of hands-on NLP engineering in production pipelines - Proficiency in Python - Experience with NLP libraries such as spaCy and NLTK - Familiarity with HuggingFace Transformers and deep learning techniques - Knowledge of LLM API integration and data pipeline development - Experience working with JSON schemas and Pydantic - Any Graduate degree The preferred skills for this role include experience with legal, regulatory, or policy documents, familiarity with knowledge graphs or graph databases like Neo4j or RDF, expertise in document parsing tools such as pdfplumber, Docling, or Apache Tika, domain knowledge in fintech or healthcare NLP, and exposure to information extraction benchmarks like CoNLL, DocRED, or SciERC. As an experienced NLP engineer, you will be responsible for designing and building end-to-end text extraction pipelines for a variety of document types, including policy, regulatory, fintech, and healthcare documents. Your role will involve extracting key entities, structuring policy clauses and obligations, fine-tuning BERT and RoBERTa models for NER and text classification tasks, and leveraging LLM APIs using prompt engineering and structured output extraction. Additionally, you will build scalable Python pipelines for processing high volumes of PDF, DOCX, and HTML files, define and enforce JSON schemas, and evaluate model performance to ensure extraction quality. Key Responsibilities: - Design and build end-to-end text extraction pipelines for complex document types - Extract key entities and structure policy clauses and obligations - Fine-tune BERT and RoBERTa for NER, text classification, and relation extraction tasks - Leverage LLM APIs using prompt engineering, tool/function calling, and structured output extraction - Build scalable Python pipelines for high-volume processing of PDF, DOCX, and HTML files - Define and enforce JSON schemas to ensure outputs are compatible with knowledge graph ingestion - Evaluate model performance and implement feedback loops to improve extraction quality Qualifications Required: - 5+ years of hands-on NLP engineering in production pipelines - Proficiency in Python - Experience with NLP libraries such as spaCy and NLTK - Familiarity with HuggingFace Transformers and deep learning techniques - Knowledge of LLM API integration and data pipeline development - Experience working with JSON schemas and Pydantic - Any Graduate degree The preferred skills for this role include experience with legal, regulatory, or policy documents, familiarity with knowledge graphs or graph databases like Neo4j or RDF, expertise in document parsing tools such as pdfplumber, Docling, or Apache Tika, domain knowledge in fintech or healthcare NLP, and exposure to information extraction benchmarks like CoNLL, DocRED, or SciERC.

One address, no account. We’ll tell you when matching roles go live.

More at CareerNet Technologies Pvt Ltd

Related open roles

View all roles