Source description
About the role
You will be responsible for designing, building, and optimizing production-grade systems that ground LLM responses in enterprise knowledge. Your key tasks will include: - Designing and implementing robust RAG pipelines, including ingestion, parsing, enrichment, indexing, retrieval, reranking, and answer generation. - Choosing and tuning retrieval strategies to maximize recall and precision for real enterprise queries. - Building citation/grounding mechanisms and response policies to ensure traceable, trustworthy outputs. Furthermore, you will be required to: - Implement and optimize vector and hybrid search over structured and unstructured data. - Develop reranking strategies and fusion methods to improve ranking quality. - Establish query understanding and rewriting techniques to enhance retrieval robustness. In addition, you will need to: - Define an evaluation harness for retrieval and generation using offline datasets and online telemetry. - Implement automated regression tests and quality gates for new prompts, retrievers, and model updates. - Create feedback loops using human review and lightweight labeling to enhance relevance over time. Moreover, you will be expected to: - Optimize latency and throughput using caching, batching, and efficient retrieval/index configurations. - Instrument the full pipeline with logs, metrics, traces, dashboards, and alerting. - Drive cost-aware design across embedding, retrieval, and generation. Your qualifications should include: - Bachelor's degree in Computer Science, Engineering, Data Science, Human-Computer Interaction, or a related field with 5+ years of relevant experience; OR a Master's/PhD with 3+ years of relevant experience. - Strong programming skills in Python and experience with LLM/RAG development in production environments. - Experience with vector databases or search engines, retrieval concepts, and evaluation methods. - Experience building scalable services and APIs, with attention to reliability and performance. - Strong understanding of data processing pipelines, metadata design, and information retrieval fundamentals. Preferred qualifications would include experience with ranking techniques, document parsing, observability practices for AI systems, ACL-aware retrieval, and prompt/tooling libraries. In the first 6-12 months, you will be expected to achieve: - A standardized RAG pipeline that improves answer relevance while reducing hallucinations and unresolved queries. - A repeatable evaluation framework with quality gates to prevent regressions. - Meaningful latency and cost reductions via caching, adaptive retrieval, and efficient strategies. - Secure, compliant retrieval that enforces access control without sacrificing search quality. You will be responsible for designing, building, and optimizing production-grade systems that ground LLM responses in enterprise knowledge. Your key tasks will include: - Designing and implementing robust RAG pipelines, including ingestion, parsing, enrichment, indexing, retrieval, reranking, and answer generation. - Choosing and tuning retrieval strategies to maximize recall and precision for real enterprise queries. - Building citation/grounding mechanisms and response policies to ensure traceable, trustworthy outputs. Furthermore, you will be required to: - Implement and optimize vector and hybrid search over structured and unstructured data. - Develop reranking strategies and fusion methods to improve ranking quality. - Establish query understanding and rewriting techniques to enhance retrieval robustness. In addition, you will need to: - Define an evaluation harness for retrieval and generation using offline datasets and online telemetry. - Implement automated regression tests and quality gates for new prompts, retrievers, and model updates. - Create feedback loops using human review and lightweight labeling to enhance relevance over time. Moreover, you will be expected to: - Optimize latency and throughput using caching, batching, and efficient retrieval/index configurations. - Instrument the full pipeline with logs, metrics, traces, dashboards, and alerting. - Drive cost-aware design across embedding, retrieval, and generation. Your qualifications should include: - Bachelor's degree in Computer Science, Engineering, Data Science, Human-Computer Interaction, or a related field with 5+ years of relevant experience; OR a Master's/PhD with 3+ years of relevant experience. - Strong programming skills in Python and experience with LLM/RAG development in production environments. - Experience with vector databases or search engines, retrieval concepts, and evaluation methods. - Experience building scalable services and APIs, with attention to reliability and performance. - Strong understanding of data processing pipelines, metadata design, and information retrieval fundamentals. Preferred qualifications would include experience with ranking techniques, document parsing, observability practices
More at KLA