Source description
About the role
On-Premises LLM & Vector Database Implementation Consultant
Location: Philadelphia, PA | Hybrid (3+ days/week onsite) Position Type: Contract | Experience: 5-10 years
[About the Role] Join a pioneering team as an On-Premises LLM & Vector Database Implementation Consultant. Drive the end-to-end deployment of state-of-the-art open-source large language models (LLMs) and vector databases in highly secure, enterprise-grade on-premises environments. This is a hands-on, mid-level consulting opportunity to shape cutting-edge AI infrastructure while collaborating with leading technologists and business stakeholders.
[Responsibilities]
- Deploy, integrate, and optimize open-source LLMs (e.g., Meta Llama 3, Mistral/Mixtral) in private, air-gapped environments
- Develop and implement Retrieval-Augmented Generation (RAG) pipelines using open-source vector databases (Qdrant, Chroma, Milvus, pgvector)
- Engineer Python-based solutions for LLM inference, prompt design, and system integration
- Tune model performance through CPU-based inference, quantization, and metadata management
- Ensure strict data privacy and enterprise security, including access controls and audit logging
- Deliver clear reference architecture, prototype solutions, and comprehensive documentation
- Lead knowledge transfer sessions to upskill internal engineering teams
[Required Skills and Experience]
- 5-10 years in software development with proven on-premises LLM deployment experience
- Advanced Python proficiency for LLM workflows and system integration
- Hands-on work with vector databases and RAG pipeline implementation
- Expertise in model quantization, CPU inference, and performance optimization
- Solid knowledge of enterprise security, access control, and data governance in private environments
[Preferred Skills]
- Experience with LangChain, LlamaIndex
- Familiarity with Rust, Go, or C++ for high-performance backend services
- Exposure to Docker, Kubernetes for on-premises deployments
- Knowledge of vLLM, llama.cpp, Hugging Face Transformers frameworks
- Prior work in regulated, enterprise, or security-focused environments
[Benefits]
- Work with cutting-edge AI technologies in a secure, enterprise setting
- Hybrid work model for enhanced work-life balance
- Opportunity for career growth and upskilling in advanced AI and infrastructure engineering
- Gain exposure to real-world, high-impact AI applications and top-tier professionals
[How to Apply]
Ready to make a tangible impact in the AI space? Submit your resume and a brief cover letter outlining your experience with on-prem LLM and vector database deployments. Onsite interviews required. Take the next step in your AI consulting career with us in Philadelphia!
More at Indotronix International
Related open roles
JPC - 205931 - Power BI Developer
Los Angeles
JPC - 201157 - Cloud Software Developer
United States
JPC - 206290 - Information & Application Developer Level 3
Remote · Remote Job - Yes
JPC - 201195 - Information Technology - Embedded S/W Engineer
United States
JPC - 205720 - Automation Solutions Architect
United States · Hybrid
JPC - 204931 - Senior ML Engineer – Deployment and Databricks MLOps
Austin · Onsite