Padmi

JPC - 202407 - On-Premises LLM & Vector Database Implementation Consultant

Philadelphia · HybridPosted 1 month ago
Software engineeringUnspecified
Apply at Indotronix International

Opens the source posting on iic.com

Source description

About the role

View original

On-Premises LLM & Vector Database Implementation Consultant

Location: Philadelphia, PA | Hybrid (3+ days/week onsite) Position Type: Contract | Experience: 5-10 years

[About the Role] Join a pioneering team as an On-Premises LLM & Vector Database Implementation Consultant. Drive the end-to-end deployment of state-of-the-art open-source large language models (LLMs) and vector databases in highly secure, enterprise-grade on-premises environments. This is a hands-on, mid-level consulting opportunity to shape cutting-edge AI infrastructure while collaborating with leading technologists and business stakeholders.

[Responsibilities]

  • Deploy, integrate, and optimize open-source LLMs (e.g., Meta Llama 3, Mistral/Mixtral) in private, air-gapped environments
  • Develop and implement Retrieval-Augmented Generation (RAG) pipelines using open-source vector databases (Qdrant, Chroma, Milvus, pgvector)
  • Engineer Python-based solutions for LLM inference, prompt design, and system integration
  • Tune model performance through CPU-based inference, quantization, and metadata management
  • Ensure strict data privacy and enterprise security, including access controls and audit logging
  • Deliver clear reference architecture, prototype solutions, and comprehensive documentation
  • Lead knowledge transfer sessions to upskill internal engineering teams

[Required Skills and Experience]

  • 5-10 years in software development with proven on-premises LLM deployment experience
  • Advanced Python proficiency for LLM workflows and system integration
  • Hands-on work with vector databases and RAG pipeline implementation
  • Expertise in model quantization, CPU inference, and performance optimization
  • Solid knowledge of enterprise security, access control, and data governance in private environments

[Preferred Skills]

  • Experience with LangChain, LlamaIndex
  • Familiarity with Rust, Go, or C++ for high-performance backend services
  • Exposure to Docker, Kubernetes for on-premises deployments
  • Knowledge of vLLM, llama.cpp, Hugging Face Transformers frameworks
  • Prior work in regulated, enterprise, or security-focused environments

[Benefits]

  • Work with cutting-edge AI technologies in a secure, enterprise setting
  • Hybrid work model for enhanced work-life balance
  • Opportunity for career growth and upskilling in advanced AI and infrastructure engineering
  • Gain exposure to real-world, high-impact AI applications and top-tier professionals

[How to Apply]

Ready to make a tangible impact in the AI space? Submit your resume and a brief cover letter outlining your experience with on-prem LLM and vector database deployments. Onsite interviews required. Take the next step in your AI consulting career with us in Philadelphia!


One address, no account. We’ll tell you when matching roles go live.

More at Indotronix International

Related open roles

View all roles