Source description
About the role
3+ years of software engineering with at least 1-2 years focused on LLM applications or AI systems in production
Hands-on experience building agentic workflows with tool calling, retrieval, and multi-step reasoning
Deep understanding of prompt engineering, context engineering, and how to get reliable behavior from LLMs
Experience building evaluation and quality systems for AI outputs
Strong Python skills and backend engineering fundamentals
You've shipped AI features to real users and dealt with the messy parts: hallucinations, edge cases, accuracy degradation, cost management
Based in SF.
HELPFUL EXPERIENCE: Agent frameworks: LangGraph, CrewAI, Claude Code/Codex patterns, or custom orchestration
Retrieval systems: vector databases (Qdrant, pgvector, Pinecone), reranking, hybrid search
MCP, tool-calling protocols, and third-party API integrations
Fine-tuning, LoRA, or other model adaptation methods
Evaluation frameworks and continuous quality monitoring
Experience with enterprise AI deployments (compliance, audit trails, governance)
Prior work at AI labs, AI-native startups, or applied ML teams
More at W3Global