Source description
About the role
About Codewalla Codewalla is a New Yorkbased product studio with engineering teams in India. Since 2005, weve built innovative products that scale. We work at the intersection of design, engineering, and AI developing systems shaped by real business needs and tested in the real world. Our team moves fast, thinks deeply, and cares about pushing what software can do to empower people and businesses. About The Role Were hiring a Gen AI Engineer with 68 years of engineering experience, including at least one year shipping LLM-powered features. Your north star is to turn raw LLM potential into reliable features and products. If fast feedback loops, hands-on execution, and code that ships straight into users hands sound like your kind of work, wed love to talk. What Youll Work On - Build MCP servers and agentic clients that handle user intent parsing, tool orchestration, and structured response generation. - Architect effective RAG pipelines with chunk decay, latency budgeting, and cost-aware vector search. - Automate evaluation pipelines that test LLM outputs for relevance, accuracy, and coherence. - Work closely with DevOps to codify and deploy infrastructure using CDK or Terraform. - Set up observability dashboards for prompt performance, latency, and failure traceability. - Continuously refine prompts, embeddings, and model behavior based on user feedback and regression tests. What Makes You a Great Fit - 5 to 8 years of full-stack or backend development experience, with at least 1 year building AI-powered or LLM-based applications. - AI-native mindset : test fast, trace deeply, pause to reframe when needed. - Strong Python and TypeScript skills. - Experience with either AWS or GCP stacks, such as : - AWS : Lambda, Bedrock, DynamoDB, OpenSearch Vector Search. - GCP : Cloud Functions, Vertex AI, Firestore, BigQuery, Vector Search. - Familiarity with LangChain, Bedrock SDK, and vector database schema design. - Understanding of prompt design, embeddings, and agen .