Source description
About the role
Founding Engineer, AI Infra
San Francisco Bay Area, CA
Goaly
Hybrid
Full-time
About Goaly
At Goaly, our mission is to make custom AI affordable for every business. Our founding team comes from the front lines of top AI labs and tech giants (Meta MSL, TikTok AI, Google DeepMind, xAI, Microsoft Research, etc.), where we built large-scale training infrastructure powering trillion-parameter models and scaled GenAI models to a global user base. Now, we are building something we wish we had before: a platform that makes training and adapting custom AI affordable for all modern companies, not just Big Tech. Our north star is ambitious: for a domain-specific task, reach 90% of SOTA performance at less than 10% of the cost. To get a taste of what we are doing, see our first tech blog.
About the Role
You will sit at the intersection of systems engineering and applied ML, building specialized infrastructure that keeps large language and multimodal models fast, reliable, and cost-effective. You will partner with research, product, and infra teams to ship production-ready platforms for training and serving AI at scale.
Key Responsibilities
- Efficiency & performance: Improve LLM training and inference efficiency through better memory utilization, optimized parallelism, and kernel-level innovations (e.g. FlashAttention, CUDA/Triton).
- Training & RL robustness: Build scalable, stable training and RL pipelines with strong reproducibility, observability, and debuggability.
- Serving & inference optimization: Design and tune high-throughput, low-latency model serving systems, including quantization, caching, and speculative decoding.
- Scalability & infrastructure: Own end-to-end training and inference infrastructure — from data ingestion and checkpointing to multi-GPU and multi-cloud orchestration.
- Production enablement: Work closely with researchers and product engineers to turn new algorithms into reliable, production-ready systems.
Requirements
- 5+ years building or operating ML infrastructure at scale, ideally supporting large language or multimodal models.
- Deep understanding of GPU architecture, distributed training frameworks (PyTorch, DeepSpeed, Megatron, Ray), and parallelism strategies.
- Hands-on experience running inference stacks (vLLM / SGLang, TGI, Triton) and optimizing them via low-level profiling.
- Strong software engineering fundamentals in Python and one of C++/Rust/Go, with clean, reliable code shipped to production.
- Working knowledge of modern data pipelines, feature stores, and vector databases used in production AI systems.
- Comfort automating infrastructure with Kubernetes, Terraform/Pulumi, and observability stacks (Prometheus, Grafana, OpenTelemetry).
Bonus Points
-
Experience deploying open-source LLMs (Llama 3, Qwen, DeepSeek) or training custom foundation models.
-
Contributions to ML systems tooling (compilers, kernels, inference runtimes) or open-source infrastructure projects.
-
Background in reinforcement learning, evaluation harnesses, or alignment tooling that hardens production AI systems.
-
Ready to apply?
-
Powered by
-
First name *
-
Last name *
-
Email *
-
LinkedIn URL
-
Phone number
-
Location *
-
Resume *
-
Click to upload or drag and drop here
-
Are you legally authorized to work in the United States? *
-
Yes
-
No
-
Will you now, or in the future, require sponsorship for employment visa status (e.g. H-1B visa status)? *
-
Yes
-
No
-
Are you currently able to work onsite in a hybrid capacity in the San Francisco Bay Area, or willing to relocate to work onsite if offered this role? *
-
Yes
-
No
-
Voluntary Self-Identification
To comply with government reporting requirements, we invite candidates to participate in the self-identification survey below. Your completion of this form is entirely optional, and your decision will neither influence the hiring process nor any subsequent stages. Any information you choose to share will be kept confidential and stored in a secure file. As outlined in our Equal Employment Opportunity policy, we uphold a commitment to non-discrimination based on any protected group status specified in applicable laws.
Gender
Please select
Race
Please select
Race and ethnicity descriptionsExpand
Voluntary Self-Identification of Veteran Status
VEVRAA requires Government contractors to take affirmative action to employ and advance in employment protected veterans. To help us measure the effectiveness of our outreach and recruitment efforts of veterans, we are asking you to tell us if you are a veteran covered by VEVRAA. If you believe that you belong to any of the following categories of protected veterans, please indicate by making the appropriate selection.
Veteran status descriptionsCollapse
Disabled veteran
A veteran who served on active duty in the U.S. military and is entitled to disability compensation (or who but for the receipt of military retired pay would be entitled to disability compensation) under laws administered by the Secretary of Veterans Affairs, or was discharged or released from active duty because of a service-connected disability
Recently separated veteran
A veteran separated during the three-year period beginning on the date of the veteran's discharge or release from active duty in the U.S military, ground, naval, or air service
Active duty wartime or campaign badge veteran
A veteran who served on active duty in the U.S. military during a war, or in a campaign or expedition for which a campaign badge was authorized under the laws administered by the Department of Defense
Armed Forces service medal veteran
A veteran who, while serving on active duty in the U.S. military ground, naval, or air service, participated in a United States military operation for which an Armed Forces service medal was awarded pursuant to Executive Order 12985 (61 Fed. Reg. 1209).
Veteran status
Please select
Apply
Req ID: R16
More at Exponential
Related open roles
Senior Software Engineer (Robotics Systems & Infrastructure)
San Francisco Bay Area · Onsite
Research Scientist/Engineer, Efficient ML Systems
San Francisco Bay Area · Hybrid
Robotics Software Engineer at BudBreak Innovations
San Francisco Bay Area · Hybrid
Full Stack Engineer at BudBreak
San Francisco Bay Area · Hybrid
Senior Robotics Software Engineer at Spacer Robotics
San Francisco Bay Area · Onsite
Chief Technology Officer at LauchPad Build AI
Los Angeles · Onsite
