Source description
About the role
About the job: You will play a key role in building and optimising high-performance AI kernels for a next-generation compute platform. This role focuses on enabling efficient execution of AI and LLM workloads by developing, profiling, and optimising kernels across varying hardware configurations.You will be responsible for:Developing AI/LLM kernels and operators for efficient inference on a specialised compute platformOptimising kernel performance across different hardware configurations and workloadsProfiling and analysing performance across compute, memory, and parallelism to identify bottlenecksOptimising low-level C/C++ code to maximise hardware utilisationCollaborating across the AI inference stack, including runtime, compiler, and system layersContributing to improvements in toolchain, compiler, and runtime componentsSupporting internal teams and external stakeholders with technical insights and documentationIdeal CandidateYou have a Bachelors or Masters degree in Computer Science, Electrical Engineering, or a related fieldYou have 5+ years of experience in AI kernel development and performance optimisationYou have experience profiling model and kernel inference performanceYou have hands-on experience with at least one of the following: CUDA, DSP, NEON, or TritonYou have strong proficiency in C/C++ and Python; exposure to assembly is a plusYou have strong problem-solving, debugging, and communication skillsYou are comfortable working close to hardware and across system layersWhats on OfferCompetitive compensation with meaningful equityHigh-impact role in a deeply technical, low-bureaucracy environmentOpportunity to work on cutting-edge AI systems and long-term career growthAbout the employerOur client is a Silicon Valleybased deep-tech company building a new compute architecture for real-time AI at the edge. Founded by engineers from leading research backgrounds, the focus is on solving the gaps in current neural processing approaches through tight integration of hardware and software.The platform is built to run both neural network inference and conventional compute workloads efficiently across a wide range of edge devices. Unlike typical accelerators that only handle parts of an ML graph, this architecture supports end-to-end execution, including both neural network graph code and standard C++ DSP and control code, enabling greater flexibility and performance in real-world deployments. Who can apply: Only those candidates can apply who: have minimum 7 years of experience are Computer Science Engineering students Salary: Competitive salary Experience: 7 year(s) Deadline: 2026-09-23 23:59:59
More at A Snaphunt Client