Source description
About the role
As a Principal Software Architect at Applied Materials, you will be responsible for designing and implementing robust, scalable infrastructure solutions combining diverse processors such as CPUs, GPUs, and FPGAs. Your primary focus will be on analyzing and partitioning workloads to the most appropriate compute unit, collaborating with cross-functional teams to translate requirements into architectural/software designs, and coding and developing quick prototypes to establish your design with real code and data. You will also be expected to be a subject matter expert in the HPC domain, conducting performance tuning, capacity planning, and monitoring GPU metrics for reliability. Key Responsibilities: - Design and implement robust, scalable infrastructure solutions combining CPUs, GPUs, and FPGAs - Analyze and partition workloads to the most appropriate compute unit - Collaborate with cross-functional teams to translate requirements into architectural/software designs - Code and develop quick prototypes to establish your design with real code and data - Be a subject matter expert in the HPC domain - Conduct performance tuning, capacity planning, and monitoring GPU metrics for reliability - Evaluate and recommend appropriate technologies and frameworks to meet project requirements - Lead the design and implementation of complex software components and systems - Ensure software systems are scalable, reliable, and maintainable Qualifications: - 12 to 18 years of experience in implementing robust, scalable, and secure infrastructure solutions combining CPUs, GPUs, and FPGAs - Working experience with GPU inference servers like Nvidia Triton - Proficiency in C/C++, Data Structures, Algorithms, and complexity analysis - Experience in developing Distributed High-Performance Computing software using Parallel programming frameworks like MPI, UCX, etc. - Proficiency in GPU programming using CUDA, OpenMP, OpenACC, OpenCL, etc. - In-depth experience in Multi-threading, Thread Synchronization, Inter-process communication, and distributed computing fundamentals - Experience in performance profiling at the application and system level - Familiarity with microservices architecture, containerization technologies, and low-latency message queues - Excellent problem-solving and analytical skills - Strong communication and collaboration abilities - Ability to mentor and coach junior team members - Experience in Agile development methodologies Additional Qualifications: - Experience in HPC Job-Scheduling and Cluster Management Software - Good knowledge of Low-latency and high-throughput data transfer technologies - Good knowledge of Parallel processing and DAG execution Frameworks If this opportunity excites you and you have a passion for designing and developing cutting-edge HPC systems and heterogeneous computing infrastructure, then we encourage you to apply. Join us at Applied Materials to be a part of a leading global company at the forefront of material innovation that changes the world. As a Principal Software Architect at Applied Materials, you will be responsible for designing and implementing robust, scalable infrastructure solutions combining diverse processors such as CPUs, GPUs, and FPGAs. Your primary focus will be on analyzing and partitioning workloads to the most appropriate compute unit, collaborating with cross-functional teams to translate requirements into architectural/software designs, and coding and developing quick prototypes to establish your design with real code and data. You will also be expected to be a subject matter expert in the HPC domain, conducting performance tuning, capacity planning, and monitoring GPU metrics for reliability. Key Responsibilities: - Design and implement robust, scalable infrastructure solutions combining CPUs, GPUs, and FPGAs - Analyze and partition workloads to the most appropriate compute unit - Collaborate with cross-functional teams to translate requirements into architectural/software designs - Code and develop quick prototypes to establish your design with real code and data - Be a subject matter expert in the HPC domain - Conduct performance tuning, capacity planning, and monitoring GPU metrics for reliability - Evaluate and recommend appropriate technologies and frameworks to meet project requirements - Lead the design and implementation of complex software components and systems - Ensure software systems are scalable, reliable, and maintainable Qualifications: - 12 to 18 years of experience in implementing robust, scalable, and secure infrastructure solutions combining CPUs, GPUs, and FPGAs - Working experience with GPU inference servers like Nvidia Triton - Proficiency in C/C++, Data Structures, Algorithms, and complexity analysis - Experience in developing Distributed High-Performance Computing software using Parallel programming frameworks like MPI, UCX, etc. - Proficiency in GPU programming using CUDA, OpenMP, OpenACC, OpenCL, etc. - In-depth experience in
More at Applied Materials