Experience: 10+ years
What You'll Do - Serve as the hands-on technical lead for integrating OEM and white-label HPC AI/ML, general purpose compute, storage, and network hardware into Lambda's HPC platform reference architectures. - Drive the end-to-end process of new product introduction (NPI) for hardware systems, including system bring-up, documentation, vendor technical engagement, production readiness, and closure of hardware risks. - Identify, debug, and resolve hardware issues across different hardware engineering domains during hardware NPI; support closure of critical fleet issues that require hardware design, vendor corrective action, or platform configuration changes. - Partner with HPC architects to translate platform blueprints into concrete hardware selections and system configurations. - Partner with the supply chain team on new vendor evaluation and QBR/HBR feedback on established vendors. - Own the hardware platform through NPI, working with PMO to de-risk execution, drive cross-functional closure of hardware readiness issues, and ensure platforms reach production on schedule. - Collaborate with the quality team and fleet reliability team during hardware NPI and after production to continuously improve product quality and reliability at scale. - Work cross-functionally with fleet engineering, deployment, operation and datacenter engineering teams to ensure on-time delivery and deployment, quality, compatibility, performance, and scalability of new systems. - Serve as the technical lead to evaluate, enable, and prototype new hardware in labs. - Review BOMs to ensure configuration accuracy, component compatibility, and alignment of key commodities and components to Lambda platform requirements.
You - 5 years of technical lead experience on hardware NPI and deployment for HPC, data center, or cloud infrastructure products, familiar with hardware NPI processes. - Possess deep knowledge and hands-on experiences in one or many of the following hardware platforms: AI/ML, general compute (x86 and ARM), storage systems, or network switches. - Broad hardware engineering domain knowledge in one or many of the below areas: electrical, thermal, mechanical, power, signal integrity, safety, compliance, reliability and manufacturing. - Are comfortable working hands-on in labs to enable and bring up new hardware systems. - Experiences in identifying, triaging and root causing hardware issues during NPI and at scale in the fleet. - Experience in PLM systems and BOM structure. - Collaborate well cross functionally to deliver production-ready hardware solutions. - Strong ownership and can do attitude, self-starter who feels comfortable working in ambiguity.
Nice to Have - 10+ years of technical lead experience on hardware NPI and deployment for HPC, data center, or cloud infrastructure products, familiar with hardware NPI processes. - Experience supporting AI/ML infrastructure and accelerated compute hardware (e.g., NVIDIA, AMD, Intel). - Experience in rack scale server development and liquid cooling designs. - Exposure to fleet observability, BMC/BIOS/Network configuration and automation. - Background in performance tuning, benchmarking, and systems validation workflows. - Can interpret platform-level architecture requirements and select or adapt OEM and white-label solutions to fit. - Prior experience contributing to reference designs or large-scale infrastructure blueprints. - Are experienced with vendor-led product development cycles and can drive hardware evaluation, risk mitigation, and feedback into roadmap decisions.
Search Senior HPC Platform Hardware Engineer jobs near San Jose Office (Zanker) (Remote) → Browse all live jobs
This posting was published by Lambda on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.