Experience: 5+ years
What You'll Do - Ensure new server, storage and network infrastructure is properly racked, labeled, cabled, and configured. - Troubleshoot hardware and software issues in some of the world's most advanced GPU and Networking systems. - Document and update data center layout and network topology in DCIM software - Work with supply chain & manufacturing teams to ensure timely deployment of systems and project plans for large-scale deployments - Manage a parts depot inventory and track equipment through the delivery-store-stage-deploy-handoff process in each of our data centers - Partner with HW Support teams to ensure data center hardware incidents with higher level troubleshooting challenges are resolved, reported on and solutions are disseminated to the large operations organization. - Work with RMA team to ensure faulty parts are returned and replacements are ordered - Follow installation standards and documentation for placement, labeling, and cabling to drive consistency and discoverability across all data centers You - Have strong past experiences with critical infrastructure systems supporting data centers, such as power distribution, air flow management, environmental monitoring, capacity planning, DCIM software, structured cabling, and cable management - Be familiar with carrier DIA circuit test and turn ups, fiber testing and troubleshooting - Basic knowledge of cable optics and the different types of use - Solid understanding of single and three phase power theories - PDU balancing and why it is important - Familiar with multiple cable media types and their uses - Knowledge of cold isle and hot isle containment - Solid understanding of server hardware and boot process - Ability to structure, collaborate and iteratively improve on complex maintenance MOPs. - Working with product management, support, and other teams to align operational capabilities with company goals. - Translating business priorities into technical and operational requirements. - Supporting cross-functional projects where infrastructure plays a critical role. - Are action-oriented and willingness to train junior staff on best practices - Are willing to travel for bring up of new data center locations as needed (25%-30%) Nice to Have - Have 5+ years experience with critical infrastructure systems supporting data centers, such as power distribution, air flow management, environmental monitoring, capacity planning, DCIM software, structured cabling, and cable management - Experience with/or knowledge of network topology and configurations and 400gb Infiniband architectures. - Experience with/or knowledge of DDP or SCM cluster storage systems. - Have 5+ years working with and reporting from a ticketing systems like JIRA and Zendesk - Advanced experience with Linux administration - Experience with High Performance Compute GPU systems (air or water cooled) - especially Nvidia NVL72 Salary Range Information This is a salaried non-exempt role, eligible for overtime.
The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
Search Data Center Operations Systems Engineer III (Los Angeles) jobs near Vernon, CA - Data Center → Browse all live jobs
This posting was published by Lambda on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.