SpaceX

Site Reliability Engineer, GNC

$125KFull-time · Hawthorne, CA
✓ Verified live on the employer's own system · added 103 days ago
Save search
Junior · 2+ yrs exp

Requirements

Education: Bachelor's degree

Experience: 2+ years

What this role involves

Site ReliabilityReliability EngineeringAnsibleTerraformIncident ResponseKubernetesDevops

Skills & tools

OperationsData AnalysisDevopsProgrammingWeb DevelopmentTroubleshootingPythonManagement

Benefits — mentioned in this posting

Equity / stockBonus / commission401(k) / retirementHealth, dental & visionFamily / parental leavePaid time off
Apply on company site ↗ See your fit → free

Full job description

we are seeking a Site Reliability Engineer to operate and scale custom-built, mission-critical products for the Guidance, Navigation, and Control (GNC) teams.

GNC teams at SpaceX are responsible for vehicle design, trajectory design and optimization, high-fidelity vehicle simulation, software and control algorithm development, while also supporting both launch and on-orbit operations across multiple vehicle programs. In this role, you will work closely with GNC teams across SpaceX to maintain and improve a suite of critical GNC-focused tools and infrastructure that must scale reliably to enable a multiplanetary future.

These systems include on-prem services, large-scale Monte Carlo simulations on our high-performance computing (HPC) cluster, automated data analysis pipelines, continuous integration systems for rocket and simulation software, GNC analysis infrastructure, and vehicle configuration verification tools.

The ideal candidate is flexible, possesses broad skills spanning product operations and software development, and thrives in a fast-paced, high-impact environment.

- Deploy, upgrade, operate, and scale a suite of mission-critical GNC products and services

- Work with SpaceX HPC team to monitor and maintain an HPC cluster consisting of tens of thousands of CPUs.

- Closely collaborate with GNC software engineers to create highly operable and maintainable products

- Monitoring and incident response for web applications and services

- Manage the underlying computational infrastructure of GNC in collaboration with IT stakeholders

- Engage in and improve the whole lifecycle of services from whiteboard to operational

- Make data-driven recommendations for future hardware purchases

- Provide end-user support to GNC engineering for products by becoming an expert on analysis applications and support users in troubleshooting and pointing to features

- Develop or improve GNC web apps and tools for better usability, maintainability, and robustness

- Demo and document new software changes such as operating system upgrades, shared filesystem changes, or major tool rollouts

- Focus on performance bottlenecks and performance improvement techniques

- Bachelor's degree in computer science, information systems/IT, engineering, math, or scientific discipline and 2+ years of software development experience OR 4+ years of professional experience building software with site reliability or DevOps in lieu of a degree

- 1+ years of experience with Python and Python based development frameworks

- 2+ years of systems administration, site reliability engineering, or DevOps experience

- 2+ years of experience with Python and Python-based development frameworks

- Expertise with Docker, Vagrant, and Kubernetes or similar technologies

- Extensive Experience with configuration management tools such as Ansible, Puppet, Terraform

- Experience with build systems (Make, Bazel / Pants / Buck, Gradle) and package management tools (pip, npm)

- Strong understanding of virtualization and hypervisor technologies

- Experience with automatically managing dozens or hundreds of servers

- Experience scaling web applications and optimizing applications for performance

- Experience with managing on-prem infrastructure, including direct experience managing GPU fleets

- Experience with high-performance computing systems or large-scale data analysis systems

- Must be comfortable working with mission-critical and sensitive systems, with a sense of urgency appropriate to the responsibilities

- An active clearance may provide the opportunity for you to work on sensitive SpaceX missions; if so, you will be subject to pre-employment drug and random drug and alcohol testing

- Willing to work extended hours and weekends when needed to meet critical deadlines

Pay Range: Site Reliability Engineer/Level I: $125,000.00 - $145,000.00/per year Site Reliability Engineer/Level II: $145,000.00 - $175,000.00/per year

Base salary is just one part of your total rewards package at SpaceX. You may also be eligible for long-term incentives, in the form of company stock, stock options, or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan.

You will also receive access to comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short and long-term disability insurance, life insurance, paid parental leave, and various other discounts and perks. You may also accrue 3 weeks of paid vacation and will be eligible for 10 or more paid holidays per year.

Employees accrue paid sick leave pursuant to Company policy which satisfies or exceeds the accrual, carryover, and use requirements of the law.

More jobs at SpaceX

Similar jobs near Hawthorne, CA

Tell me when more Site Reliability Engineer, Client Platform jobs post near Los Angeles, CA We re-check every listing against the employer’s own board — no résumé needed.

Search Site Reliability Engineer, GNC jobs near Hawthorne, CA → Browse all live jobs

This posting was published by SpaceX on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.