Baseten

Software Engineer- Model Performance Systems

Full-time · San Francisco (Remote)
✓ Verified live on the employer's own system · added 214 days ago
Save search

Skills & tools

Machine LearningGitTroubleshootingDevopsCommunicationsTeam LeadershipPythonC Plus Plus
Apply on company site ↗ See your fit → free

Full job description

We are looking for Senior Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will not just be building the automated "speedometer and diagnostic" suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work.

- Benchmarking : Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving).

- DevEx Improvement : Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation.

- Tool Development : Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis.

- System Profiling : Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack.

- Monitoring & Observability : Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance.

- Continuous Integration : Automate performance testing via CI/CD pipelines to catch regressions and build release workflow automation for the model runtimes stack.

- Optimization Automation : Build tools to find the "Pareto frontier"-identifying the absolute best configuration (latency vs. cost vs. quality) for a given model and workload.

This is a mid-senior, high leverage role. We care about your technical depth, strong communication skills to drive cross-team efforts, ability to navigate vague requirements and mentor other engineers. We want to talk to you if you have:

- A Love for Systems & Hardware : You aren't just interested in the AI; you want to understand GPU memory subsystems, InfiniBand, and how data moves across a cluster.

- An Automation Mindset : You believe that if a task has to be done twice, it should be scripted. You have a passion for stress-testing and fuzzy testing to find the "breaking point" of a system.

- Mathematical Curiosity : A desire to understand the underlying math of Transformers and how it translates into FLOPs and memory requirements.

- Technical Toolkit : Familiarity with Python, and an eagerness to master the NVIDIA software stack. C++ familiarity is good to have.

- Direct Impact : Your tools will be the gatekeeper for what defines "good" performance for our customers.

- Deep Learning (Literally) : You will gain world-class expertise in GPU orchestration and LLM inference that few engineers in the industry possess.

- High Ownership : As the lead of a small team, you will have the autonomy to build tools from scratch and contribute to open-source projects.

More jobs at Baseten

Similar jobs near San Francisco (Remote)

Tell me when more Software Engineer, Private Computing jobs post near San Francisco (Remote) We re-check every listing against the employer’s own board — no résumé needed.

Search Software Engineer- Model Performance Systems jobs near San Francisco (Remote) → Browse all live jobs

This posting was published by Baseten on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.