We are looking for Senior Software Engineers to join our team. This is a specialized, high-impact role sitting at the intersection of high-performance computing (HPC) and Large Language Model (LLM) engineering. You will not just be building the automated "speedometer and diagnostic" suite for our next-generation AI infrastructure; you will be defining the roadmap, driving key technical decisions, and taking full ownership of the future of this work.
- Benchmarking : Evaluate, run and automate standard LLM quality benchmarks (GSM8K, MMLU) alongside custom performance suites for specific workloads (e.g., long-context window, KV cache reuse, disaggregated serving).
- DevEx Improvement : Develop and maintain internal GPU-enabled development environments (similar to GitHub Codespaces). You will ensure the team has seamless, high-performance "dev machines" optimized for model experimentation.
- Tool Development : Build and contribute to open-source tools such as InferenceMAX and genai-bench to automate model evaluation, benchmarking and analysis.
- System Profiling : Use profilers like PyTorch Profiler, NVIDIA Nsight Systems and py-spy to collect performance profiles, identify bottlenecks, and debug the compute/networking stack.
- Monitoring & Observability : Develop real-time dashboards and alerts to monitor system health, model startup times, and runtime performance.
- Continuous Integration : Automate performance testing via CI/CD pipelines to catch regressions and build release workflow automation for the model runtimes stack.
- Optimization Automation : Build tools to find the "Pareto frontier"-identifying the absolute best configuration (latency vs. cost vs. quality) for a given model and workload.
This is a mid-senior, high leverage role. We care about your technical depth, strong communication skills to drive cross-team efforts, ability to navigate vague requirements and mentor other engineers. We want to talk to you if you have:
- A Love for Systems & Hardware : You aren't just interested in the AI; you want to understand GPU memory subsystems, InfiniBand, and how data moves across a cluster.
- An Automation Mindset : You believe that if a task has to be done twice, it should be scripted. You have a passion for stress-testing and fuzzy testing to find the "breaking point" of a system.
- Mathematical Curiosity : A desire to understand the underlying math of Transformers and how it translates into FLOPs and memory requirements.
- Technical Toolkit : Familiarity with Python, and an eagerness to master the NVIDIA software stack. C++ familiarity is good to have.
- Direct Impact : Your tools will be the gatekeeper for what defines "good" performance for our customers.
- Deep Learning (Literally) : You will gain world-class expertise in GPU orchestration and LLM inference that few engineers in the industry possess.
- High Ownership : As the lead of a small team, you will have the autonomy to build tools from scratch and contribute to open-source projects.
Search Software Engineer- Model Performance Systems jobs near San Francisco (Remote) → Browse all live jobs
This posting was published by Baseten on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.