Baseten

Software Engineer - Model Performance

Full-time · San Francisco (Remote)
✓ Verified live on the employer's own system · added 864 days ago
Save search

Requirements

Education: Bachelor's degree or related field

Skills & tools

Machine LearningTroubleshootingProgrammingPythonC Plus Plus
Apply on company site ↗ See your fit → free

Full job description

We are looking for a Software Engineer focused on ML performance to join our dynamic team. This role is ideal for someone who thrives in a fast-paced startup environment and is eager to make significant contributions to the exciting field of LLM Inference. If you are a backend engineer who thrives on making things faster and is excited about open-source ML models, we look forward to your application.

You'll get to work on these types of projects as part of our Model Performance team:

- Baseten Embeddings Inference: The fastest embeddings solution available

- Implement, refine, and productionize cutting-edge techniques (quantization, speculative decoding, kv cache reuse, chunked prefill and LoRA) for ML model inference and infrastructure.

- Deep dive into underlying codebases of TensorRT, PyTorch, TensorRT-LLM, vllm, sglang, CUDA, and other libraries to debug ML performance issues.

- Apply and scale optimization techniques across a wide range of ML models, particularly large language models.

- Collaborate with a diverse team to design and implement innovative solutions.

- Bachelor's, Master's, or Ph.D. degree in Computer Science, Engineering, Mathematics, or related field.

- Experience with one or more general-purpose programming languages, such as Python or C++.

- Familiarity with LLM optimization techniques (e.g., quantization, speculative decoding, continuous batching).

- Strong familiarity with ML libraries, especially PyTorch, TensorRT, or TensorRT-LLM.

- Proficiency in enhancing the performance of software systems, particularly in the context of large language models (LLMs).

- Deep understanding of software engineering principles and a proven track record of developing and deploying AI/ML inference solutions.

More jobs at Baseten

Similar jobs near San Francisco (Remote)

Tell me when more Software Engineer, Monetization ML Infrastructure jobs post near San Francisco (Remote) We re-check every listing against the employer’s own board — no résumé needed.

Search Software Engineer - Model Performance jobs near San Francisco (Remote) → Browse all live jobs

This posting was published by Baseten on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.