Experience: 3+ years
Baseten's Model Performance (MP) team is responsible for ensuring the models running on our platform are fast, reliable, and cost-efficient. As part of this team, you'll focus on Model API's - the infrastructure powering our hosted API endpoints for the latest open-source models. This work spans distributed systems, model serving, and developer experience.
You'll join a small, high-impact team operating at the intersection of product, model performance, and infra, helping to define how developers interact with AI models at scale.
- Design, build, and operate the Model APIs surface with focus on advanced inference capabilities: structured outputs (JSON mode, grammar-constrained generation), tool/function calling and multi-modal serving
- Profile and optimize TensorRT-LLM kernels, analyze CUDA kernel performance, implement custom CUDA operators, tune memory allocation patterns for maximum throughput and optimize communication patterns across multi-GPU setups
- Productionize performance improvements across runtimes with deep understanding of their internals: speculative decoding implementations, guided generation for structured outputs, custom scheduling and routing algorithms for high-performance serving
- Build comprehensive benchmarking frameworks that measure real-world performance across different model architectures, batch sizes, sequence lengths, and hardware configurations
- Productionize performance improvements across runtimes (e.g.TensorRT, TensorRT-LLM): speculative decoding, quantization, batching, and KV-cache reuse.
- Instrument deep observability (metrics, traces, logs) and build repeatable benchmarks to measure speed, reliability, and quality.
- Implement platform fundamentals: API versioning, validation, usage metering, quotas, and authentication.
- Collaborate closely with other teams to deliver robust, developer-friendly model serving experiences.
- 3+ years experience building and operating distributed systems or large-scale APIs.
- Proven track record of owning low-latency, reliable backend services (rate-limiting, auth, quotas, metering, migrations).
- Infra instincts with performance sensibilities: profiling, tracing, capacity planning, and SLO management.
- Comfortable debugging complex systems, from runtime internals to GPU execution traces.
- Strong written communication; able to produce clear design docs and collaborate across functions.
- Experience with LLM runtimes (vLLM, SGLang, TensorRT-LLM) or contributions to open-source inference engines (vLLM, TensorRT-LLM, SGLang, TGI)
- Knowledge of Kubernetes, service meshes, API gateways, or distributed scheduling.
- Background in developer-facing infrastructure or open-source APIs.
- We value infra-leaning generalists who bring strong engineering fundamentals and curiosity. ML experience is a plus, but not required.
Search Software Engineer - Model Products jobs near San Francisco (Remote) → Browse all live jobs
This posting was published by Baseten on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.