Baseten

Software Engineer - Model Products

Full-time · San Francisco (Remote)
✓ Verified live on the employer's own system · added 302 days ago
Save search
Mid-level · 3+ yrs exp

Requirements

Experience: 3+ years

Skills & tools

Rest ApisDistributed SystemsMachine LearningRecordkeepingSalesManagementTroubleshootingCommunications
Apply on company site ↗ See your fit → free

Full job description

Baseten's Model Performance (MP) team is responsible for ensuring the models running on our platform are fast, reliable, and cost-efficient. As part of this team, you'll focus on Model API's - the infrastructure powering our hosted API endpoints for the latest open-source models. This work spans distributed systems, model serving, and developer experience.

You'll join a small, high-impact team operating at the intersection of product, model performance, and infra, helping to define how developers interact with AI models at scale.

- Design, build, and operate the Model APIs surface with focus on advanced inference capabilities: structured outputs (JSON mode, grammar-constrained generation), tool/function calling and multi-modal serving

- Profile and optimize TensorRT-LLM kernels, analyze CUDA kernel performance, implement custom CUDA operators, tune memory allocation patterns for maximum throughput and optimize communication patterns across multi-GPU setups

- Productionize performance improvements across runtimes with deep understanding of their internals: speculative decoding implementations, guided generation for structured outputs, custom scheduling and routing algorithms for high-performance serving

- Build comprehensive benchmarking frameworks that measure real-world performance across different model architectures, batch sizes, sequence lengths, and hardware configurations

- Productionize performance improvements across runtimes (e.g.TensorRT, TensorRT-LLM): speculative decoding, quantization, batching, and KV-cache reuse.

- Instrument deep observability (metrics, traces, logs) and build repeatable benchmarks to measure speed, reliability, and quality.

- Implement platform fundamentals: API versioning, validation, usage metering, quotas, and authentication.

- Collaborate closely with other teams to deliver robust, developer-friendly model serving experiences.

- 3+ years experience building and operating distributed systems or large-scale APIs.

- Proven track record of owning low-latency, reliable backend services (rate-limiting, auth, quotas, metering, migrations).

- Infra instincts with performance sensibilities: profiling, tracing, capacity planning, and SLO management.

- Comfortable debugging complex systems, from runtime internals to GPU execution traces.

- Strong written communication; able to produce clear design docs and collaborate across functions.

- Experience with LLM runtimes (vLLM, SGLang, TensorRT-LLM) or contributions to open-source inference engines (vLLM, TensorRT-LLM, SGLang, TGI)

- Knowledge of Kubernetes, service meshes, API gateways, or distributed scheduling.

- Background in developer-facing infrastructure or open-source APIs.

- We value infra-leaning generalists who bring strong engineering fundamentals and curiosity. ML experience is a plus, but not required.

More jobs at Baseten

Similar jobs near San Francisco (Remote)

Tell me when more Software Engineer, Private Computing jobs post near San Francisco (Remote) We re-check every listing against the employer’s own board — no résumé needed.

Search Software Engineer - Model Products jobs near San Francisco (Remote) → Browse all live jobs

This posting was published by Baseten on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.