SimpliSafe

Staff Software Engineer, ML Infrastructure

$147K–$215KFull-time · Boston, MA
✓ Verified live on the employer's own system · added 81 days ago
Save search
Senior · 8+ yrs exp

Requirements

Experience: 8+ years

Skills & tools

SecurityDistributed SystemsTeam LeadershipMachine LearningDevopsCode ReviewManagementRecordkeeping

Benefits — mentioned in this posting

Remote / flexibleBonus / commission
Apply on company site ↗ See your fit → free

Full job description

We’re a high-tech home security company that’s passionate about protecting the life you’ve built and our mission of keeping Every Home Secure. And we’ve created a culture here that cares just as deeply about the career you’re building. Ours is a no ego culture of collaboration and innovation where those seeking their next challenge can find big opportunities and make a huge impact on the lives of all those who we protect.

We don’t just want you to work here. We want you to grow and thrive here.

We’re embracing a hybrid work model that enables our teams to split their time between office and home. Hybrid for us means we expect our teams to come together in our state-of-the-art office on two core days, typically Tuesday, Wednesday, or Thursday – working together in person and choosing where they work for the remainder of the week.

We all benefit from flexibility and get to use the best of both worlds to get our work done.

Well, we’re growing and thriving. So, we need smart, talented, and humble people who share our values to join us as we disrupt the home security space and relentlessly pursue our mission of keeping Every Home Secure.

We're looking for a Staff Software Engineer to join our Cloud ML team — the team that owns both the cloud-side ML infrastructure and the applied ML research that powers SimpliSafe's intelligent home security products. This is a senior individual contributor role for a distributed systems expert who wants to apply that craft to one of the most demanding problem domains in the company.

You'll partner closely with other Staff and Principal engineers to drive architecture, mentor across the team, and set the technical direction for our ML platform. The work spans two of our most demanding workloads: real-time computer vision inference that processes video from cameras and doorbells across our customer base, and LLM/GenAI infrastructure that will power our future generation of intelligent applications.

Both are, fundamentally, distributed systems problems — high-throughput, low-latency, multi-tenant, GPU-aware, and unforgiving of regressions.

This role is for someone who has built and operated large-scale distributed services in production — high-QPS APIs, real-time platforms, low-latency serving systems — and is excited to bring that depth to ML infrastructure. Prior ML experience is a plus, not a prerequisite. If you've shipped systems that serve a lot of traffic, scale gracefully, and stay up at 3am, we want to talk to you.

  • Drive architecture decisions for our Kubernetes-based ML platform — anchored on Ray for inference, alongside KServe, Triton, and vLLM — across real-time and batch workloads.
  • Lead deep technical reviews on system design, capacity planning, and reliability for the highest-stakes ML systems at SimpliSafe.
  • Identify and remove the systemic bottlenecks in our ML deployment infrastructure — whether that's serving reliability, deployment friction, observability gaps, scaling, or cost.
  • Own the design and evolution of cloud-side inference systems that process live video and events from SimpliSafe devices in real time.
  • Drive throughput, latency, and cost improvements (batching strategies, GPU utilization, autoscaling, multi-model serving) for production CV models.
  • Build the feedback loops between cloud inference, edge devices, and the data flywheel that improves model quality over time.
  • Help shape how SimpliSafe serves LLMs in production — model serving patterns, KV-cache and batching strategies, evaluation pipelines, guardrails, and cost controls.
  • Partner with applied ML engineers to take new GenAI-powered product features from prototype to scaled deployment.
  • Mentor engineers across the team through design reviews, code reviews, pairing, and written guidance — a meaningful uplift on everyone you work with.
  • Establish and evangelize best practices for model lifecycle management (registry, deployment, monitoring, rollback, drift) and on-call.
  • Write the documentation, runbooks, and architectural decision records that make the platform legible and durable.
  • Lead incident response and postmortems for critical ML systems; turn lessons learned into platform-level improvements.
  • Define SLOs, observability standards, and on-call practices for ML services in production.
  • 8+ years of software engineering experience, with a clear track record of building and operating large-scale distributed systems in production.
  • Deep expertise in high-throughput, low-latency services — ad serving, recommendations, real-time APIs, online platforms, or similar — including the operational reality of running them at scale.
  • Strong production experience on Kubernetes and AWS (EKS, S3, IAM, networking) and with Kafka, containerized deployments, CI/CD, and infrastructure-as-code.
  • Demonstrated experience with the building blocks of high-scale systems: load balancing, autoscaling, batching, caching, multi-tenancy, queuing, and capacity planning.
  • Proficiency in Python is required; experience with a systems language (Go, C++, Rust) for performance-sensitive components is a plus.
  • Staff-level technical leadership : ability to drive ambiguous, cross-cutting initiatives, align senior stakeholders, and elevate the engineers around you without formal authority.
  • Strong written and verbal communication — you can make complex technical tradeoffs legible to ML scientists, product, and other infra teams.
  • ML exposure is preferred — having deployed or operated production ML systems, worked closely with ML teams, or built ML-adjacent infrastructure. Exceptional distributed systems engineers without direct ML experience are encouraged to apply; we'll help you ramp.
  • Hands-on experience with Ray , KServe , Triton , vLLM , or other ML serving stacks.
  • Hands-on experience with LLM serving in production (vLLM, TGI, TensorRT-LLM, SGLang) — KV cache management, continuous batching, speculative decoding, quantization for serving.
  • Experience building real-time video or streaming pipelines (Kafka, Kinesis, Flink, or similar) at scale.
  • Experience operating GPU-based inference systems — GPU-aware scheduling, multi-model serving, accelerator utilization optimization.
  • Familiarity with ML fundamentals — how models are trained, evaluated, versioned, deployed, monitored, and rolled back in production.
  • Experience with model lifecycle tooling (MLflow, Weights & Biases, model registries, drift detection, shadow deployments).
  • Open source contributions to distributed systems or ML infrastructure projects.
  • Experience operating in environments with strong security and compliance requirements .

The Cloud ML team owns the full surface area — infrastructure and applied research — which means your work as a Staff infra engineer directly shapes what's possible for the science. You'll have unusual leverage: the platform you build determines how fast SimpliSafe can ship intelligent features, and the features we ship directly impact whether someone's home is safer tonight than it was yesterday.

The target annual base pay range for this role is $146,600 to $215,100.

Beyond base pay, we offer a Total Rewards package that may include participation in our annual bonus program, equity, and other forms of compensation, in addition to a full range of medical, retirement, and lifestyle benefits. More details can be found here .

We wholeheartedly embrace and actively seek applications from all individuals, no matter how they identify. We are committed to cultivating a diverse and inclusive workplace, and we believe our work is enriched when we incorporate a multitude of perspectives, backgrounds, and experiences. We want everyone who works here to thrive and contribute to not only our mission of keeping every home secure, but also to making our workplace safe and supportive for others.

If a reasonable accommodation may be needed to fully participate in the job application or interview process, to perform the essential functions of a position, or to receive other benefits and privileges of employment, please contact careers@simplisafe.com .

More jobs at SimpliSafe

Similar jobs near Boston, MA

Tell me when more Staff Software Engineer, Maritime jobs post near Boston, MA We re-check every listing against the employer’s own board — no résumé needed.

Search Staff Software Engineer, ML Infrastructure jobs near Boston, MA → Browse all live jobs

This posting was published by SimpliSafe on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.