Reflectionai

Member of Technical Staff - Distributed Systems Engineer

Full-time · New York, NY
✓ Verified live on the employer's own system · added 150 days ago
Save search
Senior

Skills & tools

DevopsManagementOperationsDistributed SystemsProgrammingTroubleshootingCommunicationsMachine Learning
Apply on company site ↗ See your fit → free

Full job description

Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all.

Build and operate a company-wide foundations platform that accelerates every team by providing reliable, scalable developer infrastructure, SRE capabilities, and high-throughput data ingestion tooling enabling Reflection to move faster as we scale.

Build and operate the core shared services that power our research, training, and production environments. These systems form the foundational platform that multiple teams depend on for model development, deployment, and evaluation, unifying data, compute, and workflow management across the stack while enabling rapid experimentation and reliable production systems.

- Build and operate shared services that multiple teams rely on across research and production workflows.

- Define and uphold reliability targets through SLIs, SLOs, and healthy on-call practices.

- Maintain strong operational readiness with runbooks, incident playbooks, and capacity planning.

- Ensure correctness and performance under load, addressing consistency, tail latency, and failure modes.

- Develop APIs, SDKs, and internal platforms that enable high-velocity experimentation and iteration.

- Reduce operational burden through better tooling, standardization, and platform patterns that scale across teams.

- Container Abstractions: Containers-as-a-Service, Kubernetes abstraction layers, container orchestration, reproducible environments, multi-tenant isolation.

- Distributed Systems Architecture: Sharding, replication, coordination services, high-concurrency systems, concurrency control.

- Reliability & Performance: Idempotency, retries, backpressure, SLI/SLO design, tail latency optimization, service reliability engineering.

- Strong software engineering background with experience shipping production-grade systems.

- Experience designing APIs, services, or developer platforms that handle large-scale data or compute.

- Comfortable navigating complex codebases, debugging hard problems, and optimizing for reliability and speed.

- Thrive in a high-agency, fast-paced startup environment; bias toward action and impact.

- Excited about zero to one challenges, building new systems rather than maintaining legacy ones.

- Collaborative, clear communicator, and comfortable working across research and infra boundaries.

- Motivated by creating the software backbone for the world's most capable open-weight AI systems.

More jobs at Reflectionai

Similar jobs near New York, NY

Tell me when more Member of Technical Staff - Distributed Systems Engineer jobs post near New York, NY We re-check every listing against the employer’s own board — no résumé needed.

Search Member of Technical Staff - Distributed Systems Engineer jobs near New York, NY → Browse all live jobs

This posting was published by Reflectionai on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.