Our ultimate goal is to build perfect search over all the world's information, far beyond Google. If you want to build massive-scale ML systems that will define the way the new AI world consumes information, this is the place for you.
As a Data Engineer, you'll architect and build the data infrastructure that powers everything we do—from crawling billions of pages to training our embedding models to serving real-time search. You'll have enormous autonomy in designing systems that scale to hundreds of petabytes. If you've ever wanted to build data pipelines at a scale that most companies only dream about, this is your chance.
- Deep understanding of lakehouse architectures (Delta Lake, Iceberg, Hudi) and when to use them
- Experience building and operating large-scale distributed data processing pipelines
- Hands-on experience with streaming data systems (Kafka, Flink, or similar)
- Familiarity with Ray, Spark, or ClickHouse at production scale
- An obsessive focus on reliability and building systems that don't page you at 3am
- Experience with Lance or other vector-native storage formats
- Background in GPU-accelerated data processing (RAPIDS, cuDF)
- Design a lakehouse architecture that handles 100+ PB of web crawl data
- Build streaming pipelines that process billions of documents per day for real-time indexing
- Architect the data layer for our embedding training infrastructure on Ray
- Scale our ClickHouse deployment to handle analytical queries across petabytes of search logs
Search Software Engineer, Distributed Data Systems jobs near San Francisco, CA → Browse all live jobs
This posting was published by Exa on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.