- Finetuning small language models - Improving the quality of existing data using scalable approaches. Examples include: making sure URLs are associated the right company, we have the correct HQ address, we have mapped parents-subsidiary using techniques like LLM validation, SERP, and triangulating across sources. - Adding new signals: this usually involves scrubbing, matching and normalizing new signals and matching to our existing ontology - Pushing solutions into production environments, which may involve touching data pipelines and/or backend systems
- Located within Americas timezones
- ML/Data: PyTorch, Huggingface, Gemma models, LORA, VLLM, Skypilot, Marimo - Languages & Frameworks: Python, FastAPI, React, Typescript - Cloud Platform: Google Cloud Platform (GCP) - Databases: PostgreSQL, DuckDB - Infrastructure: Cloud Run - Product/Design: Figma, Vercel V0
- Transforming noisy datasets into high-quality data products - Running expensive analytics computations efficiently - Managing the complexity of a growing number of data sources, machine learning models, and large data operations - Create a great PLG experience with upsell pathways
- Medical, dental, and vision (US) - 401k (US) - Target 4 weeks PTO
Search Data Scientist/Machine Learning Engineer jobs near Remote → Browse all live jobs
This posting was published by Sumble Inc on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.