Education: Bachelor's degree or related field
Experience: 2+ years
We are seeking a highly skilled Core ML Engineer to design, develop, and optimize machine learning systems that power next-generation AI platforms and applications. This role focuses on model development, inference optimization, and scalable ML infrastructure, enabling production-grade AI capabilities across enterprise systems.
The ideal candidate combines strong software engineering fundamentals with deep ML expertise, and thrives in building robust, high-performance systems at scale.
* Design and implement machine learning models and pipelines for production use
* Build scalable training → evaluation → deployment workflows
* Optimize model inference for latency, throughput, and cost
* Implement advanced techniques such as caching, quantization, batching, and routing
* Benchmark and profile models across diverse workloads and hardware environments
* Integrate ML/LLM models into APIs, microservices, and applications
* Build and maintain model-serving infrastructure (e.g., vLLM, ONNX, custom runtimes)
* Collaborate with platform and infrastructure teams for scalable deployment
* Design data pipelines for ingestion, preprocessing, feature engineering, and validation
* Improve data quality and model reliability through systematic evaluation
* Partner with product, platform, and hardware teams to deliver end-to-end ML solutions
* Participate in design reviews and contribute to system architecture decisions
- Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 2+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
Master's degree in Computer Science, Engineering, Information Systems, or related field and 1+ year of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.
PhD in Computer Science, Engineering, Information Systems, or related field.
Strong programming skills in Python and at least one systems language (C++/Rust/Go)
* Machine learning fundamentals (supervised, unsupervised, deep learning)
* Large Language Models (LLMs), multimodal models, or generative AI
* Distributed computing and GPU/accelerator environments including model serving and efficient cache/state management (e.g. KV cache, embeddings) across disaggregated systems
* Agentic and multi-step AI workflows, tool integration, orchestration, and multi-component pipelines
* Model optimization techniques (quantization, distillation, caching)
* Vector databases and search systems (OpenSearch, Qdrant, etc.)
* Cost-aware system design - model routing (small vs. large models), dynamic batching, and caching strategies
Search Senior Engineer - Machine Learning jobs near San Diego, CA → Browse all live jobs
This posting was published by Qualcomm on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.