Education: Doctorate
We are looking for a Principal GenAI Data Engineer to join our team. This is a Hybrid role based in San Jose, CA or Bellevue, WA (3 days in office), reporting to the Senior Manager, Enterprise AI Data Platform in the IT Data Strategy department. We are seeking an experienced technical leader to drive the design and implementation of enterprise-grade Generative AI data ingestion, knowledge preparation, and platform architectures that enable scalable, production-ready GenAI applications.
This role focuses on architecting robust pipelines and platforms for ingesting, processing, governing, and serving structured and unstructured enterprise data for AI/LLM workloads. The ideal candidate combines deep expertise in enterprise data architecture, unstructured data pipelines, GenAI platform engineering, and strong software engineering skills in Python.
Architect enterprise-scale GenAI data platforms for ingestion, transformation, enrichment, and serving of structured and unstructured data
Design scalable pipelines for enterprise knowledge ingestion from diverse data sources including documents, SaaS platforms, knowledge bases, collaboration tools, and databases
Define architecture for metadata extraction, chunking, enrichment, embeddings generation, and knowledge preparation workflows
Design AI-ready data models and storage strategies for vector, graph, and hybrid knowledge systems
Architect scalable unstructured data processing pipelines for text, images, PDFs, tables, and multimodal content
You thrive in ambiguity. You're comfortable building the path as you walk it. You thrive in a dynamic environment, seeing ambiguity not as a hindrance, but as the raw material to build something meaningful.
You act like an owner. Your passion for the mission fuels your bias for action. You operate with integrity because you genuinely care about the outcome.
True ownership involves leveraging dynamic range: the ability to navigate seamlessly between high-level strategy and hands-on execution.
You are a problem-solver. You love running towards the challenges because you are laser-focused on finding the solution, knowing that solving the hard problems delivers the biggest impact.
You are a high-trust collaborator. You are ambitious for the team, not just yourself. You embrace our challenge culture by giving and receiving ongoing feedback—knowing that candor delivered with clarity and respect is the truest form of teamwork and the fastest way to earn trust.
You are a learner. You have a true growth mindset and are obsessed with your own development, actively seeking feedback to become a better partner and a stronger teammate. You love what you do and you do it with purpose.
Foundational understanding of AI/ML technologies and experience leveraging, securing, or positioning AI-driven solutions to optimize outcomes within your functional domain
Experience building distributed/scalable data pipelines for AI workloads
Strong understanding of unstructured data extraction and processing pipelines
Experience with vector databases, graph databases, and metadata/knowledge storage systems
Hands-on experience with clustering, entity recognition algorithms, and modern retrieval strategies (including RAG, search, and agentic AI workflows)
Advanced experience architecting real-time distributed vector search infrastructure and multi-modal knowledge graph pipelines for enterprise-grade Retrieval-Augmented Generation (RAG) applications
Experience with LLMOps / GenAIOps frameworks such as LangSmith, Evaluation Framework like Arize Phoenix, Weights & Biases, or MLflow
Familiarity with Agent Frameworks like LangGraph, CrewAI, or Google ADK
Search Principal GenAI Data Engineer jobs near Bellevue, WA +2 more → Browse all live jobs
This posting was published by Zscaler on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.