Senior Principal AI Architect/Engineer Job Locations US-TX-Plano ID 2026-451268 Category IT Position Type Salaried Relocation No Overview The AI Observability Architect is a senior technical leader responsible for designing, deploying, and operating an enterprise-grade, production-ready AI observability platform that spans the full spectrum of modern agentic AI — from large language model (LLM) workflows and multi-agent orchestration to physical AI systems, reinforcement learning harnesses, multi-modal pipelines, and agentic marketplaces.
This role serves as the strategic and engineering authority for end-to-end telemetry, tracing, safety, and quality signals across heterogeneous agent frameworks and platforms. The architect leads the convergence of AI observability with safety & security (including red teaming), Responsible AI (RAI), data science, physical AI, memory/skills engineering, agent fleet management, self-evolving harnesses, reinforcement learning, agent-to-agent protocols (A2A, UCP, AP2), and continuous quality engineering — making this a uniquely broad and high-impact role within the AI Solutions & Platforms organization.
The role also owns OpenTelemetry (OTEL) integration across third-party agentic platforms (Salesforce AgentForce, ServiceNow, Microsoft Agent 365, and others), enabling unified observability and governance at enterprise scale.
Responsibilities Agentic AI Observability Architecture at Scale (30%) Define and own the enterprise observability architecture for AI agents, LLMs, multi-agent workflows, and physical AI systems — covering planner/executor loops, tool/function calls, RAG retrieval chains, and memory/state transitions. Build and operate unified telemetry pipelines incorporating metrics, logs, distributed traces, semantic/vector signals, and real-time event streaming (Kafka) at enterprise scale.
Instrument OpenTelemetry (OTEL) across heterogeneous platforms including Salesforce AgentForce, ServiceNow, Microsoft Agent 365, and internal frameworks — delivering protocol-level observability for agent ecosystems including MCP, A2A, UCP, and AP2. Design and implement observability for Agent Fleets, multi-modal pipelines, physical AI systems, and self-evolving reinforcement learning harnesses — including signal capture for reward shaping and policy evaluation.
Deliver dashboards, alerting, SLO/SLA management, incident runbook automation, and RCA tooling that drive measurable reliability improvements and reduce MTTR across agentic services. Establish cost telemetry and FinOps observability for AI workloads — token consumption, inference cost allocation, and GPU/compute efficiency across cloud environments (Azure, AWS, GCP).
Safety, Security & Red Teaming (15%) Lead observability-driven red team exercises targeting agentic AI systems — instrumenting attack surfaces, adversarial prompt injection vectors, model evasion attempts, and multi-agent trust boundary failures. Design telemetry pipelines that capture safety-critical signals: guardrail trigger rates, policy violation events, PII exposure risks, prompt leakage, and agent hallucination rates.
Partner with Security and RAI teams to embed threat modeling, zero-trust agent authentication, and behavioral anomaly detection into the observability platform. Instrument secure policy enforcement layers across agent-to-agent communication protocols (A2A, UCP, AP2) and maintain audit-ready traceability for all AI decision events.
Develop and maintain a Security Observability Playbook covering incident classification, escalation paths, and forensic trace retention policies for agentic AI systems. Responsible AI (RAI) & Governance (10%) Integrate RAI signal capture — fairness, bias detection, explainability, and safety metrics — directly into observability pipelines, making compliance measurable and audit-ready.
Deliver governance dashboards that surface RAI compliance posture across all active AI agents and LLM deployments, aligned with global regulatory standards. Support risk assessments, gap analyses, and governance frameworks with real-time observability insights — enabling proactive risk mitigation rather than reactive audit responses.
Collaborate with RAI CoE and Legal/Compliance teams to define data retention, consent logging, and model decision traceability standards embedded in the telemetry architecture. Quality Engineering for Agentic Solutions — Post Go-Live & Continuous QE (10%) Own the Continuous Quality Engineering (CQE) framework for post-production agentic solutions — defining and tracking quality metrics across accuracy, latency, agent success rate, tool-call fidelity, and user outcome measures.
Build automated quality gates within CI/CD pipelines that leverage observability data to detect regressions, drift, and degradation in agent performance — preventing silent failures in production. Instrument and monitor Skill Evaluations (evals) across the Memory, Skills, and MCP harness stack — providing traceability from eval results to production behavior.
Partner with product and business stakeholders to define SLA-backed quality benchmarks and deliver automated alerting when quality thresholds are breached. Drive root-cause analysis for quality failures using distributed trace data, enabling rapid iteration and continuous improvement cycles for agentic solutions. Memory, Skills, MCP & Harness Engineering Observability (10%) Design and implement observability for the agent memory layer — episodic, semantic, and working memory read/write operations — providing latency, accuracy, and drift monitoring across memory backends.
Instrument MCP (Model Context Protocol) server interactions, tool registrations, skill invocations, and context injection pipelines with full trace propagation and semantic tagging. Own observability for self-evolving harness and reinforcement learning (RL) systems — capturing reward signals, policy update events, environment state transitions, and learning convergence metrics.
Monitor harness execution fidelity, skill eval pass/fail rates, and regression signals across training, fine-tuning, and inference workflows — feeding data back into the quality engineering loop. Data Science Observability & Hardcore Python Engineering (5%) Lead a team of senior Python engineers building high-performance, production-grade observability tooling — including custom OTEL exporters, semantic trace enrichers, signal aggregators, and anomaly detection pipelines.
Apply data science methods — statistical process control, time-series anomaly detection, clustering, and causal inference — to transform raw telemetry into actionable AI operational intelligence. Build and maintain Python-native SDKs and libraries that simplify observability onboarding for agent developers across the organization.
Establish code quality standards, testing frameworks, and peer review practices for the observability engineering team — embedding software craftsmanship into the team culture. Agentic Marketplace, Registry & Ecosystem Observability (5%) Instrument the Agentic Marketplace and Agent Registry platforms — providing usage telemetry, adoption metrics, capability health scores, and dependency mapping for registered agents and skills.
Design observability APIs and SDK hooks that allow marketplace-registered agents to self-report health, performance, and behavioral signals into the central observability platform. Monitor inter-agent communication patterns across the marketplace ecosystem — identifying latency hotspots, circular dependencies, and protocol mismatches in agent-to-agent (A2A) workflows.
Deliver a Marketplace Observability Dashboard surfacing agent catalog health, adoption trends, quality scores, and incident history — supporting marketplace governance and curation decisions. Integration, Deployment & CI/CD Automation (5%) Build and maintain CI/CD pipelines for observability services and agent operations center components, incorporating automated testing, deployment gates, and rol
Search Senior Principal AI Architect/Engineer jobs near Plano, TX → Browse all live jobs
This posting was published by PepsiCo on their own careers system and is shown here with a direct link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.