If you are a software engineering leader ready to take the reins and drive impact, we’ve got an opportunity just for you. As a Director of Software Engineering
Payment Data Platform Infrastructure at JPMorgan Chase within the Commercial and Investment Banking
Data Analytics Payment Team , you own the infrastructure, reliability, and security posture underpinning the BRIE data platform and the NEO agent runtime. You set technical direction across infrastructure engineering, site reliability, and security operations, and you are accountable for these mission-critical systems running securely, stably, and at scale across a multi-region AWS and on-premises estate where the stakes and regulatory bar are high. Job responsibilities
Provides overall direction, oversight, and coaching for a team of engineering managers and senior technologists spanning infrastructure engineering, SRE, and SecOps for both BRIE and NEO
Owns the reliability strategy for the platform estate
SLOs, error budgets, capacity planning, and disaster recovery — sustaining active-active, multi-region operation across AWS and on-prem while meeting the platform’s high-availability commitments
Sets the security operations agenda: threat detection and response, vulnerability and patch management, secrets and key management, and evidence for audit, risk, and regulatory reviews, maintaining CPOF compliance across all environments
Champions infrastructure-as-code, immutable deployments, and platform automation so provisioning, scaling, and remediation are repeatable, reviewable, and auditable
Sets direction and governance for agentic AI-enabled engineering and SDLC/TLM automation within a technical area to drive measurable improvements in speed, quality, and operational outcomes (e.g., AI-orchestrated delivery workflows, release readiness controls, automated test modernization, and incident triage acceleration), while establishing guardrails for validation, security, resiliency, traceability, and reuse across teams.
Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation and support capacity unlock initiatives at scale.
Leads incident command for major events, drives blameless postmortems, and closes the loop on systemic remediation to raise overall operational stability
Anticipates the infrastructure needs of the data platform and agent runtime — compute, storage, networking, and GPU serving/training capacity — and translates them into a funded, prioritized roadmap
Makes decisions that influence resourcing, budget, tooling, and vendor selection across the infrastructure organization, and is accountable for those outcomes
Leads evaluation sessions with cloud providers, vendors, startups, and internal teams to probe architectural designs, technical credentials, and applicability within existing systems and information architecture
Partners with platform, product, data, and InfoSec stakeholders, and communicates reliability, security, and cost trade-offs to senior leadership Required qualifications, capabilities, and skills
Formal training or certification on software engineering concepts and 8+ years applied experience, including significant time leading infrastructure, SRE, or platform organizations, with experience managing managers
Experience leading teams of technologists and managing budget, resourcing, and delivery across multiple concurrent workstreams
Hands-on background in large-scale distributed systems: system design, application development, testing, and operational stability
Deep expertise operating production systems across public cloud and on-premises data centers, including multi-region, active-active resilience and disaster recovery
Demonstrated ownership of a security posture in a regulated environment
SecOps, identity and access, secrets management, audit, and regulatory compliance
Experience leading adoption of agentic AI-enabled engineering practices (using enterprise-authorized tools within the work environment) across teams, including defining operating expectations (human-in-the-loop validation, quality gates), measuring outcomes, and ensuring secure handling of sensitive inputs/outputs.
Strong understanding of responsible AI use and control expectations in engineering workflows, including data sensitivity, resiliency/security implications, and governance; ability to influence leaders on safe scaling patterns and reuse.
Proficient in all aspects of the Software Development Life Cycle
Advanced understanding of agile methodologies such as CI/CD, Application Resiliency, and Security
Advanced in one or more programming language(s), with the ability to guide critical technical decisions
Demonstrated proficiency in software applications and technical processes within a technical discipline (e.g., cloud, artificial intelligence, machine learning, distributed data systems)
Preferred qualifications, capabilities, and skills
Deep AWS experience across multi-account, multi-region architectures, paired with on-premises data center operations
Expert-level Kubernetes and containerized platform operations, including multi-tenant isolation (siloed and pooled) and noisy-neighbor controls
Proficiency with Infrastructure as Code (Terraform) and GitOps-driven, immutable deployment pipelines
Experience operating policy and authorization at the infrastructure layer (e.g., OPA/Rego, OpenFGA) and policy-as-code workflows and experience operating Databricks (Spark) and real-time stream processing with Apache Flink at production scale
Strong observability and SRE tooling background (e.g., Splunk, OpenTelemetry, Prometheus/Grafana) with SLO/error-budget practice
Experience running GPU compute fleets for ML serving and training (e.g., NVIDIA L40S/H100, SageMaker) and associated capacity and cost governance and experience with regulated-environment compliance frameworks and evidencing controls to risk, audit, and InfoSec
Familiarity building and running Java Spring Boot GraphQL services for high-availability, low-latency APIs and familiarity with the data platform stack — lakehouse and open table formats (Apache Iceberg), streaming and batch pipelines, and OLAP/columnar stores
This posting was published by JPMorgan Chase on their own careers system and is shown here with a direct
link to apply there. Employers: for corrections or removal, contact jobs@veritahire.com.