Skip to content
Back to jobs

[Remote] Executive Director, AI Infrastructure & Platform Engineering

frontendnode-production.up.railway.app
Anywhere
Posted Jun 20, 2026
remotefull-time

Key details

Work type
remote
Employment
full time

Job Description

Note: The job is a remote job and is open to candidates in USA. CVS Health is dedicated to shaping a more connected and compassionate health experience.

They are seeking an Executive Director for AI Infrastructure & Platform Engineering, responsible for leading the development and operational excellence of their AI compute platform, ensuring high availability and reliability for frontier AI workloads.

Responsibilities Define and execute the long-range vision and strategy for AI infrastructure and platform engineering, with availability (>99.99%), reliability, and platform performance as the primary measures of successRecruit, hire, develop, and retain a high-performing engineering organization spanning infrastructure, network, platform reliability, observability, security, 24/7 operations, change and release management, and FinOpsEstablish clear ownership, accountability, and performance expectations across all functional teams; foster a culture of operational excellence, engineering rigor, and continuous improvementProvide executive-level communication to senior leadership on platform status, milestones, risk posture, and strategic initiativesOwn the physical layer of the AI compute environment — GPU compute, storage, network fabric, capacity planning, and hardware lifecycle accountabilityDirect bare-metal Kubernetes and OpenShift operations, including cluster administration, GPU quota governance, infrastructure-as-code adoption, and availability baseline enforcementGovern high-performance network fabric operations — RoCE v2, spine-leaf topology, lossless Ethernet tuning, congestion management, and segmentationEstablish and enforce operational baselines across every layer of the stack — hardware, fabric, platform, and workload — with deviations detected, escalated, and resolved within defined SLAsDirect Innovation POD strategy to develop self-healing and autonomous capabilities that proactively prevent service degradation before it impacts availabilityBuild and sustain a high-performing 24/7 operations model — designed for sustainable, predictable coverage with no mandatory overtime and measurable team health and retentionDrive end-to-end observability across the physical and platform layers, with continuous feedback loops connecting monitoring data to incident response, change decisions, and improvement cyclesOversee change management so every modification is risk-assessed, monitored during rollout, and baseline-validated post-deploymentEnsure configuration consistency and drift detection across all platform components to prevent baseline degradation over timeLead GPU FinOps governance — utilization optimization, tenant quota enforcement, and cost reduction — in partnership with the Finance organizationEmpower the Security SRE Lead to maintain a world-class security posture across the infrastructure and platform layers, with robust compliance to frameworks including HIPAA and NIST AI RMFGovern access controls, audit logging, vulnerability management, and network segmentation across the AI compute environmentLead the operational transition from program-launch staffing to permanent CVS-owned operations — governing phased handoffs, competency validation, and milestone sign-offs to ensure minimal disruption to platform availability and business operationsEstablish and lead the long-term operating model by institutionalizing key technical, architectural, and delivery leadership capabilities into permanent CVS roles, ensuring the organization is fully self-sustaining at program closeOwn vendor relationships, contract performance, and accountability across the hardware, networking, platform, and managed-services stackManage budget ownership for the AI infrastructure and platform engineering organization, including capital planning and operational expense governance Skills 10+ years of engineering leadership experience, with substantial time directly owning physical infrastructure at data center scale — including hardware lifecycle, capacity planning, and facility coordination (power, cooling, rack-and-stack execution)Hands-on production ownership of bare-metal Kubernetes or OpenShift.

Managed cloud services (EKS, GKE, AKS) alone do not substitute for the practitioner expertise this role requiresFluency with high-speed cluster fabrics — RoCE v2, InfiniBand, EVPN-VXLAN, or carrier-grade equivalent — and the operational discipline these fabrics require (PFC, ECN, lossless tuning, congestion management)5+ years leading multiple technical teams simultaneously, including 24/7 operations organizations, with measurable team health, retention, and performance outcomesProven success establishing and enforcing operational baselines, SLO / SLI / error-budget frameworks,

Company & context

Evidence is labeled so you can tell internal community data from public sources.

Context may refresh in the background.

Trust-check this listing

Verify scam risk and ghost-job signals before you apply.

Related roles

Browse more remote Software Engineer jobs.

Senior QA Automation Engineer | Remote

frontendnode-production.up.railway.app
remoteAnywherefull-time
Posted Jun 22, 2026
View details

Compliance Automation Engineer Remote - United States

frontendnode-production.up.railway.app
remoteAnywherefull-time
Posted Jun 22, 2026
View details

Senior QA Automation Engineer | Remote

frontendnode-production.up.railway.app
remoteAnywherefull-time
Posted Jun 22, 2026
View details

Compliance Automation Engineer Remote - United States

frontendnode-production.up.railway.app
remoteAnywherefull-time
Posted Jun 22, 2026
View details

Source: Google Jobs • Last updated 3w ago