Senior SDET remote from NYC, Seattle or the SF Bay Area
Key details
- Work type
- remote
- Employment
- full time
Job Description
Senior SDET — Quality & Reliability Engineering About the Role We're hiring a Senior SDET to help drive the next evolution of quality engineering, combining automated functional validation, production health monitoring, and AI-driven failure analysis.
This role focuses on ensuring our Disaster Recovery (DR) and production environments are not just available, but fully functional, continuously validated, and increasingly capable of self-diagnosis.
The ideal candidate brings strong automation skills, systems thinking, and a genuine passion for improving reliability across complex distributed systems.
What You'll Do You'll design and build automated quality and validation systems that strengthen confidence in both production and disaster recovery readiness.
Partnering closely with QA, SRE, and Engineering, you'll validate critical business workflows, improve observability, reduce alert noise, and accelerate incident detection and resolution.
This is a role at the intersection of test automation, production reliability, and intelligent diagnostics, helping advance the shift from traditional QA toward more autonomous, resilient quality systems.
Responsibilities Design and implement automated functional health checks for DR and production environments using synthetic transactions and API validation Build continuous validation pipelines that verify end-to-end business workflows such as authentication, transactions, and system integrations Develop intelligent alerting based on functional failures and customer-impacting behavior, not solely infrastructure metrics Integrate observability signals (logs, metrics, traces) with automated test frameworks to improve system visibility and diagnosis Develop AI/ML-driven approaches to detect failure patterns, correlate issues across services, and identify probable root causes Build systems that recommend or trigger automated remediation to support early-stage self-healing Partner cross-functionally with QA, SRE, and Engineering to improve reliability, incident response, and recovery readiness Define, measure, and report on functional SLAs, service health indicators, and quality metrics Contribute to disaster recovery drills, readiness exercises, and automated validation efforts What You Bring 5+ years in QA Automation, SDET, Software Engineering, or a related technical discipline Strong experience building and maintaining automated test frameworks (Playwright, Jest, SuperTest, REST API testing tools) Experience in cloud environments, preferably AWS Familiarity with observability platforms such as Datadog, New Relic, CloudWatch, or Splunk Strong programming skills in TypeScript, Python, Java, or similar Experience designing end-to-end test strategies for distributed systems and production-like environments Strong problem-solving skills across application, infrastructure, and workflow layers Nice to Have Synthetic monitoring, production validation, or proactive health-checking systems AI/ML techniques for anomaly detection, log analysis, or failure correlation CI/CD pipelines, release automation, and validation gates Understanding of microservices architecture and distributed system failure modes SRE concepts, SLIs, SLOs, error budgets, production-readiness practices
Company & context
Evidence is labeled so you can tell internal community data from public sources.
Context may refresh in the background.
Trust-check this listing
Verify scam risk and ghost-job signals before you apply.
Related roles
Senior SDET
Senior SDET - GovTech Industry (Colombia)
QA Automation Lead [gn] Data Intelligence
Source: Google Jobs • Last updated 2w ago