Skip to content
Back to jobs

Staff Engineer, Platform Engineering – Operational Health

AudioEye
Anywhere
Posted May 27, 2026
remotefull-time

Key details

Work type
remote
Employment
full time

Job Description

Job Description: Conduct comprehensive audits of infrastructure, deployment processes, incident patterns, and on-call burden Identify foundational issues causing operational pain: fragile systems, deployment friction, poor observability, and architectural weaknesses Establish baseline metrics for system health and operational efficiency Prioritize improvements systematically based on impact to on-call burden and reliability Design and implement solutions that address the root causes of incidents Eliminate single points of failure in critical paths Implement patterns for graceful degradation and rapid recovery Build comprehensive observability, logging, metrics, and tracing infrastructure Identify and automate repetitive manual work that burdens operational staff Establish organizational standards and expectations for reliability Design and maintain runbooks, playbooks, and incident response processes Create feedback loops from incidents to systemic improvements Support CI/CD pipelines that enable safe, frequent deployments Develop operational tooling (dashboards, alerts, automation) Reduce friction in how engineers interact with infrastructure and deploy code Scale infrastructure responsibly as the organization grows Optimize platform reliability, performance, and cost through capacity planning, workload tuning, and architectural tradeoffs Own the technical strategy and roadmap for infrastructure, reliability, and operational posture Mentor engineers across the organization on operational thinking and reliability engineering Lead architecture reviews and establish technical standards Build organizational practices and knowledge that outlast any single engineerRequirements: 5–8+ years of software engineering experience with demonstrated expertise in infrastructure, operations, reliability engineering, or SRE Staff-level technical depth with ownership of significant platform decisions at scale Proven ability to manage complex trade-offs and defend technical direction Experience carrying the pager; understanding of on-call burden from lived experience Demonstrated pattern recognition; ability to design systemic solutions rather than one-off fixes Expert-level infrastructure-as-code (Terraform); comfortable with complex multi-environment deployments Docker and container orchestration expertise (Kubernetes or similar); experience designing or improving deployment pipelines Deep familiarity with Node.js and TypeScript in production environments Substantial AWS experience (EC2, RDS, Lambda, CloudWatch, VPC, load balancers, auto-scaling, managed services) Understanding of AWS architecture patterns and cost optimization Proficiency in scripting/automation languages (Python, Go, Bash); ability to write maintainable, scalable automation Hands-on experience designing and implementing comprehensive monitoring, logging, metrics collection, and distributed tracing Ability to instrument systems for visibility and rapid diagnosis Experience optimizing systems at scale and handling growth challenges Strategic thinking about capacity planning, resource requirements, and cost management Understanding of relational databases, including performance considerations and optimization Knowledge of networking concepts, caching patterns, and load balancing strategies Understanding of distributed systems concepts and failure modes Comfortable with AI/LLM tools (Claude, ChatGPT, Cursor, etc.) and uses them effectively in day-to-day work Ability to validate AI/LLM output and use it as a jump-off point (ask better questions, iterate, and verify with real evidence) Experience experimenting with AI for infrastructure automation, code generation, and problem-solving Proactive thinking about how AI can accelerate platform engineering work Experience leading the design of resilient systems and conducting architecture reviews Proven ability to elevate other engineers' capabilities and establish technical standards Strong communication skills; able to explain technical complexity to diverse audiences Ability to write clear ADRs and technical documentation that guides future decisions Demonstrated ability to influence technical direction without formal authorityBenefits: Work with a talented but humble team Competitive compensation and equity Weekly paid family meal 401k, medical, dental, and vision insurance Flexible PTO Policy 15.5 company-paid holidays including Juneteenth, MLK Day and a 1-week company shut down To support remote work conditions, AudioEye provides each employee a one-time stipend of $300

Company & context

Evidence is labeled so you can tell internal community data from public sources.

Trust-check this listing

Verify scam risk and ghost-job signals before you apply.

Related roles

Browse more remote Software Engineer jobs.

Staff Backend Product Software Engineer, Core Sync

Dropbox
remoteAnywhere in the World
Mar 7, 2026
View details

Senior Staff Software Engineer - STCE

Samsara
remoteCanadafull-timeCA$154,800 - CA$255,420
Posted 1w ago
View details

Associate Staff Engineer, Sap-Product-Lifecycle-Mgmt

Nagarro
remoteAnywhere in the World
Mar 5, 2026
View details

Staff Software Engineer

Sanderson
remoteCardiff$80,000 - $90,000
Posted 1w ago
View details

Source: Google Jobs • Last updated May 28, 2026