Remote Site Reliability Engineer
Key details
- Compensation
- $48,000 - $88,000
Job Description
Salary: £48,000 - 88,000 per year
Requirements
- Around 4+ years of experience in Site Reliability Engineering (SRE), DevOps, or System Engineering.
- Strong knowledge of cloud platforms (AWS, Azure, or GCP) and cloud-native architectures.
- Experience with observability and monitoring tools (Prometheus, Grafana, ELK, Datadog, New Relic).
- Proficiency in Infrastructure as Code (IaC) tools such as Terraform, CloudFormation, or Pulumi.
- Hands-on experience with containerization and orchestration (Docker, Kubernetes, Helm).
- Strong Linux system administration and networking fundamentals.
- Experience with incident management, debugging, and root cause analysis.
- Proficiency in scripting (Bash, Python, or Go) for automation and system monitoring.
- Knowledge of load balancing, failover strategies, and distributed systems.
- Understanding of security best practices, access control, and compliance requirements.
- Strong communication skills and the ability to collaborate with cross-functional teams.
Responsibilities
- Design and implement scalable, reliable, and fault-tolerant systems across cloud environments.
- Develop and maintain observability tools, including monitoring, logging, and alerting (e.g., Prometheus, Grafana, Datadog, ELK).
- Automate infrastructure provisioning, deployment, and incident response using Infrastructure as Code (IaC) tools like Terraform or CloudFormation.
- Optimize system performance, scalability, and incident response workflows to improve uptime.
- Work closely with development and DevOps teams to improve system design for reliability.
- Conduct root cause analysis (RCA) and implement preventative measures to minimize failures.
- Ensure high availability by designing and maintaining load balancing, failover, and disaster recovery strategies.
- Improve CI/CD pipelines to enhance deployment speed while maintaining stability.
- Optimize cloud cost and resource utilization for AWS, Azure, or Google Cloud Platform (GCP).
- Participate in on-call rotations to quickly address system failures and minimize downtime.
- Work onsite in our Lisbon office and collaborate closely with cross-functional teams; we can support relocation efforts.
Technologies
- AI
- AWS
- Azure
- Bash
- CI/CD
- Cloud
- Datadog
- DevOps
- Docker
- ELK
- GCP
- Grafana
- Hardware
- Helm
- Support
- Kubernetes
- Linux
- Load Balancing
- Prometheus
- Python
- Security
- Terraform
More
We are Air Apps, a family-founded company born in Lisbon in 2018 with offices in Lisbon and San Francisco. We are on a mission to build the worlds first AI-powered Personal & Entrepreneurial Resource Planner (PRP) and have remained self-funded while reaching over 100 million downloads worldwide. We offer a dynamic, fast-paced environment with Apple hardware for work, an annual bonus, health and life insurance, transportation budget, Coverflex benefits, childcare support, an Air Conference to connect and grow with the team, a pension fund, Urban Sports Club membership, and free meals at the hub. We are committed to diversity, inclusion, and equity, and we welcome applicants from all backgrounds.
last updated 27 week of 2026
Company & context
Evidence is labeled so you can tell internal community data from public sources.
Range from 18 indexed roles at this employer: $48,000 - $88,000(mid ~68000)
Context may refresh in the background.
Trust-check this listing
Verify scam risk and ghost-job signals before you apply.
Related roles
Browse more remote Software Engineer jobs.
Remote DevOps Engineer Backend
Remote Site Reliability Engineer
Remote Backend Engineer
Source: DevITJobs • Last updated 2w ago