SRE (Linux, Firmware & Server Infrastructure)
Key details
- Compensation
- $30,000 - $70,000
Job Description
Salary: £30,000 - 70,000 per year
Requirements
- We require strong Linux administration and troubleshooting expertise, including process and service management, system logs and diagnostics, networking fundamentals, and package/configuration management.
- We require a solid understanding of server hardware and infrastructure, including disks, RAID/HBA controllers, NICs, firmware interactions, and hardware failure modes with OS-level symptoms.
- We require proven experience with firmware management and upgrades, disk encryption, and secure server configurations.
- We require hands-on experience with remote management and lights-out technologies such as iDRAC, iLO, RACADM, Redfish, or similar APIs.
- We require a strong track record of incident ownership, including triage, mitigation, cross-team coordination, stakeholder communication, and driving issues through to resolution.
- We require experience working with vendor diagnostics, logs, support bundles, vendor escalation processes, and engineering engagement.
- We require excellent written and verbal communication skills, with the ability to explain technical issues to both technical and non-technical stakeholders.
- We require strong documentation skills, including creating runbooks, procedures, troubleshooting guides, knowledge base articles, and RCA reports.
- We prefer scripting and automation experience with Python, Bash, or Ansible.
- We prefer familiarity with configuration management and automation frameworks.
- We prefer exposure to virtualisation and containerisation technologies such as VMware, KVM, Docker, or Kubernetes.
- We prefer experience with monitoring, observability, alerting systems, log analysis, and alert tuning.
- We prefer understanding of SRE principles and metrics including SLOs, SLIs, error budgets, MTTR, and MTTD.
- We value a methodical, detail-oriented approach to troubleshooting, strong ownership, and the ability to work effectively in high-pressure, incident-driven environments.
- We require a collaborative mindset and the ability to work across global teams and vendors.
- We prefer candidates who are proactive about continuous improvement and operational excellence.
Responsibilities
- We own and manage end-to-end incident resolution for platform and hardware-related issues, including triage, mitigation, escalation, and post-incident review.
- We diagnose and troubleshoot Linux OS-level issues caused by hardware faults, firmware changes, or configuration inconsistencies.
- We manage firmware lifecycle processes, including upgrades, validation, and issue remediation.
- We work with disk encryption technologies and logging frameworks to ensure system integrity and auditability.
- We maintain and troubleshoot server configuration settings, including BIOS-level parameters across multiple hardware vendors, with a strong focus on Dell.
- We use out-of-band management tools such as iDRAC, iLO, RACADM, and Redfish APIs for remote diagnostics and recovery.
- We analyse vendor logs, support bundles, and telemetry data to identify root causes and remediation paths.
- We engage directly with hardware vendors and engineering teams, managing escalations and driving timely resolutions.
- We contribute to continuous improvement initiatives that reduce incident recurrence and operational toil.
- We produce and maintain high-quality documentation, including runbooks, troubleshooting guides, and knowledge base articles.
- We participate in post-incident reviews and support improvements in reliability metrics such as MTTR, MTTD, and SLOs.
Technologies
- Ansible
- Bash
- Docker
- Firmware
- Hardware
- Support
- KVM
- Kubernetes
- Linux
- Python
- VMware
- DevOps
More
We are hiring a Senior Platform Reliability Engineer for a 6-month contract based in Glasgow with a hybrid working model of 3 days onsite. This is a hands-off hardware role within a high-performing enterprise infrastructure team, focused on improving the reliability of critical platforms at scale. The day rate is negotiable and the role is inside IR35 via umbrella. You will work closely with internal teams and external vendors in a fast-paced, incident-driven environment, helping to resolve complex platform and hardware issues and contributing to operational excellence and reliability improvements.
last updated 20 week of 2026
Company & context
Evidence is labeled so you can tell internal community data from public sources.
Range from 30 indexed roles at this employer: $30,000 - $143,000(mid ~90013)
Context may refresh in the background.
Trust-check this listing
Verify scam risk and ghost-job signals before you apply.
Related roles
Browse more remote Backend Developer jobs.
Network Infrastructure Automation Engineer
Infrastructure Automation Engineer
ACI Network Engineer - UK
Source: DevITJobs • Last updated May 21, 2026