Remote ML Infrastructure Lead
Key details
- Compensation
- $63,000 - $97,000
Job Description
Salary: £63,000 - 97,000 per year
Requirements
- Proven experience in a senior MLOps, ML Platform, ML Infrastructure, Platform Engineering or Machine Learning Systems role
- Strong hands-on background in software engineering and cloud infrastructure, ideally with direct experience supporting production machine learning environments
- Experience building and operating systems that support the full ML lifecycle, from experimentation and training through to deployment and monitoring
- Strong knowledge of Python and sound engineering principles, including testing, automation and code quality
- Strong experience with cloud platforms such as GCP
- Experience with Docker, Kubernetes and modern containerised deployment patterns
- Strong experience with CI/CD pipelines, infrastructure-as-code and workflow orchestration
- Experience with tools such as Airflow or similar platform and orchestration technologies
- Good understanding of model observability, data quality, feature pipelines, lineage and reproducibility
- Experience designing scalable infrastructure for ML workloads, including training, batch inference and real-time serving
- Strong appreciation of reliability, security, governance and operational excellence in customer-facing or production-critical systems
- Ability to operate across both strategic and hands-on technical work
- Strong communication skills and the ability to work effectively across engineering, product and data teams
- Experience supporting computer vision, deep learning, LLM or other compute-intensive ML workloads
- Experience with GPU infrastructure, distributed training or high-performance compute environments
- Familiarity with feature stores, model registries and automated retraining pipelines
- Experience building internal developer platforms or self-service ML tooling
- Experience in regulated, high-security or high-availability environments
- Experience leading or mentoring engineers in a scale-up or high-growth technology business
- Familiarity with responsible AI, model governance or risk controls in production ML setting
Responsibilities
- Lead the design and evolution of our ML platform, infrastructure and MLOps capability
- Build and maintain scalable, reliable and secure systems for model training, testing, deployment, monitoring and lifecycle management
- Develop the infrastructure and tooling that enable ML Engineers, Data Scientists and Researchers to work efficiently and ship models with confidence
- Design robust workflows for CI/CD, model versioning, reproducibility, experimentation, feature management and release management
- Own and improve the production environment for machine learning systems, ensuring strong standards for availability, performance, observability and resilience
- Define and implement monitoring across model and platform layers, including system health, data quality, drift, latency, throughput and cost efficiency
- Build or optimise internal self-service tooling and platform capabilities to reduce friction for teams working on ML use cases
- Partner closely with ML, Data, Software and Platform Engineering teams to productionise models and improve the end-to-end ML development lifecycle
- Support the scaling of infrastructure for both training and inference workloads, including high-throughput, real-time or compute-intensive use cases where relevant
- Drive best practice in governance, security, compliance, auditability and operational rigour across the ML lifecycle
- Improve the efficiency and cost-effectiveness of ML systems, including cloud resource usage, compute environments and deployment patterns
- Mentor engineers and act as a technical leader across ML platform and operations topics
- Help define the roadmap for ML enablement, ensuring the platform can support current needs while scaling for future growth
Technologies
- AI
- Airflow
- CI/CD
- Cloud
- Computer Vision
- Docker
- GCP
- Support
- Kubernetes
- LLM
- Machine Learning
- Python
- Security
- DevOps
- IAM
- LESS
More
We are iProov, a science-based biometric identity company that helps security-conscious organizations streamline secure remote onboarding and authentication for digital and physical access. Our award-winning liveness technology and iSOC help defend against deepfakes and generative AI threats while delivering scalable user experiences. We work with leading governments and enterprises worldwide and value diversity, inclusion, psychological safety, and equal opportunities. This is a hybrid role based at WeWork Waterloo, reporting to our Chief Scientific Officer, with a negotiable base salary, company performance bonus, share options, and a wide range of benefits including flexible hybrid working, annual leave, growth shares, learning and development support, private health and wellbeing options, and other employee perks.
last updated 28 week of 2026
Company & context
Evidence is labeled so you can tell internal community data from public sources.
Range from 8 indexed roles at this employer: $63,000 - $97,000(mid ~80000)
Context may refresh in the background.
Trust-check this listing
Verify scam risk and ghost-job signals before you apply.
Related roles
Browse more remote Data Scientist jobs.
Senior Security Engineer (AI & DevSecOps)
Remote Tech Ops Analyst
Remote ML Engineer
Source: DevITJobs • Last updated 1w ago