AI/ML LLM Engineer
Key details
- Compensation
- $8,400 - $36,000
Job Description
Salary: £8,400 - 36,000 per year
Requirements
- We want hands-on experience building, training, or fine-tuning transformer-based LLMs such as LLaMA, Mistral, or similar.
- We want practical experience with LoRA, SFT, and RLHF/DPO using Hugging Face Transformers and PEFT.
- We want strong Python skills and experience building synthetic data generation pipelines.
- We want AWS SageMaker experience for model training, hosting, and inference.
- We want exposure to real-time STT/TTS pipelines.
- We want you to be comfortable working independently with daily standups and Jira-based task tracking.
- We will not consider you if you have only used OpenAI APIs and have never built or fine-tuned a base model.
- We will not consider you if you are currently engaged in other contracts that will compete for your time.
- We will not consider you if you cannot commit to the daily standup and monthly deliverable rhythm.
Responsibilities
- We are building the large language model that powers the inCall AI-powered receptionist for NHS GP practices.
- We will integrate OpenAI GPT into the live 3CX call pipeline from day one so the full voice flow is testable immediately.
- We will build a fine-tuned LLaMA 3 8B model from scratch, including architecture design, 10,000+ synthetic NHS GP training examples, LoRA/SFT fine-tuning, RLHF/DPO, and SageMaker deployment.
- We will connect AWS Transcribe (en-GB) with custom NHS GP vocabulary for real-time speech-to-text.
- We will integrate ElevenLabs for text-to-speech.
- We will download and register LLaMA 3 8B or Mistral 7B in SageMaker Model Registry.
- We will complete LoRA fine-tuning on a synthetic NHS GP dataset using Hugging Face Transformers and PEFT on ml.g4dn.xlarge.
- We will deploy a SageMaker endpoint, incall-poc-llama-endpoint, targeting at least 70% accuracy on a 100-example test set.
- We will conduct human evaluation of GPT versus LLaMA conversational quality.
- We will integrate the NHS PDS FHIR API for patient context.
- We will demonstrate all eight core NHS GP call scenarios using the LLaMA model.
- We will replace OpenAI GPT with the LLaMA SageMaker endpoint in the live pipeline.
- We will deliver a full production readiness report and technical documentation.
Technologies
- AI
- API
- AWS
- Flow
- Support
- JIRA
- Python
- Cloud
- LLM
More
We are inCall, an AI-powered receptionist for NHS GP practices, and this is Stage 1 of our four-stage product roadmap. Our broader plan adds AI triage and auto-routing, then clinical routing support, and ultimately a fully autonomous AI call handler. We operate a cost-per-call licensing model across GP practices, and an engineer who delivers a clean, well-documented LLaMA proof of concept will be well placed for the next contract. This is a remote role with a structured hiring process that includes an application form, video submission, group interview, one-to-one technical interview, and offer stage. We offer pay of £700.00-£3,000.00 per month.
last updated 26 week of 2026
Company & context
Evidence is labeled so you can tell internal community data from public sources.
Range from 4 indexed roles at this employer: $8,400 - $70,000(mid ~30400)
Context may refresh in the background.
Trust-check this listing
Verify scam risk and ghost-job signals before you apply.
Related roles
Browse more remote Software Engineer jobs.
Source: DevITJobs • Last updated 3w ago