Skip to content

Qualitative Data Categorisation Improvements

LLM-based multi-label text classifier for NHS patient experience comments, replacing a legacy ML model with a RAG-augmented approach requiring no model training

Diagram showing the classification pipeline: patient comments pass through RAG retrieval of 20 similar examples from 13,000 labelled comments, then to Claude Sonnet in batches of 50, producing categories, segments and sentiment. Below, an example shows the comment "The staff were great but I couldn't get parked" split into two segments: "The staff were great" labelled Staff manner and Positive, and "I couldn't get parked" labelled Parking and Negative. Performance metrics show Macro F1 of 0.776 versus 0.70 for the legacy model.

Patient experience teams across the NHS collect thousands of free-text comments via the Friends and Family Test (FFT). These comments need categorising against 31 themes from the Qualitative Data Categorisation (QDC) Framework to identify patterns and drive service improvement.

This is a multi-label classification problem — a single comment can be assigned several categories simultaneously. For example, "The staff were great but I couldn't get parked" covers both "Staff manner" and "Parking".

The legacy approach used a trained sklearn/BERT ensemble that required manual retraining and achieved ~0.70 weighted F1.

We replaced this with an LLM-based classifier using Claude Sonnet. The system uses retrieval-augmented few-shot prompting — for each comment being classified, it retrieves semantically similar examples from a corpus of 13,000 labelled comments and includes them as in-context calibration. Comments are processed in batches of 50 per API call for efficiency. This approach requires no model training, and making category changes is as simple as updating a prompt.

Beyond classification, the LLM approach enables segmented sentiment highlighting — each comment is broken into its component clauses, with each segment assigned its own category and sentiment. This means "The staff were great but I couldn't get parked" produces two segments: "The staff were great" (Staff manner, Positive) and "I couldn't get parked" (Parking, Negative). This gives patient experience teams a much richer, more actionable view of feedback than flat category labels alone.

Results

The LLM approach achieves:

  • Macro F1: 0.776 (vs ~0.70 for the legacy model)
  • Weighted F1: 0.777
  • Multi-label classification across 31 categories with no training data pipeline
  • Per-segment sentiment and evidence extraction
Output Link
Open Source Code & Documentation (Currently Private) Link
Technical Methodology (Currently Private) Link
Original Framwork Documentation Link