U.S. flag

An official website of the United States government

NCBI Bookshelf. A service of the National Library of Medicine, National Institutes of Health.

Rivas C, Tkacz D, Antao L, et al. Automated analysis of free-text comments and dashboard representations in patient experience surveys: a multimethod co-design study. Southampton (UK): NIHR Journals Library; 2019 Jul. (Health Services and Delivery Research, No. 7.23.)

Cover of Automated analysis of free-text comments and dashboard representations in patient experience surveys: a multimethod co-design study

Automated analysis of free-text comments and dashboard representations in patient experience surveys: a multimethod co-design study.

Show details

Chapter 1Background and introduction

Patient experience surveys

Increasing attention is being paid by health-care providers to the patient-reported experience, which is at the core of all that the NHS does.1 Patient feedback on their experiences has the potential to drive care-quality improvements, highlight system failures and improve safety, reduce patient harm and increase satisfaction with health-care provision.24

Patient experience is often determined formally through surveys. These are considered to be important indicators of the quality of health service provision and service improvement priorities59 because they involve patients’ – or sometimes their carers’ – own evaluations. The NHS has led the way internationally in mandating a national patient survey programme in England since 2001.10 There are currently hundreds of such surveys administered at various levels – locally, nationally and internationally – across a range of institutional settings and patient groups.10 The Cancer Patient Experience Survey (CPES), the survey at the core of this study, is an example of a condition-specific patient experience survey (PES). There are many informal sources of patient experience that further add to the knowledge base on NHS websites and other websites dedicated to patient feedback, such as PatientOpinion (www.careopinion.org.uk), iWantGreatCare (www.iwantgreatcare.org), NHS Choices (www.nhs.uk/pages/home.aspx), blogs, social media and online fora.

It is important for patient feedback data to be properly used, and this requires appropriate tools for the collection, analysis, presentation, engagement with and understanding of the data if they are to improve health services effectively. Surveys mostly comprise closed questions, that is, questions with a fixed set of possible answers. This includes yes/no and multiple-choice responses, Likert scales, visual analogue scales and rating questions; the defining factor is that all of these are measures that can be easily quantified for meaning-making. Methods of analysing these have been well developed.11 Using these, the CPES, which began in 2010, is widely acknowledged as the most successful national PES in enabling and embedding service improvement.12 This has been achieved by providing trusts with tailored (trust-specific) statistical feedback of responses to the CPES closed questions, which benchmarks their performance against that of all other trusts.13 More generally, there is evidence that PESs are also effective when their findings are linked to government-led campaigns, targets and incentives,6 such as Care Quality Commission (CQC) assessments.

Patient experience survey free-text comments

To add context to closed-question responses, it is common practice for open-ended questions to be provided within PESs for respondents to leave free-text comments.7 CPES respondents, for example, are offered three such questions; one in three patients writes comments.9,12 The ≥ 70,000 free-text comments produced by the CPES each year are anonymised by Quality Health, then selectively provided to NHS trusts, but in unstructured raw data form (unlike the closed-question responses), and so cannot be easily linked to the quantitative analysis or used in any systematic way nationally. The National Institute for Health Research (NIHR) commissioning brief [Health Services and Delivery Research (HSDR) programme 14/156] for this study noted that few organisations have sufficient analytical capacity to interpret large numbers of such complex data. The brief further noted that there was uncertainty as to how to present such data in a meaningful and granular way to stimulate local action, and this is where the study focus was placed.

Current limitations in the usefulness of free-text comments

Critically, there is currently no system to efficiently and usefully analyse and report free-text responses in PESs.7 The conventional approach is manual thematic analysis (i.e. researchers reading through all of the text and assigning topic or theme codes to each comment). Such work requires considerable staff resource and can take months; thus, the number and costs of analysing these data each year using qualitative methods currently precludes its systematic analysis or reporting.7 The usefulness of the data thus depends on individual willingness and capacity.7 Even when analyses are undertaken, as with the 2013 London CPES free-text comments, there is a significant time lag; the London data were released in June 2013, the analysis was completed in December 2014 and was then published in June 2015.14 There are several problems with such delays. First, the services the comments relate to may have changed considerably and most comments may no longer be relevant. Second, staff could become demoralised if they received misplaced negative feedback on services that they have improved in the meantime. Third, patients themselves worry that their data are not effectively used. Fourth, because the manual approach is resource-heavy in human labour, financial cost and time, it cannot practically be reapplied to new waves of data.

Overall then, because these data are not presented to trusts in a structured and easily accessed and assimilated form, they are unlikely to have much of an impact on service change and commissioning decisions. However, free-text comments potentially provide rich insights into patient experiences that underpin, illustrate and complement the closed-question responses.7,12,15 The HSDR programme brief reflects the current need for better use of such data and highlights this as being critical in informing service provision in a challenged NHS.16,17

Alternative approaches to the analysis of free-text comments

Despite the limitations of manual thematic analysis, there have been few previous attempts to analyse large-survey, unstructured, free-text responses in health care, including cancer services, to inform clinical and health-care practice.7,12,15,18,19 There are notable exceptions,14,1834 but these all represent one-off analyses (see Chapter 2).

NVivo version 9 (QSR International, Warrington, UK) and other qualitative data analysis (QDA) packages, including dedicated text-analysis software, such as QDA Miner (Provalis Research, Montreal, QC, Canada), are often used to organise manual thematic analyses. This might be considered particularly helpful with large data sets. However, in the experience of the study team, NVivo is unstable when handling a large amount of text; the software creates considerable metadata from relatively small text inputs, resulting in file sizes of several gigabytes. These restrictions became problematic when there were around 4600 comments in previous work.19 Rapid automated QDA outputs that might be thought useful include automated coding, word clouds, concordance and word-frequency tables and associated statistics (against a comparator dictionary or data set). However, each of these uses basic word-search technology and ignores semantics and syntactics. Automated codes augment but cannot replace even basic manual coding. Word-cloud software generates ‘pictures’ of words from word-frequency lists, but only includes the more frequent words and does not take account of synonyms or closely related words. Depending on the ‘cloud’ shape chosen by the user, the words selected will vary considerably. Word-frequency tables enable the most used words and significant differences between data sets to be determined, but can be hard to interpret. Concordance software enables the user to click on a word to access its use in context, so that a more detailed understanding can be obtained. All of these approaches enable the big picture to be quickly obtained, but present an overwhelming number of irrelevant as well as relevant words, and more so the larger the sample size. Many of the words will be filler words (devoid of semantic content, but used in natural talk and text for flow or to hold attention) and many will be meaningless out of context. Text lists and word clouds will also contain many words with similar meanings that will appear to be of lower frequency because they have not been combined, when they may together represent something very significant.35 If survey text responses are fairly simple – lists of symptoms, for instance – statistics-based solutions such as these may be useful. With even slightly more complex text responses, as with PESs, both specificity and sensitivity will be compromised.

A possible solution to the problem is to develop analyses of the data using more advanced computational approaches, which some people refer to as artificial intelligence, although the definition of this term is discipline dependent.36 Text mining has been previously tried with the 2013 Wales CPES (WCPES) free-text comments.19 The so-called template-based (otherwise known as memory- or instance-based) approach, a form of machine learning, was used; this has often been used for comparable classification tasks.37 This approach might be expected to be an improvement over manual methods and basic QDA approaches. The work was certainly useful in showing the challenges to quality of life that are often faced by cancer survivors, and associated gaps in health care.19 This led to a service model showing links between factors that negatively affected patient quality of life and potentially mediating factors.

However, a manual thematic analysis of 80% of the data was still needed to be undertaken using NVivo, meaning that the analysis process still took many months and, therefore, was only slightly faster than a fully manual approach. This is because memory- or instance-based learning approaches such as this classify new data by comparison against stored, ready-labelled instances. The version that was used, from R statistics software (The R Foundation for Statistical Computing, Vienna, Austria), relied on the k-nearest neighbour (KNN) algorithm, which is based on regression analysis and is one of the simplest machine learning approaches. In general, supervised learning methods take pairs of data points (X, Y) as input, in which X are the predictor variables (features) and Y is the target variable (label). The supervised learning method then uses these pairs as training sets and learns a model F, where F(X) is as close to Y as possible. This model F is then used to predict Ys for new data points X that are input into the system, using a statistical similarity measure such as Euclidean distance or cosine similarity. The KNN algorithm works out the distances between the new data point and all stored data points, determines from these which stored points are closest (the KNNs) and then labels the new data point using either a majority or a weighted vote based on this calculation. In other words, the closer the new data point is in its aggregate distance score to one that is already stored, the more likely it is to receive the same label as the stored data point. As this approach uses stored examples for these calculations (i.e. as templates to ‘learn’ or work out where new data will be classified), it cannot analyse data accurately if (1) the new data do not match these templates well or (2) the distances between the templates are too great. This means that for good sensitivity and specificity, this approach requires approximately two-thirds of the data to be manually analysed and labelled as templates,38 as was previously found.32 The remaining one-third of the data will then be matched accurately.

Although the previous study32 ultimately achieved a sensitivity of 78%, a precision of 83.5% and an overall F-score of 80% (see Chapter 8 for a further explanation of these scores), this approach is of limited use for repeated survey analyses. First, as explained above, it is not sufficiently automated to be useful for repeated waves of data and still does not possess the benefits of immediacy. Second, the algorithms that were used were pre-set by the R software developer and could not be modified. The ‘black-box’ approach was hard to evaluate,39 domain non-specific (i.e. not developed for health-care data per se) and insensitive to unexpected data. Each time new data need analysing, such algorithms need more training with new terms and words that may have arisen in the interim, such as new treatment names, because the algorithms are limited to the ‘controlled vocabulary’ of the templates.39

Information retrieval

Given these limitations, a memory-based template approach to text mining is not appropriate for repeated waves of data, such as those generated by the annual CPES. The solution is to customise a rule-based IR or knowledge-extraction approach. This approach has been used worldwide in biomedicine. Herland et al.40 provided a detailed description of some of the recent uses in health care, whereas Ordenes et al.41 explored the widespread use of this approach to analyse comments made about commercial products and provide customer feedback.

Despite the significant use of rule-based IR over many years, at the time that funding was applied for, it had not been tried for PES free-text comment data. Indeed, it is still believed that this is the first such application; however, since funding was obtained, some teams and commercial organisations have developed analyses for the Friends and Family Test on patient satisfaction (a different concept from experience) and social media and website patient feedback.42,43

Instead of using templates, rule-based IR methods involve a stage of lexicosyntactic preprocessing and development of gazetteers (sets of lists containing words and phrases for specific entities or concepts). This is followed by cascades of pattern-matching syntactic (grammar and other language structure) applications designed to find the targets for IR – in the case of the study, comments that match particular themes – in combination with rules that look up terms in the gazetteers.

The approach was based on the GATE open-source system, the developers of which have labelled their process as text or knowledge engineering,44 whereas the term rule-based IR is used in this report to encompass the whole analysis process; the study’s ‘text engineers’ have added further applications and lookup rules or search queries to the basic GATE pipeline (for an explanation of these terms, see Chapter 4). To develop the gazetteers and rules, theme labels and data have been used from previous manual (human hand-coded) and template-based machine learning projects that the study team has been involved in.

The aim in relation to IR of text has been to explore how a highly transferable rule-based approach might perform; however, this performance could be further improved by using the rule-based outputs as templates for template machine learning.

Benefits of the approach

This approach has many benefits. Lookup rules can tag or annotate comments as belonging to quite nuanced literal themes, by combining gazetteers using Boolean logic. The functions of the syntactic applications include part-of-speech (POS) tagging, in which an individual word (X in the XY annotation used above) is tagged on the basis of the POS it represents (e.g. whether it is a noun, verb or adjective). This can, for example, distinguish between the different uses of the test in ‘the test came back negative’ and ‘I will test the new treatment’. Importantly, words (X) are not independent of each other, because if the previous word was an adjective, it is far more likely that the next word will be a noun than a verb. This enables new words to be more likely to be categorised correctly than not, given these various points of information. Such features of IR have been used to develop an analytical process that can cope with the fragmented phrases and sentences that typify survey free text. The approach also makes use of semantic (meaning-making) information, such as co-location. Thus, for example, if a comment says ‘The nurse was present. She did not speak’, the system will know that ‘She’ refers to the nurse because that is the relevant POS in the nearest sentence. Sentiment analysis has also been used, in which the label Y is used to reveal the sentiment (positive or negative) of each comment. These and other advantages of rule-based IR over template-based machine learning, and the rationale for choosing it, can be summarised as follows:

  • External knowledge repositories [such as WordNet®, version 3.1 (Princeton University, Princeton, NJ, USA)], can be used in categorisation, so that the system can cope with novel data linked in these repositories to existing labels.
  • New words can be added to gazetteers in the time that it takes to type them.
  • By writing rules that are not dependent on specific words, but that take into account the arrangements of words within syntactic and semantic contexts, the software can handle new data with good sensitivity and specificity.
  • It is not required to annotate a large number of data to start the process, in contrast to memory-based machine learning. Thus, once the rules are written, the system can be used in virtual real time for future iterations of PESs with no modification to minimal modification (a matter of hours rather than months). This gives it potential speed advantages in future years.
  • The system can be made domain specific, increasing the relevance of its outputs,45 as the rules are written by specific individuals.
  • New rules could be written very quickly and added to the system in the future and, hence, the approach makes for a potentially more accurate, more adaptable and more future-proofed system than template-based machine learning.
  • The system could be easily modified by others adapting the system in the future; transferability to other health-care surveys was an important element in the HSDR programme brief.
  • Rules could be specifically written to tackle some specific issues with free-text comments (see Chapter 4).
  • The approach is amenable to real-time data processing.
  • It does not require training examples (using resource-consuming manual data categorisation) or redevelopment of the whole system each time novel data are added.

Previous research has shown that linguistics-based IR models such as this can outperform manual (human) categorisation of customer feedback reviews,46 an area in which IR is particularly used. Therefore, it shows promise in the analysis of patient experience feedback.

Engagement with the data: making it meaningful for all

The health sciences lag behind other disciplines in the use of computational approaches, such as IR.47 It has been suggested that this owes much to the chasm between person-centred care and the depersonalised approach of computing.47 In other words, it is unclear how it can be ensured that computational analyses benefit individual patients. There has been little work to bridge the chasm and use the outputs of computational approaches to directly improve health care for patients. Accordingly, when study design commenced, the pre-funding patient and public involvement (PPI) input included the concern that the results would be mechanistic and reductionist, with the survey respondents and the people and health-care processes they wrote about in their free-text comments being reduced to mere numbers and data points.48 There was concern that automation, generalisations and breadth would be favoured over the insights the comment boxes were meant to provide.

The study design therefore prioritises PPI (see Chapter 10) and includes a number of co-design-type substudies with patients, carers and health-care professionals that feed into the main analysis and the design of a digital toolkit with a dashboard or visual summary of the data, as well as the presentation of the newly structured (i.e. thematically grouped) free-text comments themselves. This is what Halford and Savage49 call a symphonic approach, with small qualitative and quantitative studies being used in a complementary fashion with large data set (‘big data’) computational approaches. For this reason, the outputs should lead to improvements in the health care of people with cancer in ways that are meaningful to them and incorporate and represent their perspectives and language.50,51

Aims and objectives

The primary aim was therefore to improve the use and usefulness of PES free-text comments in order to drive improvements in the patient experience, and to do so, it was considered important to use a mixture of rule-based IR and complementary smaller studies.

Therefore, it was intended to develop and validate a novel use of rule-based IR to provide rapid automated thematic analysis of large amounts of survey free text (using the CPES as the first case), and to develop and validate a linked ‘dashboard’ (which later became a toolkit) to display the results in a summary format that can be drilled down to the original free text and used by patients and staff alike.

The toolkit, and the use of co-design with stakeholders, gives the approach added value over a simple thematic analysis, in addition to the speed and automaticity that rule-based IR enables. The display is intended to illuminate service gaps and areas in which the patient experience can be practically improved at the team, NHS trust and national levels.

As one of the secondary aims, the transferability of the approach was explored. Manuals ensured transferability to free text in other health surveys more generally. The rule-based IR process (i.e. the topic-oriented automated free-text analysis) and the toolkit were designed to be quickly reproducible and easily modifiable across health-care topics.

The objectives were to integrate co-design and implementation science into the approach, output thematic analyses into a digital display, produce recommendations on toolkit design, validate the approach and ensure transferability. The work primarily asked:

  • Is the novel approach a valid, accurate way to analyse large-volume CPES free-text responses?
  • Can the approach be transferred to similar surveys on other health-care topics?
  • Is co-design (as defined here) with mixed UK stakeholders (patients, their partners/carers, NHS managers and clinicians) feasible and effective for the approach?
  • Is Normalisation Process Theory (NPT) useful for the approach?

Overview of the study

The study comprises three main stages with subcomponents in each, which are summarised in Figure 1. Parts of this figure are repeated throughout the report to orient the reader.

FIGURE 1. Overview of the study.

FIGURE 1

Overview of the study. The flow back from stage 3 to stage 2 represents the final refinements that were made to outputs as a result of stage 3 work.

Stage 1: preliminary (scoping) work

Review: see Chapter 2

A scoping review determined the key health-care dashboard design principles.

Dashboard-scoping survey and term and theme mindmapping survey: see Chapter 3

Patients, carers and health-care professionals completed an online survey to input into the text-analytics work that asked them to mindmap relevant terms and themes. In addition, they completed a separate survey on dashboard design to find out what potential end-users considered to be the key features for a good dashboard.

Rapid review of themes: see Chapter 3

Through a rapid review of the literature and incorporation of themes from previous relevant research, a draft taxonomy of themes was constructed, which was further developed in stage 2.

Stage 2: main development phase

Rule-based information retrieval work: see Chapter 4

The rule-based IR work involved modifications to GATE and further programming to solve issues with the analysis of PES data.

Prototype toolkit development: see Chapter 5

Prototypes were developed from stage 1 findings and refined through stage 2 as new data were collected and analysed.

Group workshops and interviews: see Chapter 6

The workshops incorporated group consensus concept-mapping techniques for naming and validating themes, and design discussions around dashboard prototypes, as a result of which the single-screen dashboard developed into a fuller toolkit. The interviews explored the issues in more depth.

Stage 3: validation and evaluation phase

Discrete choice experiment: see Chapter 7

A discrete choice experiment (DCE) was used to objectively validate a set of core features required of the toolkit for it to be taken up within health care, and it included a cost–benefit analysis based on the findings.

Validating the rule-based information retrieval: see Chapter 8

The sensitivity, specificity and accuracy of the approach was tested, as well as its transferability.

Toolkit evaluation through structured walk-through techniques: see Chapter 9

Techniques were used that are standard in technology evaluations to consider the usability and usefulness of the analysis and the toolkit and, combined with NPT, the work that these would do within health-care settings.

Theory

Normalisation Process Theory was used throughout the study. NPT is used to understand actions related to the implementation, embedding and integration of new technology or complex interventions.52,53 It comprises four core ideas or constructs: (1) coherence, (2) cognitive participation, (3) collective action and (4) reflexive monitoring. Each of these constructs is subdivided into four and, thus, there are 16 constructs in total. The NPT toolkit website [www.normalizationprocess.org/npt-toolkit.aspx (accessed 9 October 2018)] contains example questions that can be used to explore these, and represents the results graphically in radar plots.

Copyright © Queen’s Printer and Controller of HMSO 2019. This work was produced by Rivas et al. under the terms of a commissioning contract issued by the Secretary of State for Health and Social Care. This issue may be freely reproduced for the purposes of private research and study and extracts (or indeed, the full report) may be included in professional journals provided that suitable acknowledgement is made and the reproduction is not associated with any form of advertising. Applications for commercial reproduction should be addressed to: NIHR Journals Library, National Institute for Health Research, Evaluation, Trials and Studies Coordinating Centre, Alpha House, University of Southampton Science Park, Southampton SO16 7NS, UK.
Bookshelf ID: NBK543268

Views

  • PubReader
  • Print View
  • Cite this Page
  • PDF version of this title (14M)

Other titles in this collection

Recent Activity

Your browsing activity is empty.

Activity recording is turned off.

Turn recording back on

See more...