U.S. flag

An official website of the United States government

NCBI Bookshelf. A service of the National Library of Medicine, National Institutes of Health.

Rivas C, Tkacz D, Antao L, et al. Automated analysis of free-text comments and dashboard representations in patient experience surveys: a multimethod co-design study. Southampton (UK): NIHR Journals Library; 2019 Jul. (Health Services and Delivery Research, No. 7.23.)

Cover of Automated analysis of free-text comments and dashboard representations in patient experience surveys: a multimethod co-design study

Automated analysis of free-text comments and dashboard representations in patient experience surveys: a multimethod co-design study.

Show details

Chapter 8Evaluation of the rule-based information retrieval

Image 14-156-15-figu11

This short chapter considers the validity and reliability of the text analytics using statistical methods. The WCPES data and transferability performance are considered.

Methods

Statistics

A standard contingency table was used to classify performance (Table 11). For a binary classification problem, the table has two rows and two columns. Across the top are the observed class labels and down the side are the predicted class labels. Each cell contains the number of predictions made by the classifier that fall into that cell. Using this table, standard calculations were made for accuracy, precision, sensitivity and the F-score.

TABLE 11

TABLE 11

Contingency table for performance measurement

Classification accuracy is a common performance statistic. It is the number of correct predictions made divided by the total number of predictions made, expressed as a percentage. However, accuracy statistics alone should not be relied on, because, for example, it may be desirable to compromise accuracy to achieve greater predictive power. This is known as the ‘accuracy paradox’:

Accuracy=[true positive (TP)+true negative (TN)]/[TP+TN+false positive (FP)+false negative (FN)].
(10)

Precision is the number of TPs divided by the number of TPs and FPs. It is also called the positive predictive value. Low precision (< 1) can mean that the system throws up a large number of FTPs:

Precision=TP/(TP+FP).
(11)

Sensitivity is the number of TPs divided by the number of TPs and the number of FNs. This is also known as the recall or TP rate:

Sensitivity=TP/(TP+FN).
(12)

The F-score shows the balance between precision and sensitivity. Thus:

F-score=2×[(precision×sensitivity)/(precision+sensitivity)].
(13)

When calculating these statistics for the WCPES data, the study team used the half data set that had been reserved for this purpose (having built the analytics from the other half).

Adapting the approach for the analysis of Life After Prostate Cancer Diagnosis study data

The LAPCD study is a UK-wide patient-reported outcomes study that has generated information to improve the health and well-being of men with prostate cancer.151 Prostate cancer survivors (18–42 months post diagnosis) in all four UK countries (≈70,000) were identified through cancer registration systems and were sent postal surveys. A survey instrument was developed, covering a range of generic and cancer-specific patient-reported outcome measures, with additional items covering treatments received, sociodemographic characteristics and the patient perspective on their disease, treatment and experiences. The survey was structured into sections covering specific issues facing prostate cancer survivors [e.g. treatment decision-making (TDM) and decision regret; emotional well-being; and coping and self-management]. In addition to the closed-response items, eight free-text response questions were included at the end of each survey section for respondents to add further detail or to capture other relevant issues not covered in the section.

The adapted GATE software was used to identify themes within the responses to two free-text questions. First, at the end of section 3 of the questionnaire, which elicited 10,358 comments, a free-text question asks respondents the following:

  • Please add anything else that you would like to tell us about your diagnosis, treatment and the decision-making process.

Second, a final ‘catch-all’ free-text question, to which 6682 respondents provided comments, asks the following:

  • Is there anything else you would like to tell us about what life has been like for you following your prostate cancer?

This catch-all question was analysed as a direct test of transferability.

With regard to the first question on TDM, the LAPCD data research team involved four members of the study UAG to help inform the development of a gazetteer. Following a teleconference between the research team and the UAG on 30 March 2017, a random sample of 400 free-text responses to the TDM question were equally divided between the UAG members. The UAG members then read the comments and identified words and phrases that indicated comments describing the TDM process experienced; the treatment options with which they were presented; the level of involvement they had in the decision-making process; the amount of information and preparation they received concerning potential side effects; and their subsequent satisfaction with the treatment received. UAG members returned these TDM lists of words/phrases to the research team members, who then collated them and categorised them under subheadings.

Immediately preceding the TDM free-text question in the LAPCD data questionnaire is a closed question asking respondents to indicate their level of involvement in the decision-making process that determined their treatment (Box 3).

Box Icon

BOX 3

Closed question in the LAPCD data survey that is associated with the TDM free-text question

Comments provided by respondents were categorised on the basis of the level of involvement that patients indicated they had in the TDM process. This was done to identify common themes within the comments pertaining to levels of TDM involvement.

This information was used to develop gazetteers and rules to test for transferability.

patientopinion.org data

The study team obtained 10,000 comments from patientopinion.org (now known as Care Opinion) to test the approach on a different form of patient feedback free-text comments with often lengthy narratives.

Results and discussion

With the current approach, the following statistics for the WCPES data were calculated:

  • accuracy = 86%
  • precision = 88%
  • sensitivity = 96%
  • F-score = 92%.

The system was able to handle comments such as:

I am convinced [word unreadable] helped me. The special [word unreadable] drinks/desserts could be available post-surgery also.

It annotated this as belonging to hospital resources, drinks, food and helpful.

WordNet led to some overcategorisation through its use of metonyms, but also meant that the system was surprisingly accurate in ambiguous cases. For example, it was able to code patient-centred care well, even though this might be seen as a more conceptual code.

Some FPs were subjective. For example:

The level of support available and the quality of care has been high.

The above was classified as errors and safety, which was deemed to be incorrect, but which may be considered a positive example (care quality).

Some FNs were hard to explain. For example, the following was not categorised under waiting times, even though the word ‘waiting’ is in the comment:

The care given to me and the time I spent waiting for my operation was very good, fast turnaround. In fact, having the mammogram and quick diagnosis saved my life.

This may be a function of using NLP, which has inherent problems as well as advantages:

  • NLP has problems processing noisy data, reducing the overall accuracy.
  • Choices need to be made at each stage of the pipeline, which increases flexibility, but also the potential for problems; therefore, rules sometimes lead to unexpected results.

As the study team also found, NLP is computationally resource heavy, slowing the process and meaning that computers with large memories are needed.

Overall, the approach is reliable, performs better than the previous R algorithms19 and similar systems in development, and is likely to pick up most comments, although results will also contain some comments that are not correctly themed. This is better than missing a lot of comments but always placing those it catches correctly, as it is harder to check for what is missed than what is placed wrongly. However, users of the system may perceive it to be performing less well if they see the errors, so this is something that needs to be explained in the toolkit user guide.

Statistics for the LAPCD study data are less favourable. Thus:

  • accuracy = 47%
  • precision = 68%
  • sensitivity = 47%
  • F-score = 56%.

This is because the data contain much information about daily living rather than health care. Thus, a statement by a patient to say that they do not go out much anymore because they need to go to the toilet all the time as a result of their prostate cancer was categorised by the system as a negative comment about facilities.

At the time of writing this report, the study team has not yet run calculations on patientopinion.org data, which was not part of the original remit. This will be an interesting test of the system, as the comments are much longer.

Conclusions

The approach performs well on CPES data and, although transferability statistics seem disappointing, the study team was exploring the ‘worst-case scenario’, in which no adaptations were made. In fact, the system can be easily modified to increase the accuracy of different data sets. It is because of the ease with which rules can be tweaked that rule-based IR was chosen over template-based machine learning.

It should also be noted that the study team was conservative in scoring for the LAPCD data and this raises the point that such calculations depend on the research question. Although it was felt that the toilet example was wrongly coded, at least one senior LAPCD data researcher did not, as their research question was not so much ‘How can health care be improved?’ but ‘What is this patient’s overall experience of cancer?’.

The study team will also need to explore whether or not modifications for transferability reduce the accuracy of the system for CPES data; the more complex rule-based approaches become, the more unstable they may be. Thus, it may be important to keep different uses packaged separately.

It should also be noted here that any system (computational and manual), including that of the study team, has limitations, such as problems with some comments. No result from any form of IR should be taken at face value. It is essential that results are checked and confirmed, and this involves manually delving into the text under study.

Problems for the system included:

  • redacted comments, such as ‘Was not referred to a lymphedema clinic until I attended a mobile [word unreadable] unit workshop’ (although often the sytem got these right)
  • some comments with bizarre punctuation (e.g. ‘I chose . . . ’) or only blank spaces.

The error rates that these cause are data set specific. For the WCPES, these problems accounted for 0.5% of the data. This should not affect the practical overall accuracy of the system, as these comments are very unlikely to contain useful information anyway. Therefore, none of the calculations includes these comments. This means that the calculations are based on cleaned data. This approach was chosen because the study team was using the statistics to determine the areas in which the rules and gazetteers needed refinement.

The ability of the system to handle the more conceptual code of patient-centred care needs further exploration.

Copyright © Queen’s Printer and Controller of HMSO 2019. This work was produced by Rivas et al. under the terms of a commissioning contract issued by the Secretary of State for Health and Social Care. This issue may be freely reproduced for the purposes of private research and study and extracts (or indeed, the full report) may be included in professional journals provided that suitable acknowledgement is made and the reproduction is not associated with any form of advertising. Applications for commercial reproduction should be addressed to: NIHR Journals Library, National Institute for Health Research, Evaluation, Trials and Studies Coordinating Centre, Alpha House, University of Southampton Science Park, Southampton SO16 7NS, UK.
Bookshelf ID: NBK543275

Views

  • PubReader
  • Print View
  • Cite this Page
  • PDF version of this title (14M)

Other titles in this collection

Recent Activity

Your browsing activity is empty.

Activity recording is turned off.

Turn recording back on

See more...