U.S. flag

An official website of the United States government

NCBI Bookshelf. A service of the National Library of Medicine, National Institutes of Health.

Selby PJ, Banks RE, Gregory W, et al. Methods for the evaluation of biomarkers in patients with kidney and liver diseases: multicentre research programme including ELUCIDATE RCT. Southampton (UK): NIHR Journals Library; 2018 Jun. (Programme Grants for Applied Research, No. 6.3.)

Cover of Methods for the evaluation of biomarkers in patients with kidney and liver diseases: multicentre research programme including ELUCIDATE RCT

Methods for the evaluation of biomarkers in patients with kidney and liver diseases: multicentre research programme including ELUCIDATE RCT.

Show details

Chapter 6How can monitoring impact on patient outcomes?

Much of the test evaluation literature centres on establishing key test properties such as test accuracy. However, the ultimate use of any test in clinical practice should be based on the knowledge that testing does more good than harm to patients. Comparison of patient outcomes resulting from different interventions is ideally assessed using a RCT design and the same design can be applied to the evaluations of tests. RCTs are less commonly used for assessing medical tests but are increasing in number both for diagnostic68 and for monitoring tests (see Chapter 4), such that a thorough understanding of the ways in which testing can affect patient outcomes is important.

Patient monitoring is undertaken for many purposes, most obviously within the context of ongoing treatment as the main tool for treatment titration and maintenance, with the goal being to maintain test results within certain limits of a given marker until such a time as treatment can be discontinued or an alternative treatment is needed. Our particular interest is in monitoring people who have a known disease or condition that is likely to progress or recur at some point in the future but that does not yet require treatment. Patients are usually asymptomatic (e.g. following primary treatment for the first occurrence of a disease), but may be mildly symptomatic but not yet receiving treatment or may experience symptoms of a disease that puts them at risk of developing other conditions. The primary goal is usually earlier treatment or the avoidance or delay of treatment, with the crux of monitoring being to detect the need for a change in patient management in a timely manner.

Particular challenges to evaluating the impact of monitoring tests on patient outcomes are, first, that the effect on outcomes will be relatively small, thus requiring large samples of patients to demonstrate statistically significant effects, and, second, that changing patient outcomes is reliant on patients and clinicians following potentially complex protocols both for testing and for treatment.

Over the last 10–15 years a number of framework papers related to the development and evaluation of tests for screening, diagnosis, prognosis and treatment monitoring purposes have been published, many of which have been comprehensively reviewed by previous authors.315,316 We have selected three frameworks of particular relevance to the consideration of patient outcomes in monitoring. The first, by Adriaensen et al.,317 presents a stepwise evaluation process for new screening strategies, which includes a consideration of the trade-off between the harms and the benefits from a new test. The second, by Ferrante di Ruffano et al.,68 aims to assist those evaluating diagnostic tests to understand the ways in which changes to testing strategies can affect patient outcomes. The third, by Lord et al.,318 considers the circumstances in which randomised evidence of patient impact from a new diagnostic test may be needed. With these in mind, our aim was to consider the potential impact of monitoring on patient outcomes, illustrated by our review of randomised trials of monitoring strategies.

Methods

A monitoring care pathway was outlined to identify in simple terms the points at which monitoring might affect outcomes (Figure 10). The three identified frameworks68,317,318 were reviewed in terms of their relevance to this monitoring context and 58 trials from a review of RCTs in which monitoring was carried out in at least one arm of the trial (see Chapter 4) were used for illustration purposes. The trials were grouped into three main categories in terms of the change in patient care under evaluation and the intended impact on patient outcome:

FIGURE 10. Monitoring care pathway: (a) pathway; and (b) detail of ongoing monitoring process.

FIGURE 10

Monitoring care pathway: (a) pathway; and (b) detail of ongoing monitoring process. Adapted from Ferrante di Ruffano et al. with permission. +ve, positive; –ve, negative.

  1. A new monitoring strategy compared with an existing monitoring strategy, such that the current monitoring test might be replaced by a new and more accurate test, a new test might be added to the strategy or the currently used test might be applied at a different intensity or with an alternative threshold for intervention. Depending on the associated change in patient care, the new monitoring strategy may be intended to detect patients at an earlier stage of disease, to more accurately detect those in need of treatment or to reduce the invasiveness or frequency of testing.
  2. A monitoring strategy compared with immediate treatment of all patients at risk of an adverse outcome, in which monitoring may be used to avoid or delay treatment in those who do not need it.
  3. A monitoring strategy compared with no monitoring, in which patients are usually treated on the basis of clinical presentation only and the likely aim of monitoring is the detection and treatment of disease at an earlier stage.

In the following sections, we first consider the similarities and differences between monitoring, screening and diagnosis, before broadly outlining the potential for benefit and harm from monitoring and considering the ways in which patient outcomes can be mediated by particular aspects of the monitoring care pathway according to the aim of monitoring and the change in strategy under evaluation.

Monitoring compared with screening or diagnosis

The monitoring care pathway outlined in Figure 10 bears close resemblance to that for diagnosis and for screening.68,317 In a monitoring context, (1) a test is administered according to a predetermined schedule to detect a target condition or some precursor or marker of that condition, (2) the test result is considered (often in relation to previous measurements and with the potential for repeat testing to confirm abnormal or indeterminate results), (3) the test result is considered alongside other evidence (usually including the results of further investigations) to decide whether or not therapeutic intervention is needed and (4) the necessary intervention is implemented.

Where the pathway diverges from that for diagnosis is with the added dimension of repeated testing over time and a merging of the ‘diagnostic’ and ‘management’ decisions outlined in the diagnostic care pathway of Ferrante di Ruffano et al.68 The serial nature of testing for monitoring purposes can affect patient outcomes in a number of ways, most obviously by increasing the physical and psychological burden of testing on patients but also potentially impacting on other outcomes, for example through patient and physician compliance with testing protocols. Furthermore, although a diagnostic test informs both a ‘diagnostic decision’ (often when more than one differential diagnosis may be available) and a ‘management decision’ (assisting in the choice of a range of therapeutic options), a monitoring test is often relatively less definitive, providing more of a guide to the need for changes in patient management. A positive monitoring test result frequently triggers further investigation to determine whether or not and when a particular treatment should be implemented, rather than informing the choice of one of a range of therapeutic options. In this respect, monitoring is more akin to screening, in which a test is applied repeatedly over time to detect and treat a particular clinical condition, rather than to differentiate between diagnoses, with the caveat that monitoring populations have a higher risk of the clinical event of interest occurring and that, although the ultimate goal of a new screening programme is usually a reduction in disease-specific mortality, monitoring can be implemented for a range of reasons.

Potential benefits and harms from monitoring

Figure 11 uses the concept of a 2 × 2 contingency table to illustrate that assigning potential benefits and harms from a monitoring test is not as straightforward as might be imagined. For simplicity, the following is set mainly in the context of a new monitoring strategy to allow earlier detection and treatment.

FIGURE 11. Summary of the potential harms and benefits from a new monitoring strategy.

FIGURE 11

Summary of the potential harms and benefits from a new monitoring strategy. Harms and benefits are similar to those identified from a screening context by Adriaensen et al. FN, false neagative; FP, false positive; TN, true negative; TP, true positive. (more...)

In general terms, for those patients with a ‘true’ result, benefits accrue both to those who would otherwise have experienced a poor clinical outcome but for the new test (A) and to those who would have been detected clinically and successfully treated but the new test allows this to happen at an earlier point in the disease process (B). Benefit also occurs for those with no disease and for whom a negative monitoring test result has a reassurance value, increasing a patient’s sense of control over the disease (G). Positive benefits might also be experienced by all patients, for example with the use of a less invasive test or less frequent testing (H).

Patients who have a ‘false’ monitoring test result will experience harm from the new testing strategy in a similar manner to that experienced after false-negative or false-positive diagnostic test results. False-negative test results can lead to a false feeling of security, delayed detection of disease and potentially a delay in effective treatment until the disease becomes clinically apparent (D). False-positive test results can lead to unnecessary further investigation and/or unnecessary treatment (E). For monitoring tests that aim to detect preclinical or very-early-stage disease, ‘early’ false-positive results (i.e. in patients whose disease would not have progressed to clinically overt disease within a clinically meaningful time frame) will lead to a longer period of time in a diseased state and potentially in overdiagnosis and unnecessary treatment (F). Similarly, ‘early’ positive monitoring tests in those with a true-positive test result (C) can cause harm for those patients who go on to experience a poor clinical outcome and who undergo a longer period of treatment (with associated side effects and a longer period of time in a ‘diseased’ state).

More direct harms can also be incurred by the testing experience, relating to repeated applications of the monitoring test, to any confirmatory testing and to any intervention that is implemented (I and J). These can be physical or psychological in nature and may result from either positive or negative test results, with the serial nature of testing in a monitoring context necessarily multiplying the potential impact. In particular, the ongoing monitoring process (even with repeated negative results) may raise general levels of anxiety and distress and can also have a ‘labelling’ effect that can have a negative influence on patients’ perceptions of themselves and their disease.319,320

When is randomised evidence needed?

The RCT is the gold standard approach to assessing impact on patient outcomes but, given the challenges to implementing this design for the assessment of test impact, its use requires careful consideration.

Lord et al.318 determined the need for randomised evidence for a new diagnostic test based, first, on whether or not the cases detected by the new test represent a similar spectrum of disease to those detected by the old test and, second, whether or not treatment has been shown to be, or can be assumed to be, as effective in the new group of patients, regardless of disease spectrum. Incorporating the time dimension of monitoring into this framework, for a new monitoring strategy one must determine, first, whether or not the cases detected by the new strategy represent a similar spectrum of disease, both in terms of the biological characteristic that is measured by the test and in terms of the time point in the disease process at which disease recurrence or progression is identified, and, second, whether or not treatment has been shown to be, or can be assumed to be, as effective in the new group of patients, regardless of disease spectrum and the timing of detection in relation to the stage of disease.

Notwithstanding the simple appeal of this approach, testing strategies are necessarily complex interventions, with various components and possible interactions that can combine to affect patient outcomes; even a ‘perfect’ test and highly effective treatment will not necessarily improve patient outcomes. Ferrante di Ruffano et al.68 developed a framework to consider how testing can affect health outcomes. This has been adapted for the monitoring context in Table 13, considering factors such as timing, test properties, treatment effectiveness, potential to change practice and the patient experience.68 Many of these can apply equally to patients undergoing monitoring; however, some require more or less emphasis or need to be adapted to the monitoring context.

TABLE 13

TABLE 13

Patient outcome framework for monitoring tests

We considered the extent to which these factors might affect a monitoring evaluation according to the aim of the new strategy and the way in which it fits with standard care, to demonstrate how this could inform the need for randomised evidence (Table 14).

TABLE 14

TABLE 14

Analysis of the need for randomised evidence for a new monitoring strategy

Earlier detection and treatment

Evaluations of monitoring strategies that aim to detect and treat disease at an earlier stage or time point will generally take the form of a new monitoring regime compared with an existing monitoring regime or the initiation of monitoring when no previous monitoring was undertaken. In this setting, the goal of earlier detection almost necessarily implies that those patients detected and treated by a new monitoring strategy will have a different spectrum from those previously treated. Although patient outcomes may be affected by all of the identified mechanisms, it is the clinical validity of the test(s) used, their ability to detect long-term change within a clinically meaningful time frame and the effectiveness of the treatment at the particular stage of disease that are of overarching importance. Longitudinal studies can establish test properties and identify the spectrum of patients detected by the test(s). If no randomised evidence for treatment in this group of patients exists or if it is not clear whether or not the existing evidence will apply in the new group of patients, a new RCT may be needed, as for example in a trial evaluating a lower CD4 threshold for the initiation of antiviral treatment in HIV infection.157 Alternatively, the evidence may be such that no trial is indicated. A trial of ultrasound to detect small HCCs in patients with cirrhosis of the liver found that many of the lesions detected were too small to warrant treatment, with some even regressing rather than progressing.163 This high rate of ‘early false-positive’ results could potentially have been identified in a longitudinal study without the need for a RCT.

Even when clinically valid, timely monitoring tests and effective treatments are available, the potential for a monitoring strategy to have a positive impact on patients will be influenced both by the degree to which it can substantively add clinical value over and above that of usual clinical practice and by the degree of clinical and patient confidence in the new strategy. When an individual monitoring test provides a clear guide to future management, as in the CD4 example above, the added contribution of the change in strategy might be relatively easy to discern. However, when a new monitoring test is one component of a bigger surveillance programme, for example the addition of biochemical tests and/or imaging tests to an existing surveillance programme, its added value may be more difficult to ascertain. In some circumstances, the accuracy of the confirmatory test could also mitigate any impact of the monitoring test, especially if the new test is able to detect very-early-stage disease that is not detectable by the confirmatory test.182184,321

Clinician and patient confidence in and compliance with a prescribed monitoring strategy can be vital to the success of a monitoring evaluation. For example, if clinicians have a high degree of faith in a new test, its ‘off-protocol’ use in the control arm of a trial may dilute the observed effect. This was observed in a trial of ultrasound for the detection of HCC, which also attempted to evaluate the added value of serial alpha-fetoprotein (AFP) measurements; high rates of serum AFP assay use in the two groups not randomised to AFP (60.5% and 54.8%, respectively) precluded reliable interpretation of the data and led to a final analysis restricted to ultrasound randomisation only.163

The effect of the patient experience of monitoring on outcome will depend on the nature of the test(s) involved and of the care that would otherwise have been received, especially if no monitoring was previously carried out. In the TOMBOLA (Trial Of Management of Borderline and Other Low-grade Abnormal smears) trial, for example, non-attendance was higher in the cytological surveillance arm: 10.6% did not attend the first cytological surveillance appointment whereas 6.8% did not attend for immediate colposcopy.178 Of the 10% who did not attend the first cytological surveillance appointment, 2% did not attend at all and 8% attended > 6 months after the test was due. Some tests (or subsequent confirmatory testing) will carry a risk of immediate or long-term physical harm, for example the introduction of routine endoscopic surveillance with biopsies in patients with Barrett’s oesophagus to allow the earlier detection of oesophageal carcinoma or the initiation of 6-monthly computed tomography (CT) in patients at risk of colorectal cancer recurrence, with the potential to affect patient compliance.176 Strategies that afford patients more control over a disease, however, might increase adherence to treatment regimens, as in a trial of daily foot skin temperature monitoring by patients with diabetes mellitus in which those who were compliant with the monitoring strategy for at least 50% of the time were significantly less likely to develop a foot ulcer.174

Patients will also experience a sometimes complex psychological impact from the monitoring experience. A qualitative study of patients’ attitudes to and understanding of CA-125 testing for the monitoring of ovarian cancer recurrence found that both negative and positive test results can reassure patients if the result appears to legitimise patients’ own subjective experience of the disease, with a positive result confirming their worst suspicion or a negative result providing reassurance that they are as ‘well’ as they feel.322 Positive, or rising, test results can mediate a patient’s experience of any clinical symptoms, with the knowledge that he or she is probably experiencing a relapse potentially making symptoms more unbearable or causing him or her to reinterpret earlier ‘symptoms’ that had previously been discounted.322 Similar findings in other monitoring situations have also been observed.320

Reduce the invasiveness of testing

Patients’ exposure to invasive tests can be reduced in two ways: (1) by replacing an existing test with a less invasive one or (2) by introducing a triage test to select the most appropriate patients for invasive testing. For the former option, assuming that the new test aims to detect disease at a similar stage or time point, it will be important to establish its properties in relation to the existing test and to ensure that it identifies a similar spectrum of patients. If so, treatment can be expected to be similarly effective and the impact of the test on patient outcomes will result from the less invasive nature of the test, as for example with the use of a less invasive biopsy approach in patients at risk of colorectal cancer.177

If a new triage test is introduced to prevent harm from the testing process, such as gene expression profiling as triage for endomyocardial biopsy when monitoring for acute cardiac transplant rejection162 or the introduction of cytological surveillance to reduce the number of colposcopies undertaken in women with mild dyskaryosis,178 the need for a randomised trial will be greater. Only a subgroup of patients will be selected for further investigation and treatment, such that treatment effectiveness could vary and outcomes in those no longer selected for treatment must also be assessed. In this setting, the properties of the new test (in particular the rate of false-negative results), the timing of detection and treatment effectiveness are likely to be of overarching importance. Patient and clinician confidence in the new triage test will also be needed to ensure that any potential benefits are realised in practice. An element of ‘trust’ that the new, less invasive test will accurately identify those in need of the more invasive test is required on both parts. Patient preference is not always as intuitive as it might appear, for example gene testing was favoured over endomyocardial biopsy in the example above,162 whereas patients at risk of bladder cancer recurrence preferred an immediate result from a more invasive test (cystoscopy) to the result from a less invasive (urine) test a week later.179

Reduce the volume of testing

Sometimes new monitoring strategies aim to reduce the number or frequency of tests without adversely having an impact on patient outcome, as is often the goal when monitoring for cancer recurrence.180,184 One example is the proposal to reduce the number of CT scans from five over a 3-year period to two in the first year of follow-up following primary treatment of non-seminoma testicular cancer.180 Under the new regime, there could be a change in spectrum to later-stage disease in those detected at the 12-month follow-up point and detection of those patients whose disease recurs beyond 12 months would rely on clinical relapse or detection by biochemical markers or chest radiography. In this situation, the number of cases that would be missed and the stage of disease at detection could be identified using a longitudinal study design but the clinical outcomes for the different strategies could not be directly compared without a RCT. If, however, sufficient evidence exists for treatment at the various stages of disease, this could be linked to data from a longitudinal study using a decision-analytic-type model.318

A trial of less frequent fetal surveillance of small-for-gestational-age fetuses demonstrates the sometimes complex responses of clinicians and patients to monitoring. Over half of the experimental group in this trial attended for ultrasound more frequently than scheduled and underwent additional tests of fetal well-being, suggesting that clinicians were not always comfortable with the planned reduced frequency of fetal surveillance and making it difficult to assess whether or not the apparent safety of less frequent monitoring may have been in part because of this additional surveillance.158 At the same time, 17% of women in the twice-weekly surveillance group attended less frequently than requested, suggesting a patient perception of over-frequent monitoring.185

Reduce overtreatment

In some monitoring contexts, a new test can be introduced to replace another simply to better select the right patients, for example a new deoxyribonucleic acid (DNA)-based test to detect human cytomegalovirus infection following stem cell transplantation compared with the existing antigenemia test.167 This is more akin to a diagnostic test context in which the goal is to use the most sensitive and/or specific test, as the time dimension of monitoring is less relevant than the properties of the test concerned. Although the new test may detect a different biochemical marker, if there is no indication of a differential treatment response according to the marker used, and randomised evidence exists for the effectiveness of treatment in the patients identified, then a study to establish the properties of the tests concerned may be sufficient evidence for the introduction of the new test.

To fully affect patient outcomes, however, clinicians must have confidence to act on the results of the new test. Only half of patients who tested positive on the DNA-based test in the example above actually underwent treatment, whereas, in a trial of a new galactomannan assay in patients at risk of invasive aspergillosis following stem cell transplantation, two-thirds of those treated in the experimental arm had a negative monitoring test, perhaps because clinicians had previously relied on clinical assessment as the basis for treatment decisions.167,168

Delay or avoid treatment when it is not required

The final scenario is one in which monitoring is introduced as an alternative to immediate treatment. In this circumstance, patient outcome might be affected by later treatment in the surveillance arm, further delays to necessary treatment because of false-negative test results, the effectiveness of treatment at a later stage of disease and the need for clinician and patient confidence in the monitoring regime. When immediate treatment is the standard care option, evidence for treatment at a later stage of disease may not be available, so the onus is not only on demonstrating that the monitoring test used is clinically valid and able to detect the point at which treatment is needed, but also on evidencing treatment effectiveness in the surveillance group. In a small trial in infants with mild hip dysplasia, sonographic surveillance allowed abduction treatment to be delayed or avoided with no significant difference in radiological outcomes at 1 year compared with immediate abduction treatment; < 50% of those in the surveillance arm underwent treatment during the course of the trial.158

Conclusion

The impact of a monitoring strategy is driven not only by the properties and timing of testing and the effectiveness of treatment but also by patients’ responses to the type and frequency of testing and clinicians trust in, and willingness to comply with, the monitoring protocol. Reitsma et al.323 advocate that what is more important for clinical decision-making than the level or change in a given marker is the confidence with which that marker can be used to inform patient management. A move towards a test validation paradigm is advocated, by which a number of methods (including establishing test properties) are used to determine whether or not the results of a test are meaningful in practice. In some circumstances, randomised evidence will be needed to fully assess the impact of a test, but this level of evidence will not be needed in every circumstance.

For example, the feasibility of testing and interpretability of test results can be estimated in the development phase of a test, as long as the technical properties of the test are established in clinically relevant populations rather than in laboratory-based studies alone.316 Patients’ interaction with the testing experience and their likely adherence to monitoring or to subsequent management can be assessed in qualitative studies. Feasibility or pilot studies can help identify what a new test adds to current clinical practice, particularly in terms of clinicians’ interaction with and likely adherence to a new monitoring strategy, and can help identify potential barriers to implementation, as has been recommended for trials of complex interventions.228 Estimates of key aspects of test performance can be obtained from non-randomised, preferably longitudinal studies comparing tests against a (delayed) reference standard or clinical outcome.65 The efficacy of treatment is the only mechanism that requires evaluation in a RCT per se; however, the combined effect of the individual mechanisms that come into play may be fully assessable only in a RCT.

Any decision to undertake a RCT should be informed at a minimum by good evidence of the natural history of the disease, the establishment of test properties (in terms of clinical validity and estimation of long-term change in disease status) and evidence (or lack) of treatment efficacy in those patients who are identified by the new monitoring strategy.

Copyright © Queen’s Printer and Controller of HMSO 2018. This work was produced by Selby et al. under the terms of a commissioning contract issued by the Secretary of State for Health and Social Care. This issue may be freely reproduced for the purposes of private research and study and extracts (or indeed, the full report) may be included in professional journals provided that suitable acknowledgement is made and the reproduction is not associated with any form of advertising. Applications for commercial reproduction should be addressed to: NIHR Journals Library, National Institute for Health Research, Evaluation, Trials and Studies Coordinating Centre, Alpha House, University of Southampton Science Park, Southampton SO16 7NS, UK.
Bookshelf ID: NBK513116

Views

  • PubReader
  • Print View
  • Cite this Page
  • PDF version of this title (92M)

In this Page

Other titles in this collection

Recent Activity

Your browsing activity is empty.

Activity recording is turned off.

Turn recording back on

See more...