3.5.1. Internal Validity
All three trials were DB and placebo-controlled; LIPO-010 and CTR-1011 featured appropriate randomization and allocation concealment processes, while the study by Stanley et al. 2014 did not present relevant methodological details, thus precluding an assessment of the associated risk of bias, and leaving uncertain the validity of the results. In LIPO-010 and CTR-1011, the CT scans (for VAT) were centrally reviewed and analyzed in a blinded fashion, although no details were provided about the number of individuals who examined the images, as well as any training or calibration procedures they underwent; the uncertainty regarding the reproducibility of the findings might undermine the confidence in the results. The authors of Stanley et al. 2014 did not provide information about the manner in which the CT scans were reviewed, although they indicated that, for the measurement of VAT, previous research demonstrates that single-slice CT has an estimated correlation between repeat measurements of 0.99, with errors in precision estimated at 3.9%.5 If true, these values increase the confidence in the effects of tesamorelin on VAT.
Baseline characteristics were generally similar across treatment groups in all trials, with few differences. For instance, mean VAT levels were different across the three trials; however, discussions with the clinical expert suggested that these could be due to variations in clinical management of patients arising from temporal and/or geographical differences between the studies. Furthermore, compared with participants in CTR-1011, a smaller percentage of participants in LIPO-010 had undetectable viral loads, thus indicating that they were less healthy. It is plausible that treatment effects are artificially magnified in a less healthy population, thus making the treatment appear to be better than reality. Overall, the clinical expert consulted by CDR for the purpose of this review noted that any observed inequities between trials were minor and unlikely to substantially affect treatment response.
Both LIPO-010 and CTR-1011 used an ANCOVA model to evaluate the primary efficacy outcome; i.e., the per cent change in VAT between tesamorelin and placebo. The model included treatment as a fixed effect and baseline VAT as a covariate; in CTR-1011, an additional covariate (centre) was included in the model, although no rationale was provided. The results pertaining to this outcome presented in this report were extracted from the FDA statistical review, which appeared to generate the results (for both studies) using an ANCOVA model with treatment as fixed effect and baseline VAT as covariate. It was unclear why the FDA statistical review did not include centre as a covariate in the CTR-1011 analyses, although the results were consistent with those presented by the manufacturer.
LIPO-010 and CTR-1011 did not use a true ITT population; rather than including all randomized participants, the defined ITT analysis sets across both studies included only those participants who took at least one dose of the assigned study drug; i.e., a modified ITT population. Not adhering to the true ITT principle threatens the prognostic balance that a successful randomization process creates, thus leaving uncertain the validity of the results. Nevertheless, examining the participant disposition () suggests that relatively few participants who were randomized did not contribute to the ITT analysis set: in LIPO-010, two participants (0.73%) in the tesamorelin group, and none in the placebo group; in CTR-1011, five participants (1.82%) in the tesamorelin group, and three participants (2.3%) in the placebo group. Missing data were imputed using the LOCF method, whereby baseline values were carried forward into the treatment period. However, carrying the last observation forward may have artificially stabilized VAT levels among participants who dropped out; conversely, observed data could also be biased if the probability of withdrawal is related to an increase in VAT levels. Although the number of participants for whom data were imputed was unclear in both trials, there did appear to be a large difference in the number of individuals contributing to the ITT versus PP analysis sets, which possibly suggests a substantive amount of imputation, and ultimately may limit the validity of the results (). Further, in CTR-1011, it was unclear how missing data were handled for the analysis of EQ-5D data, which was analyzed using a linear mixed model, and for which as much as 35% of participants did not contribute to the analysis — missing such a substantive portion of the trial sample leaves uncertain the robustness of the results, and increases the uncertainty of the magnitude of the treatment effects on EQ-5D. The study by Stanley et al. 2014 indicated using a modified ITT population among participants with available baseline and six-month follow-up data. Excluding participants from analyses may not preserve the integrity of randomization, especially if the percentage of exclusions is large, thus limiting the validity of the results. In Stanley et al. 2014, five participants (17.9%) randomized to tesamorelin, and six participants randomized to placebo (23.1%) did not contribute to the primary analysis. The authors of the study imputed data for participants with missing observations, and found the results to be consistent with the analysis of available data. Still, the authors did not account for four of the 26 participants (15.4%) who were excluded after randomization to placebo, of whom three declined participation prior to the baseline visit, and one developed an exclusionary pre-condition. Not accounting for these participants might mean that, unlike the randomized set of participants, the analyzed set is no longer prognostically balanced, which could undermine confidence in the results.
In LIPO-010 and Stanley et al. 2014, the rates of discontinuation appeared to be disproportional across the treatment groups, with more participants in the tesamorelin arm discontinuing from the study than those receiving placebo. As above, these imbalances might have disrupted the prognostic balance that successful randomization creates, thus leaving uncertain the validity of the results. Further, across both studies, the most common reason for withdrawal was AE. Due to the unique AEs associated with tesamorelin — e.g., injection-site reactions, arthralgia, and myalgia — participants may have been inadvertently unblinded, which itself could have affected their behaviour in the trials, and consequently impacted their responses to subjective outcomes, such as QoL.
LIPO-010 and CTR-1011 were powered based on a difference of 8% in the change in VAT at week 26 between tesamorelin and placebo, while the study by Stanley et al. 2014 was based on an estimated difference of 16.5%. The cut-off of 8% appears to have been derived from the 2004 Forum for Collaborative HIV Research, which, according to the manufacturer, established an “expected decline” of 8% in VAT for patients with HIV-associated lipodystrophy receiving an rhGH product in clinical trials of up to 26 weeks in duration.18 A Scientific Advisory Panel convened by Health Canada acknowledged that an 8% decrease in VAT is a clinically important short-term effect, although Health Canada noted that the cut-off was “largely arbitrary,” and “reportedly a ‘hybrid’ end point derived from the VAT loss seen in a previous [phase] 2 study of the rhGH product and the 5% total weight loss target the FDA has adopted for development of drugs for treating obesity.”21 The authors of Stanley et al. 2014 did not provide a rationale for estimating a treatment effect of 16.5%. Although all studies were adequately powered to evaluate the primary efficacy outcome, none of the trials were powered to assess secondary efficacy outcomes or for harms outcomes.
In LIPO-010 and CTR-1011, to help contextualize changes in body image, the manufacturer reports having derived an MID for each of the three parameters from previous lipodystrophy studies. However, it does not provide details about the specific studies to which it referred, thus leaving uncertain the validity of the chosen MID values, as well as the degree to which the results are clinically meaningful. It also reports that the instrument used to measure body image was validated,23 although the FDA had several concerns about this claim.24 In particular, the FDA’s Study Endpoint and Labeling Development (SEALD) team concluded that the instrument had “questionable content validity,” and the team raised concerns about the instrument’s ability to measure a clinically important treatment effect. The SEALD team specifically reported that the manufacturer did not address whether qualitative research was done to evaluate patient understanding of the final instrument. As a result, the instrument did not meet the standards for instrument development as recommended by the FDA, nor did it meet the standard for evidence saturation. The SEALD team also reported that internal consistency reliability data had limited relevance because single items were utilized as end points instead of summary scales. In addition, the team stated that the enrolment criteria in the trials did not pre-specify a minimum PRO score, which could have made it more difficult to demonstrate improvement. Further, with respect to the instrument used to measure QoL, the manufacturer provided only limited information regarding the interpretation of scores, and no information about the psychometric properties of the instrument, which leaves uncertain the validity of the results.
Across LIPO-010 and CTR-1011, several pre-specified supportive analyses were conducted, including testing for numerous covariate-by-treatment interactions, although no adjustments for multiple comparisons appear to have been made for these analyses. A further limitation is that, in CTR-1011, the manufacturer appeared to conduct a post-hoc analysis examining the NNRTI-by-treatment interaction; a post-hoc analysis decreases the credibility of the analyses and limits certainty in the results. Moreover, both trials used gatekeeping testing strategies that, per the manufacturer, were created for labelling purposes and to help control the experiment-wise type 1 error rate at a one-sided level of 0.05. This is a common and appropriate strategy to account for multiplicity. Only the changes in VAT and belly appearance distress were considered in the testing strategy, which limits the ability to interpret the analyses of outcomes outside the gatekeeping procedure. Without controlling for multiplicity, analyses of all outcomes other than change in VAT and belly appearance distress, as well as the subgroup analyses, should be considered hypothesis-generating and interpreted with caution. No adjustments for multiplicity were performed in Stanley et al. 2014, including for the co-primary end point, although the authors did not evaluate outcomes of interest other than change in VAT.
3.5.2. External Validity
Discussions with the clinical expert consulted by CDR for the purpose of this review highlighted that generalizability of the findings of the three trials is a major concern. Chiefly, the studies appear to have been conducted at a time when HIV patients were commonly receiving ART regimens that were associated with accumulation of visceral fat. The expert indicated that, in today’s clinical practice, HIV patients are substantially different from those who were enrolled in the three studies, and are much more likely to receive ART regimens that consist of backbones other than PIs, including integrase strand transfer inhibitors (INSTIs) and NNRTIs. To this end, of the six regimens recommended by the US DHHS to manage ART-naive patients, five are INSTI-based, and one is ritonavir-boosted PI (PI/r)–based:11
INSTI-based regimens:
Dolutegravir/abacavir/lamivudine (
DTG/
ABC/
3TC)
DTG plus tenofovir disoproxil fumarate (
TDF)/emtricitabine (
FTC)
Elvitegravir (
EVG)/cobicistat (
COBI)/tenofovir alafenamide (
TAF)/
FTC
PI/r-based regimen:
Darunavir/ritonavir plus
TDF/
FTC.
In Canada, three of the abovementioned preferred regimens are available as single-tablet regimens (STRs): Genvoya (EVG/COBI/FTC/TAF), Stribild (EVG/COBI/FTC/TDF), and Triumeq (DTG/ABC/3TC). The three remaining preferred regimens are available as multi-tablet regimens consisting of a two-drug backbone; e.g., FTC/TDF (Truvada) plus a third drug. STRs are preferred to multi-tablet regimens given their convenience advantage, which results in greater adherence and optimal treatment response. There are two additional STRs available, both of which are NNRTI-based, although they are listed as “alternative” regimens by the DHHS: Atripla (efavirenz/TDF/FTC), and Complera (rilpivirine /TDF/FTC). In LIPO-010, the number of participants receiving ART regimens that comprised integrase inhibitors was not reported, whereas fewer than 5% of participants in CTR-1011 were receiving INSTI-based therapies. In Stanley et al. 2014, the number of participants on integrase inhibitors was unclear, as the authors reported that seven individuals in the tesamorelin group (25.0%) and six in the placebo group (27.3%) were receiving entry inhibitors and integrase inhibitors.
Moreover, as a result of the large number of treatments available today, single or multiple substitutions of the components of an ART regimen offer optimal virologic suppression with fewer AEs, such as the accumulation of visceral fat. The clinical expert consulted by CDR for the purpose of this review indicated that fewer than 5% of patients in his HIV practice are affected by serious fat deposition.
Further, although the manufacturer tested the statistical significance of treatment-by-ART regimen interactions in LIPO-010 and CTR-1011, the results may be limited by the fact that individual drugs within a class were grouped and analyzed together.
Examination of the inclusion and exclusion criteria of the included studies reveals additional factors that may limit the generalizability of the findings. First, the authors of Stanley et al. 2014 provided limited demographic information about the participants enrolled in their study, thus precluding an adequate assessment of the generalizability of the study findings to a Canadian population. With respect to LIPO-010 and CTR-1011, the exclusion of certain subgroups of participants — those with type 1 diabetes and type 2 diabetes in LIPO-010, and those treated with oral antidiabetic drugs or insulin LIPO-010 and CTR-1011 — leaves uncertain the effects of tesamorelin in patients with HIV-associated lipohypertrophy who present with certain comorbidities. Diabetes is particularly important given concerns raised by the FDA that more patients treated with tesamorelin developed glucose intolerance, and had a higher risk of developing diabetes across the clinical development program.15
Finally, the durations of all three trials were insufficient (even after considering the 26-week extension phases in LIPO-010 and CTR-1011) to adequately capture some important safety outcomes, including the occurrence of diabetes and cancer, thus leaving uncertain the long-term safety profile of tesamorelin.