When May Medicine Treat Distinguishable States as Equivalent?
Medicine can distinguish ever finer biological and clinical states. Yet the decisive question is not simply what medicine can distinguish, but what it must distinguish to make sound decisions. This Scientific Essay develops a framework for determining when treating distinguishable states as equivalent is sufficient for the decision at hand—and when doing so conceals differences that alter benefit, harm, or access to care. Drawing on examples from precision oncology, pharmacogenomics, diagnostic testing, non-inferiority trials, and biosimilar regulation, it synthesises published evidence and reports no new empirical data.
Abstract
Medicine cannot act on every detectable difference. It must group patients, biological states, interventions, and outcomes in ways that make evidence transportable, yet those same groupings can obscure differences that matter for benefit or harm. This essay develops a decision-relative account of medical classification and proposes decision-sufficient differentiation as a heuristic for deciding when distinguishable states should be separated or treated as equivalent for a specified clinical decision. Its distinctive contribution is not a new account of treatment-effect heterogeneity or clinical utility, but the coupling of two evidentiary obligations within one decision-relative framework while distinguishing non-adoption from equivalence. Preserving a distinction as an action-guiding clinical category requires affirmative evidence of added decision value, whereas treating distinguishable states as equivalent requires sufficiently precise exclusion of differences beyond a prespecified clinical margin. A refinement may nevertheless be withheld because it lacks sufficient net utility, but that conclusion should not be mistaken for evidence that the underlying states are equivalent. The argument distinguishes analytical or biological difference, prognostic association, predictive relevance, and clinical utility. HLA-B*57:01 screening before abacavir illustrates a justified split; biosimilarity illustrates a justified equivalence for a defined clinical use; and evolving HER2-low and HER2-ultralow categories illustrate how the relevance of a boundary can change with therapy, assay interpretation, trial eligibility, and regulatory context. Precision is therefore not maximal differentiation, but justified differentiation and justified equivalence.
I. Medicine Begins by Treating Some Differences as Equivalent
No clinical trial can compare two wholly identical persons. Patients differ before treatment is assigned, and medicine must nevertheless make claims that travel beyond the singular encounter. It must infer from observed patients to future ones and from controlled comparisons to decisions made under less controlled conditions. Knowledge would collapse if every detectable difference had to remain analytically separate.
Clinical categories therefore gather heterogeneous instances under usable names. Diagnoses, treatment arms, endpoints, and intervention labels all group cases that are not identical in every relevant respect. These acts of grouping are not incidental defects. They are part of how medical evidence becomes possible.
Randomisation and standardisation do not deny this heterogeneity. Randomisation makes treatment assignment independent of baseline characteristics in expectation, allowing valid group-level causal comparisons under the relevant assumptions. Average effects are indispensable summaries for decisions that must be made before all individual contingencies are known. Precision medicine proceeds from this background of group-based evidence, aiming to improve decisions by using patient information more effectively rather than abandoning population inference altogether (Kosorok & Laber, 2019).
The equivalence at work here is provisional and decision-relative. It treats cases as similar enough for a specified evidentiary purpose, not as ontologically the same. This essay develops a decision-relative account of medical classification whose contribution is not to rediscover treatment-effect heterogeneity or clinical utility, but to connect two evidentiary obligations while keeping their conclusions distinct. Preserving a distinction as an action-guiding category requires evidence that it improves the decision at hand; treating distinguishable states as equivalent requires affirmative evidence that residual differences are clinically negligible for that same choice. A coarser strategy may also be retained because the refinement lacks sufficient net utility, but this practical conclusion does not establish equivalence. The criterion linking these judgments is decision-sufficient differentiation. Yet the same act that makes evidence transportable can also erase the difference on which a decision turns.
This framing does not make biological distinctions merely instrumental. A distinction may remain scientifically important for explanation, mechanism, taxonomy, or future discovery even when it does not yet justify a separate clinical category. The narrower question addressed here is when an observed distinction should govern an available medical decision. The argument therefore evaluates clinical use, not the ontological reality or broader scientific value of the distinction.
II. When Equivalence Fails
Equivalence becomes dangerous when a shared category conceals differences in expected benefit or harm. A treatment may have a well-estimated average effect and still vary in clinically important ways across patient states. Treatment effect heterogeneity refers to non-random variation in the direction or magnitude of an expected effect. It does not imply that the effect for a single individual can be directly observed; it asks whether evidence supports different expectations across clinically meaningful strata.
The PATH statement emphasises patient-centred approaches to such heterogeneity, including estimates relevant to decisions rather than isolated subgroup findings (Kent, Paulus, et al., 2020). Relative and absolute effects must also be separated. A similar relative risk reduction can produce larger absolute benefit in higher-risk patients, while a small absolute benefit may not justify toxicity, cost, or burden in lower-risk patients. Prognostic risk models and models of predictive effect modification therefore answer different questions: prognosis concerns outcome risk regardless of treatment; prediction concerns whether the expected treatment effect itself differs across marker-defined states.
HLA-B*57:01 and abacavir provide a clear case in which grouping would fail for a specified drug decision. The action is whether to prescribe abacavir; the harm is abacavir hypersensitivity, especially immunologically confirmed hypersensitivity. HLA-B*57:01 status can be measured before exposure and can change the treatment choice. In the PREDICT-1 trial, prospective screening reduced immunologically confirmed hypersensitivity reactions to abacavir (Mallal et al., 2008). Its implication is narrower than it first appears: a negative result did not exclude every clinically suspected reaction, and the finding did not make every genetic difference clinically useful. It showed that, for this intervention and endpoint, grouping carriers and noncarriers would expose a marker-defined subgroup to avoidable harm.
The case shows how one measurable distinction can break open a treatment category when measurement, harm relevance, and action converge. It also exposes the central temptation: if a hidden difference can harm, finer subdivision begins to look like the safest remedy. But that remedy has its own failure mode.
III. When Everything Becomes Different
The response to a missed treatment-relevant difference is not weak. More measurement can make medicine safer and more discriminating. Molecular profiling, imaging, risk models, pharmacogenetics, and computational prediction can reveal variation that group averages obscure. Risk stratification, validated biomarkers, and credible effect modifiers are already central to modern biomedical research. In this broad sense, precision medicine uses population evidence more selectively rather than abandoning it (Kosorok & Laber, 2019).
The difficulty begins when possible distinctions multiply faster than the evidence available to evaluate them. Smaller subgroups yield less precise estimates; multiple comparisons make chance interactions easier to mistake for signal; flexible thresholds and overfitted models may not transport to new settings. Apparent differences may also reflect baseline risk rather than genuine treatment effect modification. The problem is not more data, but the unjustified conversion of every detectable pattern into an action-guiding category.
PATH is useful here because it does not deny heterogeneity; it asks how claims about it become credible and useful. It emphasises pre-specification where feasible, patient-centred effect estimates, absolute effects, and validation (Kent, Paulus, et al., 2020). Its elaboration also shows why risk modelling and effect modelling require care: a reliable estimate of prognosis is not automatically a reliable estimate of differential treatment benefit (Kent, van Klaveren, et al., 2020).
The dilemma is therefore symmetrical in concept. Coarse categories can hide real differences in expected benefit or harm; overly fine ones can stabilise noise into labels. The more narrowly a reference class is defined, the more patient-near an estimate may appear, yet the more fragile its evidence often becomes. The problem, then, is not how many differences can be detected, but which ones should govern care.
IV. Difference Is Not Relevance
A difference can be real without improving care. Once a state can be named, stained, sequenced, scored, or modelled, it may appear to demand clinical recognition. But medicine needs more than detectability. It needs to know what kind of relevance is being claimed.
Analytical or biological difference is the first level. A laboratory method may distinguish molecular profiles, receptor-expression patterns, inflammatory signatures, or pharmacokinetic states. Such differences may be reproducible and biologically meaningful, yet they do not by themselves show that a clinical decision should change.
Prognostic association is different. A marker may identify patients with better or worse outcomes regardless of treatment, informing surveillance, counselling, or baseline risk. But prognosis is not prediction of treatment effect. A high-risk state may have worse outcomes under every option while receiving the same relative benefit from treatment; a lower-risk state may show a different treatment effect for reasons not captured by baseline prognosis.
Predictive relevance concerns whether the expected effect of an intervention differs across marker-defined states. A predictive marker need not itself be causal; it may be a proxy or signature if its decision-relevant prediction is credible, validated, and applicable. Conversely, mechanistic plausibility alone is insufficient. A persuasive biological story may fail to improve outcomes if the marker is unreliable, the effect too small, the threshold unstable, or the decision unchanged.
Clinical utility adds the final standard. A distinction must improve a decision relative to a simpler strategy, taking account of benefits, harms, uncertainty, misclassification, downstream interventions, and cost. Decision curve analysis offers a methodological analogy: prediction models should be evaluated not only by statistical accuracy but by net benefit across clinically relevant threshold probabilities (Vickers & Elkin, 2006). Predictive performance is not identical with usefulness for decision-making.
A detectable feature should guide action only when it can be measured reliably, bears on a defined choice, changes the expected balance of benefit and harm, and adds value over a simpler strategy. The question has shifted from whether difference exists to whether using it improves choosing. Yet relevance does not reside in the marker alone; it can change when the treatment, assay, or decision changes.
V. Categories Change When Decisions Change
HER2 testing in breast cancer makes this relational character of relevance visible. Its established clinical function was to identify tumours with HER2 protein overexpression or ERBB2 gene amplification for therapies directed at HER2. Tumours scored as IHC 3+, or as IHC 2+ with gene amplification by in situ hybridisation, were distinguished from tumours without such evidence. Tumours scored as IHC 0, IHC 1+, or IHC 2+/ISH-negative were traditionally treated as HER2-negative for those classical treatment decisions. This did not mean that every such tumour lacked HER2 protein; it meant that lower expression did not define the same eligibility category for the therapies, assays, and thresholds then in use.
Trastuzumab deruxtecan changed the treatment context. In DESTINY-Breast04, patients with previously treated advanced breast cancer whose tumours met the trial definition of HER2-low—IHC 1+ or IHC 2+/ISH-negative—derived clinical benefit from trastuzumab deruxtecan compared with physician’s choice of chemotherapy (Modi et al., 2022). The result did not establish HER2-low as a timeless biological subtype. It showed that a range of results previously grouped as HER2-negative for one therapeutic purpose could become relevant for another intervention.
The 2023 ASCO–College of American Pathologists guideline update recognised the therapeutic relevance of IHC 1+ and IHC 2+/ISH-negative results in this context, while concluding that it was premature to establish HER2-low as a new independent result category or stable biological entity (Wolff et al., 2023). That caution matters because the IHC 0/1+ boundary arose from trial eligibility and treatment selection, not from a newly discovered natural discontinuity. Low-level HER2 immunohistochemistry also remains vulnerable to preanalytic conditions, staining performance, interpretation, sampling, and intratumoral heterogeneity.
DESTINY-Breast06 sharpened the point. Trastuzumab deruxtecan was evaluated after endocrine therapy in metastatic breast cancer, including tumours described as HER2-low and HER2-ultralow (Bardia et al., 2024). HER2-ultralow referred to IHC 0 with faint, incomplete membrane staining within the specified boundary. Because the primary efficacy assessment concerned HER2-low disease and the HER2-ultralow analysis was exploratory, the latter supports an assay-defined treatment context but does not by itself establish a fully resolved predictive boundary. In 2025, the U.S. Food and Drug Administration approved fam-trastuzumab deruxtecan-nxki for certain unresectable or metastatic HR-positive breast cancers that are HER2-low or HER2-ultralow after prior endocrine therapy (U.S. Food and Drug Administration, 2025). The College of American Pathologists biomarker reporting template, version 1.6.1.0, likewise distinguishes within score 0 between absent membrane staining and weak, incomplete membrane staining (College of American Pathologists, 2025).
The lesson is that clinical meaning can change without nature presenting a final new taxonomy. The biological expression did not come into being because a new drug became available. What changed was which low-level expression could guide a treatment decision. The resulting category is neither arbitrary nor simply a timeless natural kind. It depends on the relation among expression, IHC scoring, specimen and assay performance, trial eligibility, treatment mechanism, endpoint, and decision threshold.
That relation has limits. Evidence supporting HER2-low or HER2-ultralow treatment contexts does not justify limitless subdivision of the HER2 continuum, especially near thresholds exposed to measurement error and sampling variation. Yet the existence of a continuum does not eliminate the need for decision categories. The question is therefore no longer where nature has drawn the final boundary, but how medicine can justify the boundary it uses.
VI. A Heuristic for Decision-Sufficient Differentiation
A boundary cannot be justified merely because a difference exists or because a category is convenient. The framework proposed here calls the evaluative process decision-sufficient differentiation. Compared with a coarser classification, a refinement is decision-sufficient when it is reliably measurable, produces a validated change in expected benefit or harm, crosses a defensible action threshold, and yields added clinical utility that outweighs the uncertainty, misclassification risk, complexity, and cost it introduces. Justified equivalence is the provisional conclusion reached when evidence is precise enough to exclude differences beyond a prespecified clinical margin and further subdivision is unlikely to improve the decision.
This framing does not reduce biology to immediate usefulness. A distinction may remain important for causal explanation, discovery, or future intervention even if it does not yet warrant different clinical action. The narrower claim is that a classification should become more granular only when the added resolution improves the choice it is meant to guide. Decision-sufficient differentiation therefore offers not another hierarchy of evidence, but a common rule for evaluating proposed refinements, claims of equivalence, and decisions to retain coarser classifications.
The heuristic can be stated operationally. Let C₀ denote the coarser classification strategy and C₁ the proposed refinement. Let U(C) denote the expected net clinical utility of using classification C for the clinical choice under consideration. That utility may be expressed as net benefit, expected utility, or another prespecified balance of benefits and harms, and it should include testing burdens, measurement error, misclassification, delay, implementation complexity, and relevant costs. The incremental value of refinement is therefore
ΔU = U(C₁) − U(C₀).
Two quantities should be prespecified where possible. The clinical margin δ is the smallest difference in expected benefit or harm judged clinically important for the specified outcome. The utility threshold τ is the minimum incremental value required to justify the added classification. Neither is universal; each depends on the clinical context, the consequences of error, and the available alternatives.
Figure 1 translates these specifications into a sequential heuristic for determining whether a distinction should be preserved, treated as equivalent, withheld from implementation, or left unresolved.

A proposed distinction can then be evaluated through the following sequence:
- Specify the decision. Define the target population, the clinical choice, the intervention or comparison, the relevant outcomes and time horizon, the coarser classification C₀, and the proposed refinement C₁. Without an explicit comparator and decision, claims that a distinction is “clinically meaningful” remain indeterminate.
- Establish measurability. Determine whether the distinction can be identified with sufficient analytical validity, reproducibility, calibration, and transportability in the intended setting. A difference that cannot be measured reliably cannot support an operational split. This conclusion does not, by itself, establish that the underlying states are equivalent.
- Estimate decision-relevant differences. Assess whether C₁ identifies differences in expected benefit or harm that exceed the prespecified clinical margin δ. The estimate should be accompanied by an uncertainty interval and, where appropriate, external validation. Failure to detect a treatment-by-subgroup interaction does not establish equivalence (Kent, Paulus, et al., 2020). For a two-sided equivalence assessment, the uncertainty interval must lie within a prespecified equivalence region, such as −δ to +δ (Piaggio et al., 2012).
- Test the action consequence. Determine whether the estimated difference would change the preferred intervention, alter its intensity or timing, or move a patient or product across a defensible action threshold. A reproducible difference may be biologically or prognostically important without being sufficient to change the decision under consideration.
- Compare net clinical utility. Estimate ΔU after incorporating the benefits of improved targeting and the harms and burdens introduced by the refinement. The distinction is decision-sufficient only if ΔU exceeds τ and remains favorable across plausible assumptions about measurement error, misclassification, outcome valuation, and implementation.
The sequence yields four defensible conclusions. A distinction should be preserved when it is measurable, validated, action-changing, and associated with incremental utility above τ. Justified equivalence may be concluded when the available evidence excludes differences beyond δ with sufficient precision and no relevant threshold crossing remains plausible. The coarser classification may also be retained without a claim of equivalence when a supported difference does not change the preferred action or when the burdens introduced by refinement keep ΔU below τ. The evidence remains unresolved when uncertainty still permits both clinically important differences and practical equivalence, or when measurement and validation are inadequate. These conclusions must not be conflated: failure to justify a split is not affirmative justification for equivalence, and failure to implement a refinement does not show that the distinguishable states are clinically interchangeable.
The evidentiary burdens for preserving a distinction, claiming equivalence, and withholding refinement are different in form, though not necessarily in stringency. Preserving a distinction requires evidence that the added resolution changes action and produces incremental utility above τ. Justified equivalence requires evidence precise enough to exclude differences beyond δ; absence of detected heterogeneity cannot establish absence of clinically important difference. Withholding a refinement on utility grounds instead requires evidence that its expected benefits do not outweigh the measurement error, misclassification, treatment divergence, delay, inequity, complexity, or cost it introduces. The consequences of error determine how cautious medicine should be: a false equivalence may conceal preventable harm, whereas an unnecessary or unreliable split may itself expose patients to inappropriate divergence of care. Accordingly, evidentiary thresholds should be calibrated to the expected consequences of false equivalence and false differentiation rather than to a general presumption that either error is inherently safer.
The three core cases illustrate different outcomes of this logic. HLA-B*57:01 concerns patient subgroups before a drug decision and illustrates a justified split. Biosimilarity concerns therapeutic products and illustrates justified equivalence for a defined clinical use: biological products may remain analytically distinguishable while the totality of evidence excludes clinically meaningful differences in safety, purity, and potency. The FDA framework accordingly permits minor differences in clinically inactive components and relies on stepwise comparative evidence rather than molecular identity or a non-significant result in a single trial (U.S. Food and Drug Administration, 2015). HER2 presents the dynamic case: a boundary that was previously unnecessary or unresolved may become decision-sufficient when therapies, assay interpretation, eligibility criteria, or regulatory decisions alter the clinical margin, action threshold, or incremental utility.
A deliberately less tractable stress test is PD-L1 expression in non-small-cell lung cancer. The relevant decision must first be specified: for example, FDA approval of first-line single-agent pembrolizumab in defined patients uses a tumour proportion score of at least 1%, but that cutoff should not be generalized to every checkpoint-inhibitor regimen or clinical setting (U.S. Food and Drug Administration, 2019). Even within a specified use, the decision sufficiency of PD-L1 categories remains conditional. Commercial assays differ in sensitivity, and immune-cell scoring is less reproducible; Blueprint Phase 2 found close agreement in tumour-cell staining among the 22C3, 28-8, and SP263 assays but systematically lower and higher staining with SP142 and 73-10, respectively (Tsao et al., 2018). Repeat sampling can also move the same patient across clinically used cutoffs: another study classified 28 of 77 patients, or 36%, discordantly across the <1%, 1–49%, and ≥50% categories on repeat testing (Naso et al., 2021). The heuristic therefore does not dictate a single correct PD-L1 boundary. It exposes why the boundary remains contestable: measurement reliability, δ, the relevant action threshold, and ΔU vary with the assay, specimen, therapeutic agent, and clinical setting. For some specified uses, the distinction may be decision-sufficient; for others, the most defensible conclusion remains evidence unresolved.
This synthesis sits beside, rather than above, existing methods. PATH evaluates claims about treatment effect heterogeneity (Kent, Paulus, et al., 2020; Kent, van Klaveren, et al., 2020); decision curve analysis asks whether prediction produces clinical net benefit (Vickers & Elkin, 2006); and biosimilarity frameworks ask whether product differences remain clinically meaningful within a regulatory domain. None alone supplies a common heuristic for adding and removing classificatory boundaries. PATH may find a claim of heterogeneity insufficiently credible without thereby establishing equivalence, while decision curve analysis may find no added net benefit without determining whether this reflects genuine decision-equivalence or inadequate measurement and validation. Decision-sufficient differentiation adds the bidirectional question: relative to a named comparator and a defined clinical choice, is the evidence sufficient to preserve the distinction, treat the states as equivalent, retain the coarser strategy without claiming equivalence, or leave their status unresolved?
The framework remains deliberately modest. It is a conceptual heuristic rather than a validated decision rule, and the cases were selected as contrasting illustrations rather than through a systematic review. The quantities δ, τ, and ΔU cannot remove judgment; they make it explicit and contestable. The framework does not replace trials, modeling, decision analysis, guideline development, or regulatory review. A next step would be to test whether prespecifying the comparator, clinical margin, action threshold, utility threshold, and misclassification costs changes biomarker evaluation, guideline development, or regulatory decisions. Precision begins where medicine can justify the distinctions it preserves, the equivalences it claims, and the refinements it declines to implement.
VII. Precision as Justified Equivalence
Medical precision is not measured by the number of distinctions a system can record, but by the justification with which they alter care. Research must test whether proposed separations are reliable, predictive where prediction is claimed, externally valid, and useful relative to simpler alternatives. Regulation must guard both against unjustified equivalence and against needless fragmentation. Clinical decision-making must know what a category is for, because a label justified for one treatment, endpoint, assay, threshold, and population may not be justified for another. Equivalence remains provisional when interventions, measurements, or benefit–harm balances change. The question, then, is not whether medicine should individualise or standardise in the abstract, but how far a category may travel before it outruns the decision it was built to support. The most precise medicine may therefore be the one that knows not only when differences matter, but when it is justified in allowing them not to matter.
References
Bardia, A., Hu, X., Dent, R., Yonemori, K., Barrios, C. H., O’Shaughnessy, J. A., Wildiers, H., Pierga, J.-Y., Zhang, Q., Saura, C., Biganzoli, L., Sohn, J., Im, S.-A., Lévy, C., Jacot, W., Begbie, N., Ke, J., Patel, G., Curigliano, G., & DESTINY-Breast06 Trial Investigators. (2024). Trastuzumab deruxtecan after endocrine therapy in metastatic breast cancer. New England Journal of Medicine, 391(22), 2110–2122. https://doi.org/10.1056/NEJMoa2407086
College of American Pathologists. (2025, June). Template for reporting results of biomarker testing of specimens from patients with carcinoma of the breast (Version 1.6.1.0). https://www.cap.org/wp-content/uploads/2026/07/Breast.Bmk_1.6.1.0.-REL_CAPCP-1.pdf?download=true
Kent, D. M., Paulus, J. K., van Klaveren, D., D’Agostino, R., Goodman, S., Hayward, R., Ioannidis, J. P. A., Patrick-Lake, B., Morton, S., Pencina, M., Raman, G., Ross, J. S., Selker, H. P., Varadhan, R., Vickers, A., Wong, J. B., & Steyerberg, E. W. (2020). The Predictive Approaches to Treatment effect Heterogeneity (PATH) statement. Annals of Internal Medicine, 172(1), 35–45. https://doi.org/10.7326/M18-3667
Kent, D. M., van Klaveren, D., Paulus, J. K., D’Agostino, R., Goodman, S., Hayward, R., Ioannidis, J. P. A., Patrick-Lake, B., Morton, S., Pencina, M., Raman, G., Ross, J. S., Selker, H. P., Varadhan, R., Vickers, A., Wong, J. B., & Steyerberg, E. W. (2020). The PATH statement: Explanation and elaboration. Annals of Internal Medicine, 172(1), W1–W25. https://doi.org/10.7326/M18-3668
Kosorok, M. R., & Laber, E. B. (2019). Precision medicine. Annual Review of Statistics and Its Application, 6, 263–286. https://doi.org/10.1146/annurev-statistics-030718-105251
Mallal, S., Phillips, E., Carosi, G., Molina, J.-M., Workman, C., Tomažič, J., Jägel-Guedes, E., Rugina, S., Kozyrev, O., Cid, J. F., Hay, P., Nolan, D., Hughes, S., Hughes, A., Ryan, S., Fitch, N., Thorborn, D., Benbow, A., & PREDICT-1 Study Team. (2008). HLA-B*5701 screening for hypersensitivity to abacavir. New England Journal of Medicine, 358(6), 568–579. https://doi.org/10.1056/NEJMoa0706135
Modi, S., Jacot, W., Yamashita, T., Sohn, J., Vidal, M., Tokunaga, E., Tsurutani, J., Ueno, N. T., Prat, A., Chae, Y. S., Lee, K. S., Niikura, N., Park, Y. H., Xu, B., Wang, X., Gil-Gil, M., Li, W., Pierga, J.-Y., Im, S.-A., & André, F. (2022). Trastuzumab deruxtecan in previously treated HER2-low advanced breast cancer. New England Journal of Medicine, 387(1), 9–20. https://doi.org/10.1056/NEJMoa2203690
Naso, J. R., Banyi, N., Al-Hashami, Z., Zhu, J., Wang, G., Ionescu, D. N., & Ho, C. (2021). Discordance in PD-L1 scores on repeat testing of non-small cell lung carcinomas. Cancer Treatment and Research Communications, 27, 100353. https://doi.org/10.1016/j.ctarc.2021.100353
Piaggio, G., Elbourne, D. R., Pocock, S. J., Evans, S. J. W., Altman, D. G., & CONSORT Group. (2012). Reporting of noninferiority and equivalence randomized trials: Extension of the CONSORT 2010 statement. JAMA, 308(24), 2594–2604. https://doi.org/10.1001/jama.2012.87802
Tsao, M. S., Kerr, K. M., Kockx, M., Beasley, M.-B., Borczuk, A. C., Botling, J., Bubendorf, L., Chirieac, L., Chen, G., Chou, T.-Y., Chung, J.-H., Dacic, S., Lantuejoul, S., Mino-Kenudson, M., Moreira, A. L., Nicholson, A. G., Noguchi, M., Pelosi, G., Poleri, C., … Hirsch, F. R. (2018). PD-L1 immunohistochemistry comparability study in real-life clinical samples: Results of Blueprint Phase 2 Project. Journal of Thoracic Oncology, 13(9), 1302–1311. https://doi.org/10.1016/j.jtho.2018.05.013
U.S. Food and Drug Administration. (2015). Scientific considerations in demonstrating biosimilarity to a reference product: Guidance for industry. https://www.fda.gov/media/82647/download
U.S. Food and Drug Administration. (2019, April 11). FDA expands pembrolizumab indication for first-line treatment of NSCLC (TPS ≥1%). https://www.fda.gov/drugs/fda-expands-pembrolizumab-indication-first-line-treatment-nsclc-tps-1
U.S. Food and Drug Administration. (2025, January 27). FDA approves fam-trastuzumab deruxtecan-nxki for unresectable or metastatic HR-positive, HER2-low or HER2-ultralow breast cancer. https://www.fda.gov/drugs/resources-information-approved-drugs/fda-approves-fam-trastuzumab-deruxtecan-nxki-unresectable-or-metastatic-hr-positive-her2-low-or-her2
Vickers, A. J., & Elkin, E. B. (2006). Decision curve analysis: A novel method for evaluating prediction models. Medical Decision Making, 26(6), 565–574. https://doi.org/10.1177/0272989X06295361
Wolff, A. C., Somerfield, M. R., Dowsett, M., Hammond, M. E. H., Hayes, D. F., McShane, L. M., Saphner, T. J., Spears, P. A., & Allison, K. H. (2023). Human epidermal growth factor receptor 2 testing in breast cancer: ASCO–College of American Pathologists guideline update. Journal of Clinical Oncology, 41(22), 3867–3872. https://doi.org/10.1200/JCO.22.02864
Editorial Disclosure
Daniel Artur Raus is Founder and Scientific Editor of Interstitium. He took no part in the editorial assessment or publication decision for this contribution. The manuscript was handled and approved by Sandra Sáez-Morales, Associate Editor.
Contributions
Daniel Raus de Baviera: Conceptualisation, Funding acquisition, Resources, Investigation, Writing original draft.
Competing Interest
The author declares that he has no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Publisher’s Note
Interstitium remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
