Forecasting disease can support prevention, but it can also burden people with uncertainty and exposure. A risk score may prompt useful screening or a change in treatment before symptoms appear. The same score can also be mistaken for a diagnosis, disclosed beyond care, or used to make a consequential decision that the person cannot understand or challenge.
Predictive health data includes genetic risk estimates, laboratory trends, imaging findings, electronic-record models, and signals from consumer devices. These tools do not merely describe the present. They create claims about a possible future, often by comparing one person with patterns found in a reference population. Ethical use therefore depends on much more than statistical accuracy.
A Prediction Is Not A Diagnosis
Risk is usually conditional. It changes with age, environment, behavior, treatment, and the population in which a model was developed. Relative risk can sound dramatic while the absolute chance of an outcome remains small. A model may also rank people correctly without estimating their actual probability well, a problem known as poor calibration.
Communication should state what outcome is predicted, over what period, for which population, and with what uncertainty. It should distinguish a screening signal from a confirmed condition. The National Human Genome Research Institute notes that polygenic risk scores estimate likelihood rather than certainty and may perform differently across populations because genomic datasets have not represented all groups equally.[1]
Useful Forecasts Need An Actionable Purpose
A prediction has a stronger justification when it leads to an intervention with meaningful evidence of benefit. If a high-risk result changes nothing beyond increasing surveillance or anxiety, collecting it may offer little value. Conversely, even an imperfect prediction can be useful when it supports a low-burden, effective preventive step and the limitations are clear.
Before deployment, an institution should define the decision the model informs, the alternatives available, and the harm of false positive and false negative results. A model used to recommend a conversation requires a different standard from one that restricts treatment. The higher the stakes and the less reversible the outcome, the stronger the evidence and human oversight should be.
Data Quality Can Become A Fairness Problem
Models learn from records produced by real health systems. Missing visits may reflect lack of transportation rather than good health. Lower spending can reflect barriers to care rather than lower need. Diagnostic codes may record unequal clinical attention. If these patterns are treated as neutral measurements, a system can reproduce disadvantage while appearing objective.
Performance should be evaluated for clinically relevant groups and care settings, not only as a single average. Where sample sizes are too small for reliable estimates, that uncertainty should be reported rather than hidden. Monitoring should include who receives an alert, what action follows, and whether outcomes improve. Equal error rates alone do not prove equal benefit.
The World Health Organization’s guidance on artificial intelligence for health places autonomy, safety, transparency, accountability, inclusiveness, and sustainability among its central principles.[2] These principles connect technical evaluation with the institutions and people affected by a prediction.
Consent Must Match The Consequence
Not every clinical calculation requires a separate consent form. Yet people should not be surprised when their data is used to infer a sensitive future condition, especially if the result may be stored, shared, or returned to them. Notice should explain the purpose, data involved, possible actions, limits of the model, and whether declining changes access to care.
Some people may prefer not to receive information about a condition that cannot be prevented or treated. This right not to know is not absolute: clinicians may face exceptional situations involving serious, preventable harm to the patient or others. But convenience for a data system is not enough to override a considered preference. Options for receiving results should be documented and revisitable.
Privacy Includes Inferences
Privacy protections often focus on raw records, although a model’s output can be more revealing than any single input. A prediction about cognitive decline, substance use, mental health, or future disability may affect employment, insurance, relationships, or a person’s willingness to seek care. Governance should cover inferred data, not only the files used to create it.
Access should be limited by purpose. Retention periods, secondary research, vendor access, and downstream sharing should be specified. Institutions should test whether a supposedly de-identified dataset can be linked back to people when combined with other information. They should also avoid collecting every available variable merely because storage is inexpensive.
People Need Explanation And Recourse
A useful explanation is not a technical dump. It tells a patient or clinician which information materially influenced the result, what the score means in practical terms, where it may be unreliable, and what happens next. The U.S. Food and Drug Administration’s guidance on clinical decision support emphasizes whether a health professional can independently review the basis for a recommendation when determining regulatory treatment of software functions.[3]
People also need a way to correct input data, request human review, and contest an adverse decision. Appeals should be timely enough to matter. Logs should show when a prediction was generated, which version of the model was used, who viewed it, and how it influenced care. Without these mechanisms, nominal human oversight can become automatic acceptance.
Governance Must Continue After Launch
Clinical practice, populations, sensors, and documentation habits change. Model performance can drift even when the software remains unchanged. The National Institute of Standards and Technology’s AI Risk Management Framework organizes risk work around governing, mapping, measuring, and managing throughout the lifecycle.[4] That approach is particularly important when a tool continually learns or is deployed beyond its original setting.
Responsible oversight assigns an owner for performance, sets thresholds for pausing use, records model and data changes, and reviews incidents as well as aggregate outcomes. Independent voices should be included when the tool affects access to care or vulnerable groups. Retirement is also a governance decision: an outdated score should not persist simply because it is embedded in workflow.
A Practical Test For Responsible Prediction
Before predictive data shapes a health decision, teams should be able to answer six questions:
- Purpose: What decision will the prediction improve, and is there an effective response?
- Validity: Does it perform and remain calibrated for the people and setting in which it will be used?
- Proportionality: Are data collection and intervention proportionate to the benefit and stakes?
- Choice: What will people be told, and can they decline or limit receipt of sensitive results?
- Recourse: Can inputs be corrected and consequential decisions reviewed by a qualified person?
- Lifecycle: Who monitors drift, unequal effects, security, model changes, and retirement?
Prediction Should Expand Agency
The ethical limit of predictive health data is reached when uncertain information is treated as destiny, hidden systems make irreversible choices, or surveillance grows without corresponding benefit. Prediction is most defensible when it gives patients and clinicians time, options, and evidence for a meaningful action.
A responsible system makes uncertainty visible and remains accountable for what follows. Its success is not the number of risks detected. It is whether people receive better care without losing privacy, fair treatment, or authority over the futures the data claims to describe.
Responsible prediction depends on both a legitimate choice and defensible data boundaries. Our guide to meaningful consent in digital medicine addresses the first requirement, while control of health device data examines the second.
Sources
- National Human Genome Research Institute, Polygenic Risk Scores.
- World Health Organization, Ethics And Governance Of Artificial Intelligence For Health.
- U.S. Food and Drug Administration, Clinical Decision Support Software Guidance.
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework.