4. Bayesian statistics and causal inference on clinical data

Bayesian networks and causal inference applied to clinical databases in collaboration with clinical departments.

Background

Clinical databases record many variables for each patient—age, lifestyle, medications, laboratory values, diagnoses and outcomes—and these variables influence one another in complicated ways. Two variables can be correlated without one causing the other; a third variable (a confounder) may be driving both. Distinguishing association from causation is essential if the analysis is to inform medical decisions.

Schematic Bayesian network of clinical variables with a highlighted confounding path, and a three-step flow from observed association to estimated causal effect
(A) A Bayesian network represents the dependencies among variables as a directed graph; the dashed path shows how a confounder can create a spurious association. Variable names are generic examples. (B) The graph structure tells us which variables must be adjusted for to estimate a causal effect. Conceptual drawing, not data.

What we study

Why a physics department does this

Bayesian statistics, graphical models and Monte Carlo sampling are the same mathematical tools we use in molecular simulation; applying them to clinical data is a natural extension of our expertise.

Back to Research Themes