Stochastic Control and Stochastic Gradient Estimation with Applications to Kidney Transplantation

Loading...
Thumbnail Image

Files

Publication or External Link

External Link to Data Files

Date

Advisor

Fu, Michael C.

Citation

Abstract

This thesis develops methods in stochastic control and stochastic gradient estimation, with applications to kidney transplantation.

The first part studies the optimal acceptance of incompatible kidneys for individual patients with end-stage kidney disease. Modern desensitization techniques allow patients to accept incompatible kidneys, improving access to organs but increasing the risk of adverse transplant outcomes. We formulate the kidney acceptance decision as an optimal stopping problem within a Markov decision process (MDP) framework. Compared with existing models, our formulation includes donor-patient compatibility as a state variable and allows retransplantation, thereby capturing trade-offs between waiting for a more compatible organ and accepting an earlier but riskier incompatible transplant. Under interpretable assumptions, we establish control limit optimal policies that are implementable and clinically explainable, and demonstrate their performance in numerical experiments.

The second part focuses on gradient estimation for optimal stopping problems that admit control limit optimal policies, including the kidney acceptance model in the first part. We develop a smoothed perturbation analysis (SPA) estimator for the gradient of the MDP value function with respect to the control limit, enabling sensitivity analysis and gradient-based policy optimization. The estimator is unbiased and has low variance, whereas existing methods either are biased due to discontinuities in the sample value function or suffer from high variance.

The third part develops a gradient estimation method for general discontinuous sample performance functions using the multidimensional Leibniz integral rule. By applying the Leibniz rule directly to indicator-type discontinuities, or combining it with the push-out technique for more general cases, we obtain unbiased gradient estimators. The new method generalizes the generalized likelihood ratio (GLR) method and provides efficient single-run estimators for a broad class of problems where existing methods are inapplicable or inefficient. Numerical experiments demonstrate its effectiveness and robustness.

Notes

Rights