Combining Supervised Pretraining and Reinforcement Learning for Scalable Low-Thrust Trajectory Design

Loading...
Thumbnail Image

Files

Publication or External Link

External Link to Data Files

Date

Advisor

Martin, John R

Citation

Abstract

This thesis considers a hybrid machine learning training framework that combines supervised pretraining with deep reinforcement learning to approximate optimal low-thrust control for heliocentric transfer and rendezvous trajectories. The primary test case is a circular two-body transfer problem in which a continuously thrusting spacecraft is transferred between heliocentric orbits of varying semi-major axis while minimizing propellant consumption. An additional experiment is conducted for an interplanetary rendezvous scenario. Mass-optimal reference trajectories are first generated using an indirect optimal control formulation based on Pontryagin’s Maximum Principle and homotopic smoothing to obtain bang-bang thrust profiles. These trajectories are then converted into a Markov decision process dataset by mapping each state and optimal control to a normalized polar state representation and continuous action vector, and encoding them as state–action–reward–transition tuples that populate the experience replay buffer of a Soft Actor-Critic (SAC) agent. Comparative experiments against a baseline SAC agent trained from scratch show improvements in average episode reward and higher-performing controllers when using pretraining in Monte Carlo validation tests. These results demonstrate that seeding off-policy reinforcement learning with mass-optimal trajectory data is an effective strategy for improving training efficiency in reinforcement learning applied to trajectory design problems.

Notes

Rights