Abstract
Abstract
Stepwise reinforcement learning (RL) performs an electromagnetic simulation after each structural change, so exploration in photonic inverse design is limited by high simulation costs. We propose delayed-gratification deep reinforcement learning (DG-DRL), which decouples reinforcement-learning decisions from simulation calls, so that a sequence of decisions covering all design variables requires only one simulation. The method combines variable-wise decisions, dynamic action masking, and delayed rewards. It excludes invalid options during design construction and calculates optical feedback after the complete structure is formed. We validate the method on three tasks: high-transmission structural-color films, near-field beam shaping, and freeform-grating beam steering. After 2,500 simulations, the number of valid designs and the color-gamut area obtained by DG-DRL in the 90–100% mean-transmittance range are 2.44 and 9.58 times those obtained by stepwise RL, respectively. For beam steering at 1100 nm and 70°, DG-DRL reaches a mean efficiency of 78.2% after 40,000 simulations, close to the 78.3% reported for RL in the literature after 2,000,000 simulations. The ablation further shows that introducing variable-wise decisions or delayed rewards alone reduces the mean efficiency of the baseline, whereas combining them increases it from 53.93% to 79.37%. These results show that jointly designing decision organization and reward timing can make more efficient use of simulation samples and provide an effective method for high-dimensional discrete photonic inverse design under limited simulation budgets.