Abstract
Abstract
Sustainable reservoir operation depends jointly on hydrological processes and on the quality of the institutions that manage them, yet most dynamical models of reservoir storage treat governance as fixed. We develop a two-state nonlinear model coupling reservoir volume $V$ to a governance level $s\in[0,1]$ that decays without maintenance and is raised by an effort policy $u=\pi(V,s,\xi)$, closing a feedback loop absent from earlier open-loop formulations. An availability factor makes non-negativity of $V$ a theorem rather than an assumption, correcting a defect we identify in an earlier open-loop version of the model, whose own numerical simulation lets storage drift to about $-15$ under reduced inflow; under the corrected model the same scenario instead converges to a small positive equilibrium. We establish well-posedness, a unique equilibrium under open-loop management, local stability via the closed-loop Jacobian, global asymptotic stability for scarcity-responsive policies (via Bendixson--Dulac and Poincar\'e--Bendixson), and a contraction certificate, based on a weighted max-norm and applicable to arbitrary Lipschitz policies, that yields global exponential stability under constant inputs and exponential seasonal entrainment under periodic ones. We show that a policy which instead rewards abundance of storage can produce a saddle-node bifurcation and a genuine governance-collapse trap, with two stable equilibria separated by a basin boundary that a numerical continuation and a direct search confirm. Building on the contraction certificate, we derive a sufficient Lipschitz-norm condition under which a neural-network policy, trained by differentiating through the reservoir dynamics, is provably stabilising for any time-varying inflow, release, and exogenous signal, and we complement it with a sharper numerical grid certificate. On a synthetic benchmark with seasonal, autocorrelated, drought-prone inflow, we find that reservoir governance under drought is where a learned policy's value is decided: a certified network trained on a distribution that explicitly balances nominal, harsh-drought, and mismatched-parameter conditions outperforms every classical baseline under harsher droughts by a wide, statistically significant margin, while retaining a formal stability certificate. By contrast, a certified network trained only on nominal conditions improves on the tuned constant-effort baseline by about $2\%$ within its own training distribution, at a certification cost of only about $1.8\%$ relative to an unconstrained network (which improves on it by about $4\%$), but is outperformed by the simpler classical baselines once droughts become more severe or the physical parameters are mismatched, showing that in-distribution gains do not automatically transfer under distribution shift even when stability itself remains certified. Comparing the two training regimes traces out a robustness--performance trade-off rather than a single dominant policy; scaling up the network and adding explicit hazard-detection inputs moved further along this trade-off rather than eliminating it. Gating the two networks together after training, by an interpretable and Lipschitz-bounded combination rule, yields a continuously tunable family of policies, one member of which improves on the tuned constant baseline in all three test environments and on every classical baseline under drought (while remaining significantly worse than the threshold and PI rules under nominal and mismatched conditions)--the paper's central empirical result. We discuss the limitations of a single scalar governance variable and exogenous inflow, and outline the calibration protocol required to extend the framework, developed here on a synthetic benchmark, to a real reservoir.