Chen et al. (2018), 'Neural Ordinary Differential Equations,' cite Pontryagin by name for the adjoint sensitivity method and equate it with reverse-mode automatic differentiation / backpropagation
"Neural Ordinary Differential Equations" (Chen, Rubanova, Bettencourt, Duvenaud; NeurIPS 2018) models a network's hidden-state evolution as a continuous ODE and trains it by treating the solver as a black box: "We treat the ODE solver as a black box, and compute gradients using the adjoint sensitivity method (Pontryagin et al., 1962)." The method solves a second, augmented ODE backward in time to carry the cost gradient — the adjoint a(t)=∂L/∂z(t) — back through the trajectory, which the paper's own framing equates with "reverse-mode automatic differentiation (also known as backpropagation)," describing the adjoint as the instantaneous analog of the chain rule.
This is a direct, named, Tier-1 citation from a widely-cited modern deep-learning paper to Pontryagin's 1962 optimal-control apparatus — a stronger and more specific link than the general "optimal control anticipated backprop" framing the vault already carries via claim-kelley-bryson-optimal-control-precursor and claim-recht-adjoints-equivalence-not-transmission (Recht's 2016 post never names Pontryagin, only Kalman and Bryson). It corroborates, at a different source and decade, the same bridge already drawn biographically in claim-pontryagin-worked-blind-via-mothers-spoken-symbol-glosses between Pontryagin's maximum-principle lineage and the adjoint/costate machinery behind backpropagation. See also backpropagation-gap and moc-backpropagation-origins.
Source
“We treat the ODE solver as a black box, and compute gradients using the adjoint sensitivity method (Pontryagin et al., 1962).”
claude-sonnet-5 · Promotion from 10-inbox/raw/2026-07-11-hop-pontryagin-backprop-bridge.md, 2026-07-12 · raw markdown