Full Reference Manuscript: Exact Instantaneous Superposition in Nonlinear Dynamics
Reference Derivations for Tangent Space Linearization and Integral Accumulation
Reference manuscript for the tangent-space series. Use the compact seven-part series as the canonical reading path.
Author
Dieter Olson
Published
January 1, 2026
WarningReference Status
This is the single-file full reference manuscript for the tangent-space material. It is retained for derivations, appendices, and historical traceability. The canonical rendered reading path is the compact seven-part Tangent-Space Series, which now carries the contraction, hybrid, and residual-aware topics in shorter source-backed form.
NoteThesis Abstract
Nonlinear dynamical systems are commonly characterized as systems where the superposition principle does not hold. This thesis challenges that characterization, arguing that it obscures a more precise mathematical fact: at each fixed state (and fixed contact/constraint mode), the instantaneous mapping from inputs to state derivatives is exactly linear in the tangent space, and the associated variational equations accumulate perturbations linearly along a chosen reference trajectory. We do not claim that motion trajectories superpose—nonlinear flows are generally not additive. Superposition here refers to the instantaneous tangent-space structure.
We develop this claim across three interconnected frameworks. Part I establishes that linearization is not an approximation but the exact infinitesimal structure of smooth dynamics—the derivative of a vector field is, by definition, a linear map operating on a linear tangent space. Part II shows how the linear variational system propagates perturbations along a reference trajectory. Part III connects these geometric foundations to optimal control through Hamiltonian structures and adjoint dynamics.
This unified perspective helps explain why linear control techniques can remain useful for nonlinear plants and why linearization can yield rigorous local conclusions once the hypotheses are stated explicitly. We illustrate with worked examples including the pendulum, demonstrate connections to LQR and Lyapunov stability, and delineate scope of validity. Throughout, we distinguish between local exactness (in tangent spaces) and finite-scale approximation limits (due to curvature, mode changes, and model scope), with residuals treated as geometric terms that must still be bounded in context.
Throughout this thesis, “exact” means exact in the infinitesimal limit. The tangent space captures the system’s instantaneous behavior with complete mathematical fidelity — this is the definition of the derivative, not an approximation. For finite perturbations, second-order residuals arise with magnitude \(O(\|\delta x\|^2)\). Part IV: Residual-Aware Control discusses how such residual terms might be bounded and used in algorithm design when the needed smoothness and estimation assumptions hold.
Part I: Geometric Foundations and Local Dynamics
1 Introduction and Motivation
1.1 The Central Thesis
Nonlinear dynamics are often described as systems in which superposition does not apply. Students of differential equations learn early that while linear systems permit the combination of solutions—if \(x_1(t)\) and \(x_2(t)\) are solutions, then so is \(\alpha x_1(t) + \beta x_2(t)\)—nonlinear systems deny this convenience. The narrative typically stops there, leaving the impression that nonlinearity and superposition are fundamentally incompatible.
This statement, while true in a global sense, obscures a more precise and mathematically exact fact:
ImportantCentral Claim
At each fixed state and fixed contact/constraint mode, nonlinear systems admit exact linear structure in the instantaneous tangent-space mapping; along a chosen reference trajectory, the associated variational equations accumulate perturbations linearly.
WarningScope: No Trajectory Superposition
We will never claim that motion trajectories superpose in nonlinear systems. Trajectory-level additivity is not implied—nonlinear flows generally do not superpose, as established in standard treatments (e.g., Slotine & Li, Applied Nonlinear Control, §3.1). Superposition here refers to instantaneous tangent-space objects (velocity and acceleration) and, in mechanics, to the linearity of accelerations with respect to applied generalized forces at a frozen state and fixed contact mode. The tangent space moves with the state, so finite-time trajectories are not additive.
NoteTerminology: “Tangent Hyperplane” as Affine Subspace
The term “tangent hyperplane” is used colloquially throughout this series to evoke the geometric intuition of a flat structure touching a curved space. More precisely, the relevant geometric object is an affine subspace of the tangent space. When the input distribution \(\text{span}\{g_i(x)\}\) has codimension greater than 1, the reachable set within the tangent space is an affine subspace, not strictly a hyperplane (which implies codimension 1). We retain the term for accessibility while noting this distinction for mathematical precision.
The purpose of this thesis is to develop that statement rigorously. The differential-geometric core is not heuristic, but its mechanics interpretations still depend on the model class, smoothness assumptions, and contact mode being held fixed. Within that scope, the argument rests on the same calculus and variational machinery used throughout modern nonlinear dynamics.
WarningWhat This Is NOT
This is not “another linearization approximation paper.”
We are not claiming that linearization is “good enough” for nonlinear systems. We are claiming something much stronger and more precise:
The tangent space is exact (not approximate) infinitesimal structure—it is the Fréchet derivative, which is the unique best linear approximation by definition
Superposition is exact (not approximate) in tangent spaces because tangent spaces are vector spaces
Approximation enters only when moving between tangent spaces—residuals measure manifold curvature, not “linearization error”
Key distinction: Standard linearization says “we accept error for tractability.” This thesis says “we exploit exact local geometry and quantify transport costs.” The exactness comes from geometry, not from “being close to the operating point.”
This perspective has profound implications:
It explains why linear control techniques such as LQR and pole placement succeed in practice, even when applied to inherently nonlinear plants
It clarifies why stability analysis via linearization (Lyapunov’s indirect method) yields rigorous conclusions about nonlinear behavior
It provides the geometric foundation for modern nonlinear control theory, where trajectories are curves on manifolds with tangent spaces attached at each point
It establishes that residuals from additivity failures are genuine geometric features (curvature), not modeling errors
1.2 Structure of This Thesis
This thesis develops the “exact superposition in the infinitesimal” framework across three interconnected parts:
Part I: Geometric Foundations and Local Dynamics - Chapter 1: Tangent spaces and the exact nature of differentiation - Chapter 2: Time evolution through moving tangent spaces - Chapter 3: Local optimal control in tangent space (LQR, Lyapunov) - Chapter 4: Residuals as manifestations of manifold curvature
Part II: Integral Superposition and Transport - Chapter 5: Integration preserves linear accumulation of variations - Chapter 6: State transition operators and variational dynamics - Chapter 7: Physical interpretations (impulse, work) as integral consequences - Chapter 8: Global residuals and error quantification
Part III: Hamiltonian Structure and Optimal Control - Chapter 9: Variational principles and Hamiltonian formulation - Chapter 10: Adjoint (costate) dynamics as backward sensitivity integration - Chapter 11: Control synthesis via DDP, iLQR, and MPC - Chapter 12: Worked examples and case studies
Each part builds on the previous, progressing from pure geometric foundations through integral accumulation to practical optimal control applications.
1.3 Scope and Prerequisites
1.3.1 Mathematical Prerequisites
This thesis assumes familiarity with: - Multivariable calculus: Partial derivatives, gradients, Jacobians - Linear algebra: Vector spaces, linear maps, eigenvalues - Ordinary differential equations: Existence and uniqueness theorems - Basic functional analysis: Normed spaces, limits, \(o(\cdot)\) and \(O(\cdot)\) notation
The thesis introduces, as needed: - Differential geometry concepts (manifolds, tangent bundles) - Variational calculus (Euler-Lagrange equations, Hamiltonian mechanics) - Control theory foundations (state-space representation, LQR, stability)
1.3.2 Scope of Validity
The mathematical framework developed here applies to:
✅ Smooth nonlinear systems: Vector fields \(f(x,u)\) that are continuously differentiable (\(C^1\))
This includes: - Mechanical systems (Newtonian, Lagrangian, Hamiltonian) - Robotic manipulators and vehicles - Aerospace systems (aircraft, spacecraft) - Chemical process control - Power systems and electrical machines
The framework requires modification or does not apply to:
❌ Discontinuous systems: Switching controllers, relay systems, variable-structure control ❌ Impact dynamics: Collisions, rigid body contact with instantaneous velocity changes ❌ Coulomb friction: Stick-slip dynamics with set-valued force laws ❌ Hybrid systems: Discrete state transitions combined with continuous dynamics ❌ State constraints: Hard boundaries with inequality constraints
For these cases, generalizations exist (differential inclusions, Filippov solutions, complementarity systems, hybrid automata) but lie outside this thesis scope.
2 Tangent Spaces and Exact Differentiation
2.1 Nonlinear Dynamics and Smoothness
Consider a nonlinear dynamical system described by the ordinary differential equation:
\[
\dot{x} = f(x,u), \quad x \in \mathbb{R}^n, \; u \in \mathbb{R}^m
\] {#thesis-eq-dynamics}
where: - \(x\) represents the state vector (positions, velocities, temperatures, concentrations, etc.) - \(u\) represents the control input or external forcing - \(f: \mathbb{R}^n \times \mathbb{R}^m \to \mathbb{R}^n\) is the vector field encoding system physics
The critical assumption is that \(f\) is continuously differentiable (\(C^1\)) in both arguments. This means:
\(f\) is continuous
All partial derivatives \(\frac{\partial f_i}{\partial x_j}\) and \(\frac{\partial f_i}{\partial u_k}\) exist
These partial derivatives are themselves continuous functions
NoteWhy \(C^1\) Smoothness Matters
This assumption is minimal—it represents the bare mathematical requirement for derivatives to exist—yet physically reasonable. Virtually all systems derived from classical mechanics satisfy this smoothness condition. The laws of physics, expressed through Newton’s equations, Lagrangian mechanics, or thermodynamic principles, generically produce smooth vector fields.
What does smoothness buy us? It guarantees that at every point in state-input space, the vector field \(f\) has a well-defined derivative. This derivative is a linear map—and it is through this linear map that superposition re-enters the picture.
2.2 Derivatives as Linear Maps
2.2.1 The Fréchet Derivative
In single-variable calculus, the derivative at a point is a number—the slope of the tangent line. In multivariable calculus, this generalizes: the derivative of a vector-valued function is not a number but a linear map. This is the Fréchet derivative, the mathematically precise formulation of differentiation in higher dimensions.
For our vector field \(f\), the Jacobian matrix \(A = \frac{\partial f}{\partial x}(x_0, u_0)\) is defined as the unique linear map satisfying:
\[
f(x_0 + \delta x, u_0) = f(x_0, u_0) + A \cdot \delta x + o(\|\delta x\|)
\tag{1}\]
The notation \(o(\|\delta x\|)\) denotes terms that vanish faster than \(\|\delta x\|\) as \(\|\delta x\| \to 0\). Precisely:
This captures what we mean by “the error becomes negligible”—the remainder term becomes arbitrarily small relative to the perturbation size.
ImportantExactness in the Limit
Equation Equation 1 states that near the point \((x_0, u_0)\), the nonlinear function \(f\) is exactly equal to a linear function plus terms that become arbitrarily small relative to the perturbation. This is not an approximation we impose for convenience—it is the definition of the derivative.
2.2.2 Directional Derivatives
An equivalent, perhaps more intuitive, definition uses directional derivatives. For any direction \(h \in \mathbb{R}^n\):
\[
A \cdot h = \lim_{t \to 0} \frac{f(x_0 + th, u_0) - f(x_0, u_0)}{t}
\tag{3}\]
This limit computes the rate of change of \(f\) as we move from \(x_0\) in the direction \(h\). The crucial point: this rate of change depends linearly on the direction \(h\).
If we double the direction vector, the rate of change doubles: \(A \cdot (2h) = 2(A \cdot h)\) (homogeneity)
If we add two direction vectors, the rates of change add: \(A \cdot (h_1 + h_2) = A \cdot h_1 + A \cdot h_2\) (additivity)
These are the defining properties of a linear map. This is not an approximation or simplification—it is a mathematical theorem following from the definition of the derivative and the smoothness of \(f\).
2.2.3 The Jacobian Matrices
At a fixed operating point \((x_0, u_0)\), we compute two derivative matrices:
These matrices have specific interpretations: - \(A\): State Jacobian (\(n \times n\))—how each component of \(\dot{x}\) depends on each state variable - \(B\): Input Jacobian (\(n \times m\))—how \(\dot{x}\) responds to control inputs
Together, \((A, B)\) define the tangent space linearization at that point.
2.3 Tangent Spaces: The Geometric Picture
2.3.1 Manifolds and Tangent Bundles
To understand “tangent space linearization,” we invoke differential geometry. A smooth manifold is a space that locally resembles Euclidean space—think of Earth’s surface, which is curved globally but flat in small neighborhoods.
At each point \(x_0\) on a manifold \(\mathcal{M}\), there exists a tangent space\(T_{x_0}\mathcal{M}\): the collection of all possible velocity vectors for curves passing through that point. For a sphere, the tangent space at any point is a flat plane touching the sphere at that point.
For dynamical systems with states in \(\mathbb{R}^n\), the tangent space at \(x_0\) is itself an \(n\)-dimensional vector space. This tangent space is where linearization lives.
2.3.2 The Tangent Space Is Linear by Construction
When we perturb the state by \(\delta x\) and the input by \(\delta u\), the resulting perturbation to \(\dot{x}\) is governed by:
\[
\delta \dot{x} = A \, \delta x + B \, \delta u
\tag{6}\]
This equation deserves careful attention. It states that in the tangent space at a fixed state, the dynamics are exactly linear. The perturbation \(\delta \dot{x}\) depends linearly on \(\delta x\) and \(\delta u\), with coefficients given by the Jacobian matrices. Note that the first-order differential (linearization) is exact at the point; finite steps accumulate higher-order terms.
ImportantNot an Approximation
This is not an approximation of the nonlinear system—it is the differential structure of the system at that state. The linearization captures the infinitesimal behavior exactly, in the same sense that the derivative of a curve gives its exact instantaneous velocity.
2.3.3 Visualization: The Moving Tangent Plane
Imagine a hiker walking along a mountain ridge: - At each point, the ground is locally flat (the tangent plane) - But the tangent plane rotates and tilts as the hiker moves - The mountain is curved even though every tangent plane is flat
This is precisely the geometric picture for nonlinear dynamics: - The state space is a manifold (possibly curved, complex topology) - At each point there is a tangent space where dynamics are exactly linear - As a trajectory evolves, it passes through a continuous family of tangent spaces
2.4 Example: The Simple Pendulum
2.4.1 System Dynamics
Consider the simple pendulum—a mass hanging from a rigid rod of length \(L\), swinging under gravity. Let: - \(\theta\): angle from vertical (downward) - \(\omega = \dot{\theta}\): angular velocity - \(g\): gravitational acceleration
This is manifestly nonlinear: the restoring torque depends on \(\sin\theta\), not \(\theta\) itself. The \(\sin\) function prevents global superposition—\(\sin(\theta_1 + \theta_2) \neq \sin\theta_1 + \sin\theta_2\).
2.4.2 State-Space Representation
Rewrite as a first-order system. Define state vector \(x = [\theta, \omega]^T\):
This is the harmonic oscillator equation. Every physics student knows this as the “small angle approximation”—but our analysis reveals it as something more fundamental.
2.4.5 Eigenanalysis
The eigenvalues of \(A\) are found from \(\det(A - \lambda I) = 0\):
These are pure imaginary numbers, indicating: - Oscillatory behavior with no damping - Natural frequency\(\omega_n = \sqrt{g/L}\) - Neutrally stable equilibrium (neither attracting nor repelling)
These eigenvalues exactly characterize the local behavior of the nonlinear system near equilibrium. For small perturbations, the linear and nonlinear trajectories are indistinguishable to first order.
2.4.6 Where Nonlinearity Emerges
The nonlinearity of the full pendulum—the amplitude-dependent period, the possibility of full rotations—emerges only as perturbations grow large enough that higher-order terms in the Taylor expansion of \(\sin\theta\) become significant:
Homogeneity: \(B(\alpha \, \delta u) = \alpha \, B \delta u\) for scalar \(\alpha\)
ImportantExact Superposition Theorem (Fixed State & Fixed Mode)
Superposition holds exactly for the instantaneous mapping to the tangent space at a fixed state \((x, \text{mode})\). If we apply two infinitesimal input perturbations \(\delta u_1\) and \(\delta u_2\) simultaneously, the resulting state perturbation is exactly the sum of the perturbations that would result from applying each input separately:
There is no interaction, no coupling, no nonlinear mixing—at the infinitesimal level. The tangent space moves with the state, so trajectories are not additive.
This is not an approximation we impose for convenience. It is a mathematical theorem, inherited directly from the properties of matrix multiplication and the definition of the derivative.
ImportantInstantaneous Force–Acceleration Superposition (Fixed State & Fixed Mode)
For a mechanical system with equations \(M(q)\ddot{q} + h(q, \dot{q}) = \tau + J^\top f\), at a fixed state \((q, \dot{q})\) and fixed contact/constraint mode \(m\), the acceleration is:
which is affine/linear in applied torques \(\tau\) and generalized forces \(f\). Therefore, instantaneous acceleration contributions from separate force/torque components add exactly.
However, mode switching (contact/friction transitions), actuator saturation, and inequality constraints (unilateral contacts) make the mapping only piecewise linear and can break global superposition. The claim is valid for fixed \((q, \dot{q}, m)\)—not universally across mode boundaries.
2.5.2 Counterfactual Decomposition
At a fixed time, suppose we apply multiple input perturbations \(\delta u_1, \delta u_2, \ldots, \delta u_k\). We can meaningfully ask: how does each input contribute to the resulting state change?
In the tangent space, the answer is unambiguous. Each input perturbation \(\delta u_i\) produces a corresponding state rate change:
\[
\delta \dot{x}_i = B \delta u_i
\]
And the total state rate change is simply the sum:
This counterfactual decomposition allows us to attribute portions of the rate of change to individual inputs. The contributions are additive and non-interacting.
NoteValidity Domain
Superposition holds for the instantaneous mapping to the tangent space at fixed \((x, \text{mode})\); the tangent space moves with the state, so trajectories are not additive. When we integrate over finite time intervals, deviations from additivity appear. These deviations arise from manifold curvature: as the state evolves, it passes through different tangent spaces. The nonlinear coupling emerges from how these tangent spaces vary, not from any breakdown of local linearity.
3 Time Evolution and Moving Tangent Spaces
3.1 Why Global Superposition Fails
If instantaneous superposition holds exactly in each tangent space, why do we say nonlinear systems violate superposition? The answer lies in time evolution: the tangent space moves with the state.
The key observation: as the system evolves along a trajectory, the operating point changes—and with it, the tangent space. The Jacobian matrices are not constants; they are functions of the current state:
At each instant, the tangent space provides an exact linear description of the local dynamics. But the tangent space itself moves as the system evolves.
3.1.1 Geometric Interpretation
Return to the mountain hiker analogy: - At each point, the ground is locally flat (tangent plane) - The tangent plane rotates and tilts as the hiker moves - The mountain is curved even though every tangent plane is flat
For nonlinear dynamics: - State space is a manifold (possibly curved, complex topology) - At each point there is a tangent space (exactly linear) - Trajectories pass through a continuous family of tangent spaces
Global superposition fails not because linearity breaks down at any particular instant, but because the linear structure itself evolves. Two solutions we might wish to superpose live in different tangent spaces at each moment in time. Adding them is geometrically meaningless—like trying to add vectors from different vector spaces.
3.1.2 Nonlinearity as Tangent Space Variation
This perspective transforms how we think about nonlinearity. The “nonlinear effects” that distinguish, say, a large-amplitude pendulum swing from a small one are not violations of linearity per se. They are consequences of:
The trajectory passing through regions where tangent spaces differ significantly
The accumulated effect of these differences over time
The curvature of the manifold encoded in how \(A(t)\) and \(B(t)\) vary
Nonlinearity is relocated from the local structure (which is exactly linear) to the global geometry (how local structures fit together).
3.2 Exactness in the Limit
3.2.1 Taylor Expansion Analysis
Consider a finite timestep \(\Delta t\). The state at time \(t + \Delta t\) can be expressed via Taylor expansion:
The terms break down as: 1. Zeroth order: \(x(t)\) — current state 2. First order: \(f(x(t),u(t)) \, \Delta t\) — tangent space prediction (linear contribution) 3. Second order and higher: \(O(\Delta t^2)\) — curvature effects
3.2.2 The Limit Argument
As \(\Delta t \to 0\), the higher-order terms become negligible compared to the first-order term:
This is the precise sense in which linearization is exact in the limit. The first-order evolution—governed entirely by the tangent space dynamics—captures the behavior of the nonlinear system up to terms that become arbitrarily small relative to the timestep.
ImportantInfinitesimal Exactness
Linearization is exact in the limit \(\Delta t \to 0\). This is not a weaker form of exactness than we find elsewhere in mathematics—it is the same sense in which derivatives, integrals, and differential equations themselves are exact.
The derivative \(\dot{x} = f(x,u)\) is defined as a limit; it tells us the instantaneous rate of change, not the change over any finite interval. Linearization inherits this same infinitesimal exactness.
The tangent space is not an approximation to the nonlinear system—it is the infinitesimal structure of the system, captured exactly.
3.2.3 Variational Equation
For a perturbation \(\delta x(t)\) from a nominal trajectory \(x(t)\), the evolution is governed by the variational equation:
\[
\frac{d}{dt}\delta x = A(t) \, \delta x + B(t) \, \delta u
\tag{19}\]
This is a linear time-varying (LTV) system. At each instant, it’s exactly linear. The time-variation encodes how the tangent space changes along the trajectory.
4 Local Optimal Control: LQR and Lyapunov Stability
4.1 Connection to Linear Quadratic Regulation
The theoretical framework developed thus far is not merely academic—it underlies some of the most successful techniques in control engineering.
4.1.1 The LQR Problem
Linear Quadratic Regulation (LQR) solves the following problem:
Find: Optimal feedback control law minimizing \(J\)
Solution: Linear state feedback \(u = -Kx\) where gain \(K\) is computed by solving the algebraic Riccati equation:
\[
A^T P + P A - P B R^{-1} B^T P + Q = 0
\tag{20}\]
with \(K = R^{-1} B^T P\).
4.1.2 Why LQR Works for Nonlinear Systems
But real systems are rarely linear. Why does LQR work so well in practice? The answer lies in the exactness of tangent space linearization.
When we apply LQR to a nonlinear system:
At each operating point \((x,u)\), compute \((A,B) = \left(\frac{\partial f}{\partial x}, \frac{\partial f}{\partial u}\right)\)
The linearization captures the exact infinitesimal dynamics at that point
The quadratic cost is similarly justified: any smooth cost function is quadratic to leading order near a minimum (via Taylor expansion)
The resulting controller operates in a sequence of tangent spaces along the trajectory. At each instant, it computes the optimal action based on the local linear-quadratic structure—and this local structure is exact, not approximate.
ImportantLQR’s Success Explained
The success of LQR for nonlinear systems stems from exploiting the exactness of the tangent-space representation, not from any assumption that the true system is globally linear.
4.1.3 Gain Scheduling
Gain scheduling makes this explicit: 1. Compute LQR gains \(K(x)\) at various operating points 2. Interpolate between gains during operation 3. Controller adapts as system moves through tangent spaces
This is implicit exploitation of the moving tangent space perspective.
4.2 Lyapunov Stability Analysis
4.2.1 Lyapunov’s Indirect Method
One of the most powerful tools for analyzing nonlinear stability is Lyapunov’s indirect method:
TipTheorem: Lyapunov’s Indirect Method
Consider \(\dot{x} = f(x)\) with equilibrium \(x_e\) (i.e., \(f(x_e) = 0\)). Let \(A = \frac{\partial f}{\partial x}(x_e)\) be the Jacobian at equilibrium.
If all eigenvalues of \(A\) have strictly negative real parts (Re\((\lambda_i) < 0\)), then \(x_e\) is locally asymptotically stable
If any eigenvalue of \(A\) has strictly positive real part (Re\((\lambda_j) > 0\)), then \(x_e\) is unstable
If all eigenvalues have non-positive real parts with some on imaginary axis, stability is inconclusive (higher-order analysis required)
4.2.2 Why This Works
This is a remarkable result: we draw conclusions about a nonlinear system—which may exhibit chaos, limit cycles, or other complex behaviors—based solely on a linear analysis.
The justification is precisely the exactness we have established:
Near equilibrium, the tangent space dynamics dominate
The linearization \(\delta\dot{x} = A\delta x\) captures the essential character of nearby trajectories
Eigenvalues of \(A\) determine whether perturbations grow or decay
For small perturbations, this linear analysis is exact to leading order
The eigenvalues are pure imaginary (on the borderline), so Lyapunov’s indirect method is inconclusive. Indeed, the pendulum equilibrium is: - Stable (nearby trajectories stay nearby) - But not asymptotically stable (they don’t converge to equilibrium) - Neutrally stable—consistent with conservative mechanical system
For the inverted pendulum at \(x_e = [\pi, 0]^T\), one eigenvalue would be positive, confirming instability.
4.2.3 Connection to Tangent Space Framework
Lyapunov stability analysis via linearization succeeds because: - The tangent space at equilibrium captures exact local dynamics - Eigenvalues encode local growth/decay rates exactly - For stability, only local behavior matters (by definition)
This is another manifestation of “exact superposition in the infinitesimal” enabling rigorous global conclusions.
4.3 Practical Implications
4.3.1 Design Philosophy
The tangent space perspective suggests a design philosophy:
Locally optimal: Design controllers that are optimal in each tangent space
Globally aware: Account for how tangent spaces vary (gain scheduling, adaptive control)
Residual monitoring: Track when nonlinear effects (curvature) become significant
4.3.2 When Linear Techniques Suffice
Linear control techniques (LQR, pole placement, loop shaping) work well when: - Operating region is small (tangent spaces don’t vary much) - Trajectory stays near equilibrium (single tangent space dominates) - Curvature is mild (residuals small)
But when we integrate over finite time, deviations from additivity appear. Consider applying two inputs \(u_1\) and \(u_2\) separately, then together, over interval \([t_0, t_1]\):
Input 1 alone: \(x_1(t_1)\) starting from \(x(t_0)\)
Input 2 alone: \(x_2(t_1)\) starting from \(x(t_0)\)
Both inputs: \(x_{12}(t_1)\) starting from \(x(t_0)\)
(The \(+x(t_0)\) term corrects for double-counting the initial condition.)
For linear systems, \(r = 0\) exactly—superposition holds globally. For nonlinear systems, \(r \neq 0\) in general. The residual quantifies the failure of additivity.
5.1.2 Geometric Interpretation
Residuals arise from manifold curvature. As trajectories evolve:
They pass through different tangent spaces
The tangent spaces themselves vary in orientation and structure
The accumulated effect depends on the path taken through state space
The residual \(r\) measures this path-dependence—it is a geometric feature, not a modeling error.
ImportantResiduals as Geometric Objects
Residuals are not approximation errors or numerical artifacts. They represent genuine second-order nonlinear coupling encoded in the curvature of the state-space manifold.
5.2 Quantifying Residuals
5.2.1 Scaling Analysis
For small perturbations, we expect:
\[
\|r\| = O(\|\delta x\|^2)
\tag{21}\]
This follows from Taylor expansion: the first-order terms (tangent space) cancel exactly, leaving second-order terms (curvature).
More precisely, for perturbations of size \(\epsilon\):
where \(H_f(x) = \nabla^2 f(x)\) is the Hessian of \(f\) (encoding the second-order curvature of the vector field).
NoteInterpretation
Residuals scale quadratically with perturbation size \(\epsilon\)
They accumulate over time (integral over trajectory)
They depend on the local curvature of the vector field (\(\|H_f\|\)), not the Jacobian
For small perturbations and short times, residuals are negligible. For large perturbations or long times, they become significant.
5.2.2 Connection to Second-Order Geometry
For control systems evolving in \(\mathbb{R}^n\) (flat space), the relevant curvature object is the Hessian of the vector field \(f\), not the Riemann curvature tensor (which vanishes identically in flat space). The residual accumulation is approximated as:
where \(H_f(x)[\delta x, \delta x]\) denotes the quadratic form defined by the Hessian of \(f\) at \(x\). This formalizes the intuition that residuals measure how the vector field “curves” away from its linear approximation along the trajectory.
5.3 Practical Implications
5.3.1 Validity Regime
The tangent space linearization is most accurate when residuals are small:
\[
\frac{\|r\|}{\|\delta x\|} \ll 1
\tag{23}\]
This provides a quantitative criterion for when linear techniques suffice.
5.3.2 Adaptive Control Strategies
Monitoring residuals in real-time allows: 1. Detecting when nonlinear effects become significant 2. Adapting control gains or switching to nonlinear methods 3. Validating model-based predictions against measurements
5.3.3 Model-Based vs. Geometric Errors
Distinguishing: - Model error: Mismatch between true physics and assumed model \(f(x,u)\) - Geometric residuals: Inevitable second-order effects from manifold curvature
Even with a perfect model, geometric residuals exist. They are features of the mathematics, not deficiencies of modeling.
Part II: Integral Superposition and Transport
6 Fundamentals of Integral Accumulation
7 Integration Preserves Linearity
7.1 From Infinitesimal to Finite Time
In Part I, we established that in the tangent space at each instant:
This is exact superposition at a single moment. The natural question: what happens when we integrate over time?
The key insight: integration is a linear operation. If pointwise superposition holds at every instant, then integrated superposition holds over intervals.
7.2 The Integration Theorem
ImportantTheorem: Integral Linearity
For variations \(\delta x_1(t)\) and \(\delta x_2(t)\) satisfying:
This is the endpoint variation—the change in state perturbation from start to finish.
7.3 What Superposes: Variations, Not Causes
WarningCritical Distinction
What superposes is the variation between endpoints, not the “fractional cause” of the final state. This is a subtle but crucial point.
Consider two inputs \(u_1(t)\) and \(u_2(t)\) applied from \(t_0\) to \(t_1\): - We can exactly compute: \(\Delta x_1 = x_1(t_1) - x(t_0)\) and \(\Delta x_2 = x_2(t_1) - x(t_0)\) - If both are applied: \(\Delta x_{12} = x_{12}(t_1) - x(t_0)\) - The integrated variation equation tells us how \(\Delta x_1\) and \(\Delta x_2\) relate to \(\Delta x_{12}\)
But we cannot meaningfully say “input 1 caused 40% of the final state”—causation is not additive in nonlinear systems. Variation is additive (to leading order).
8 State Transition Operator
8.1 Solution of the Variational Equation
The variational equation:
\[
\frac{d}{dt}\delta x = A(t) \delta x + B(t) \delta u
\tag{26}\]
is a linear time-varying (LTV) system. Its solution has the form:
These properties make \(\Phi\) a linear map that “transports” perturbations from \(t_0\) to \(t_1\) while accounting for how the tangent space evolves.
\(\Phi(t_1, t_0) \delta x(t_0)\): Initial perturbation transported to \(t_1\)
\(\int_{t_0}^{t_1} \Phi(t_1, \tau) B(\tau) \delta u(\tau) d\tau\): Accumulated effect of inputs, transported from application time \(\tau\) to final time \(t_1\)
The accumulated input effect is a linear integral—each infinitesimal contribution \(B(\tau) \delta u(\tau) d\tau\) is transported to \(t_1\) via \(\Phi(t_1, \tau)\), and the results superpose exactly.
ImportantIntegral Superposition Principle
The state transition operator formulation shows that integrated variations superpose exactly through linear transport. The nonlinearity of the original system is encoded in how \(\Phi\) varies with the nominal trajectory, not in any failure of linearity within the variational dynamics.
9 Time-Invariant Special Case
9.1 Constant Jacobian
If the system linearization is time-invariant (\(A\) and \(B\) constant), the state transition operator simplifies:
\[
\Phi(t_1, t_0) = e^{A(t_1 - t_0)}
\tag{28}\]
The matrix exponential is defined via:
\[
e^{At} = I + At + \frac{(At)^2}{2!} + \frac{(At)^3}{3!} + \cdots
\]
and can be computed via eigendecomposition if \(A\) is diagonalizable.
This is the classical result from linear systems theory. The integral term shows explicitly how contributions from different times superpose—each input at time \(\tau\) contributes \(e^{A(t_1-\tau)} B \delta u(\tau) d\tau\) to the final state, and these contributions add linearly.
This tells us: if we perturb the input at time \(\tau\) by \(\delta u(\tau)\), the state at time \(t > \tau\) changes by approximately \(\Phi(t, \tau) B(\tau) \delta u(\tau)\).
11.3 Integrated Sensitivity
The total sensitivity to an input trajectory \(u(\cdot)\) over \([t_0, t_1]\) is:
This is a linear functional mapping input functions to state changes—the embodiment of integral superposition.
12 Adjoint (Costate) Dynamics
12.1 Backward Sensitivity
Forward sensitivity tells us how the final state depends on inputs. Adjoint dynamics answer the dual question: how does a scalar cost depend on the state trajectory?
This is the adjoint equation—it integrates sensitivities backward from the final time.
12.3 Physical Interpretation
The costate \(\lambda(t)\) represents the “shadow price” or “marginal value” of the state at time \(t\): - If we perturb \(x(t)\) by \(\delta x(t)\), the cost changes by approximately \(\lambda(t)^T \delta x(t)\) - Integrating backward accumulates these marginal values - At \(t_0\), \(\lambda(t_0)^T \delta x(t_0)\) gives the total cost impact of initial condition perturbations
12.4 Duality: Forward and Backward
Forward sensitivity (\(\Phi\)): How inputs propagate to states
Backward sensitivity (\(\lambda\)): How states contribute to costs
These are dual perspectives on the same system, connected through:
If the system starts at rest (\(\dot{q}(t_0) = 0\)) and \(f\) represents damping or potential forces that integrate to small values, then approximately:
Higher-order terms arise from nonlinear coupling (e.g., \(f\) depending on \(\dot{q}\)), but the impulse contributions are exactly additive.
NoteInterpretation
The integral nature of impulse guarantees linear accumulation of force effects. Each infinitesimal force contribution \(u(\tau) d\tau\) adds linearly to the total momentum change. This is integral superposition applied to mechanical systems.
15 Power-Energy: Work Integral
15.1 Power and Work
For a mechanical system, power delivered by a force is:
Again, exact additivity due to integral linearity.
However, there’s a subtlety: \(\dot{q}(\tau)\) itself depends on the forces. So while the integral is linear in the integrand, the velocity trajectory is nonlinear in the forces. The work contributions are additive, but their relationship to the final energy is not simple due to cross-coupling.
WarningVariation vs. Causation
The work done by each force adds linearly: \(W_{12} = W_1 + W_2\). But the final kinetic energy \(E_k(t_1)\) does not decompose as “fraction from force 1” plus “fraction from force 2” because the velocity \(\dot{q}\) depends nonlinearly on both forces.
This reinforces: integrated variations superpose; fractional causation does not.
16 General Principle: Integral Accumulation
Both impulse and work illustrate a general principle:
ImportantIntegral Accumulation Principle
Along a fixed trajectory\(x(t)\), for any quantity defined as an integral of a linear function of inputs:
where \(g\) is linear in \(u\) with \(x(t)\) held fixed, the contributions from different inputs superpose exactly:
\[
\Psi(u_1 + u_2) = \Psi(u_1) + \Psi(u_2)
\]
Important caveat: This superposition holds only for variations at a fixed trajectory \(x(t)\). If \(x(t)\) itself depends nonlinearly on \(u\) (as it generally does in closed-loop or trajectory optimization settings), then the principle does not hold globally — changing \(u\) changes the trajectory, and the two contributions cannot be evaluated independently.
Impulse and work are special cases where \(g = u\) and \(g = u^T \dot{q}\) respectively. The linearity of integration ensures exact additivity of these quantities, even as the underlying dynamics remain nonlinear.
17 Global Residuals and Error Quantification
18 Second-Order Residual Analysis
18.1 Residual Definition (Revisited)
From Chapter 4, the residual from applying two inputs together vs. separately is:
where subscripts denote which inputs were applied.
18.2 Scaling With Perturbation Size
Via Taylor expansion of the nonlinear dynamics, residuals scale as:
\[
\|r(t_1)\| = O(\epsilon^2)
\tag{52}\]
where \(\epsilon\) is a measure of perturbation magnitude (e.g., \(\epsilon = \max\{\|\delta u_1\|, \|\delta u_2\|\}\)).
More precisely:
\[
\|r(t_1)\| \leq C (t_1 - t_0) \epsilon^2
\tag{53}\]
where \(C\) depends on the second derivatives of \(f\) (the Hessian) and the trajectory.
18.3 Interpretation
Residuals are quadratic in perturbation size: doubling the perturbations quadruples the residual
They accumulate linearly in time: \(\|r\| \propto (t_1 - t_0)\)
They depend on curvature: regions of high curvature produce larger residuals
NotePractical Implication
For small perturbations (\(\epsilon \ll 1\)) and short times, \(\|r\| / \epsilon \ll 1\), so the linear approximation (integral superposition) is excellent. For large perturbations or long times, residuals become significant, indicating that nonlinear effects must be accounted for.
19 Connection to Hessian and Curvature
19.1 Second Derivatives
Residuals arise from second-order terms in the Taylor expansion of \(f(x, u)\). The Hessian matrices:
In practice, monitor the ratio \(\|r\| / \|\delta x\|\): - If \(\ll 1\): Linear techniques sufficient - If \(\sim 0.1\): Nonlinear effects becoming noticeable but manageable - If \(\gtrsim 1\): Nonlinear methods required
This provides a data-driven criterion for switching between linear and nonlinear control strategies.
Part III: Hamiltonian Structure and Optimal Control
WarningWork in Progress
Part III is a draft/incomplete section. The following provides the conceptual framework and key equations but requires substantial expansion — detailed derivations, complete algorithms, and worked examples are not yet present. Treat this section as a preliminary outline, not a finished exposition.
21 Variational Principles and Hamiltonian Formulation
21.1 Action Functional and Euler-Lagrange Equations
21.1.1 Lagrangian Mechanics
Classical mechanics can be formulated via the principle of least action. Define the action functional:
where \(L\) is the Lagrangian (often kinetic minus potential energy for mechanical systems).
The actual trajectory \(x^*(t)\) extremizes \(S\) subject to the dynamics constraint \(\dot{x} = f(x, u)\). This leads to the Euler-Lagrange equations.
21.1.2 Hamiltonian Formulation
Introduce the Hamiltonian:
\[
H(x, u, \lambda, t) = L(x, u, t) + \lambda^T f(x, u)
\tag{60}\]
where \(\lambda\) is the costate (Lagrange multiplier enforcing dynamics).
These are the Hamiltonian canonical equations (Pontryagin’s Maximum Principle in continuous time).
NoteConnection to Previous Content
The costate equation \(\dot{\lambda} = -\frac{\partial H}{\partial x}\) is the adjoint equation from Chapter 6. The Hamiltonian framework unifies forward dynamics (\(\dot{x}\)), backward sensitivity (\(\dot{\lambda}\)), and optimality (\(\frac{\partial H}{\partial u} = 0\)) into a single variational principle.
The costate \(\lambda_2\) evolves deterministically, leading to at most one switch for this problem. The switching curve divides state space into “apply +1” and “apply -1” regions.
21.4.3 Example 3: LQR as Hamiltonian System
Problem: Quadratic cost with linear dynamics.
\[
\min J = \frac{1}{2}x(t_1)^T Q_f x(t_1) + \frac{1}{2}\int_{t_0}^{t_1} \left[ x^T Q x + u^T R u \right] dt
\tag{81}\]
\[
\dot{x} = Ax + Bu
\tag{82}\]
Hamiltonian:
\[
H = \frac{1}{2}(x^T Q x + u^T R u) + \lambda^T (Ax + Bu)
\tag{83}\]
Optimality:
\[
u^* = -R^{-1}B^T \lambda
\tag{84}\]
Costate equation:
\[
\dot{\lambda} = -Qx - A^T \lambda
\tag{85}\]
Key insight: Assume \(\lambda(t) = P(t) x(t)\) for some symmetric matrix \(P(t)\).
Substituting into the costate equation and matching coefficients yields the Riccati differential equation:
\[
\dot{P} = -Q - A^T P - PA + PBR^{-1}B^T P
\tag{86}\]
This is the standard LQR result, derived from the Hamiltonian formulation!
ImportantConnection to Tangent Space Framework
LQR succeeds because: 1. Local linearization (\(Ax + Bu\)) is exact in tangent space 2. Quadratic cost aligns with second-order Taylor expansion 3. Riccati equation propagates cost-to-go exactly along tangent directions 4. Optimal control is linear feedback exploiting exact local superposition
21.5 Constrained Optimization
21.5.1 Inequality Constraints
Consider control bounds \(u_{\min} \leq u \leq u_{\max}\).
Local truncation error: \(O(\Delta t^5)\) (fourth-order method).
Global error: \(O(\Delta t^4)\) over fixed time interval.
NoteLinearization at Sub-Steps
For control applications, RK4 evaluates dynamics at intermediate points. When linearizing for DDP/iLQR, we linearize around the RK4 trajectory, not just the endpoints. This requires careful Jacobian computation accounting for the RK4 structure.
22.3.2 Adaptive Step Size
For systems with varying stiffness or curvature, adaptive timesteps improve efficiency.
This concentrates computational effort where curvature is high (large \(e_k\)) and relaxes where dynamics are smooth.
22.3.3 Implicit Methods
For stiff systems (widely separated timescales), explicit methods require tiny timesteps. Implicit methods allow larger steps at the cost of solving nonlinear equations per step.
where \(M\) bounds the second derivative: \(\|\frac{d^2 x}{dt^2}\| \leq M\).
Proof sketch: 1. LTE at each step: \(\|e_k\| \leq \frac{M}{2}\Delta t^2\) 2. Error propagation: \(e_{k+1} \leq (1 + L\Delta t)e_k + \frac{M}{2}\Delta t^2\) 3. Sum geometric series with \(N = (t_N - t_0)/\Delta t\) steps 4. Result: \(O(\Delta t)\) global convergence \(\square\)
22.4.3 Higher-Order Methods
RK2: \(O(\Delta t^2)\) global error
RK4: \(O(\Delta t^4)\) global error
RK45 (adaptive): \(O(\Delta t^5)\) with automatic step control
Practical rule: Use RK4 for most control problems; adaptive RK45 for high-accuracy requirements or stiff systems.
22.5 Numerical Stability
22.5.1 Stability Region
A method is absolutely stable for \(\dot{x} = \lambda x\) (test equation) if \(|x_{k+1}/x_k| \leq 1\) for all \(k\).
Euler’s method stability region: \(|1 + \lambda \Delta t| \leq 1\), i.e., \(\lambda \Delta t\) must lie in disk of radius 1 centered at \(-1\) in complex plane.
RK4 stability region: Larger (includes larger negative real values), allowing bigger timesteps for stable systems.
Implicit methods (e.g., Backward Euler): \(A\)-stable—stable for all \(\text{Re}(\lambda) < 0\), ideal for stiff systems.
22.5.2 Stiffness and Timestep Selection
A system is stiff if it contains dynamics at vastly different timescales. Example:
This matches Euler’s truncation error—both are \(O(\Delta t)\). Using RK4 reduces integration error to \(O(\Delta t^4)\), but linearization residuals remain \(O(\Delta t)\) unless we re-linearize more frequently.
Implication for iLQR/DDP: Re-linearize at each iteration using updated trajectory; per-iteration error is \(O(\Delta t)\), but iterations refine the trajectory.
Optimize locally: \(\delta u_k^* = \arg\min Q_k(\delta x_k, \delta u_k)\) where \(Q_k\) is local quadratic cost
Result: Feedback gain \(K_k\) and feedforward term \(d_k\)
Forward pass: Apply \(\delta u_k = K_k \delta x_k + d_k\) with line search
Iterate until convergence
23.1.2 Connection to Integral Superposition
DDP exploits: - Local linearization at each step (tangent space) - Local quadratic cost (second-order expansion) - Linear transport of cost-to-go via state transition
Each iteration refines the nominal trajectory by solving a sequence of exact local problems and integrating the results.
23.2 Iterative LQR (iLQR)
iLQR is a variant of DDP with first-order cost expansion (quadratic in \(\delta u\), linear in \(\delta x\)). The backward pass computes:
where \(K_k\) and \(d_k\) come from solving local LQR problems along the trajectory.
The key insight: each local LQR problem is solved exactly (via Riccati recursion), and the global solution emerges from integrating these exact local solutions.
23.3 Model Predictive Control (MPC)
MPC applies the same principle in real-time:
At current state \(x(t)\), solve optimal control over finite horizon \([t, t+T]\)
Apply only the first control \(u^*(t)\)
Measure new state \(x(t+\Delta t)\)
Repeat with updated state and shifted horizon
MPC is “repeated DDP/iLQR” at each timestep, exploiting integral superposition in each optimization.
Certainty-equivalence: Solve deterministic problem for mean; robust if noise small (exploits local linearity).
23.12 Summary: DDP/iLQR/MPC
Key insights:
DDP/iLQR iteratively solve sequence of exact local LQR problems
Riccati recursion propagates cost-to-go backward through tangent spaces
Feedback gains\(K_k\) exploit exact local superposition
MPC applies receding horizon for robustness to disturbances
Computational efficiency\(O(Nn^3)\) enables real-time control
All methods succeed by exploiting the exact infinitesimal linearity established in Parts I and II.
24 Case Studies and Worked Examples
This chapter presents three detailed case studies demonstrating the tangent space framework in practice: spacecraft rendezvous (orbital mechanics), planar robot arm (manipulation), and quadrotor stabilization (underactuated flight). Each example includes complete problem formulation, DDP/iLQR implementation, and results analysis.
ImportantIllustrative, Not Validated
The code blocks in this chapter are abbreviated sketches (some solver bodies are stubbed), and the numeric “Results” below are representative, hand-chosen values illustrating the expected qualitative behavior — they are not the output of a validated simulation and have not been benchmarked against empirical data. Treat them as worked-example targets, not measured results.
24.1 Case Study 1: Spacecraft Rendezvous
24.1.1 Problem Formulation
Scenario: Transfer a spacecraft from initial orbit to rendezvous with target (e.g., ISS docking).
Terminal state (achieved): \([-0.02, 0.01, 0.003]\) m, \([0.0001, 0.0002, 0]\) m/s
Convergence: 8 iterations, cost reduced from \(5.2 \times 10^6\) to \(3.1 \times 10^3\)
Fuel consumption: \(\int \|u\| dt = 12.4\) N·s (equivalent to 0.124 kg propellant at \(I_{sp} = 300\)s)
NoteTangent Space Perspective
The CW equations are already linear, so tangent spaces don’t vary—perfect test case! The framework reduces to standard LQR, demonstrating that our approach subsumes linear control as a special case where all tangent spaces are identical.
For nonlinear orbital mechanics (elliptical orbits, J2 perturbations), tangent spaces vary significantly, and DDP/iLQR exploits local linearity at each point along the trajectory.
24.2 Case Study 2: Planar Robot Arm
24.2.1 Problem Formulation
Configuration: Two-link planar arm with revolute joints.
Results: - Convergence: 12 iters - Smooth trajectory respecting arm dynamics (Coriolis effects visible in joint accelerations) - Final position error: 1.2 cm (within tolerance)
ImportantNonlinear Coupling
The inertia matrix \(M(q)\)depends on configuration—tangent spaces vary! At \(\theta_2 = 0\) (straight arm), \(M\) is maximal; at \(\theta_2 = \pm 90°\), different coupling.
DDP succeeds by: 1. Linearizing at each \(q_k\) along trajectory (exact local \(M(q_k)\)) 2. Propagating cost backward through varying tangent spaces 3. Computing gains \(K_k\) that adapt to local geometry
This is impossible with global linearization (fixed \(M\)), but natural in tangent space framework.
24.3 Case Study 3: Quadrotor Stabilization
24.3.1 Problem Formulation
States: Position \(p = [x, y, z]^T\), velocity \(v\), orientation (quaternion \(q\)), angular velocity \(\omega\)
\(x = [p, v, q, \omega]^T \in \mathbb{R}^{13}\) (quaternion has 4 components but 3 DOF)
Tangent space at hover is 12D (position, velocity, attitude, rates). Away from hover, tangent space rotates with orientation—this is where framework shines.
24.3.3 DDP Implementation Sketch
def quadrotor_cost(x, u, x_target):"""Quadratic cost: state deviation + control effort""" Q = np.diag([10, 10, 50, # Position (z weighted)1, 1, 1, # Velocity100, 100, 100, # Orientation (Euler angles from quat)1, 1, 1]) # Angular velocity R =0.1* np.eye(4) dx = x - x_targetreturn0.5* dx.T @ Q @ dx +0.5* u.T @ R @ u# Run DDP for 50 steps (2 seconds at 40 Hz)X_opt, U_opt = ddp_solver(quadrotor_dynamics, x0, x_target, N=50)
Results: - Settles to 5 cm radius in 1.5 s - Smooth control (no oscillations) - Robust to 20% model uncertainty
NoteUnderactuation and Tangent Space
Quadrotor is underactuated: 6 DOF (3 position + 3 orientation), but 4 controls. Cannot independently control all directions instantaneously.
Tangent space insight: Locally, the 4 control inputs span a 4D subspace of the 6D acceleration space. DDP finds sequences exploiting this structure: - Tilt to generate horizontal force (couple position/attitude) - Integrate over time to achieve desired position
This coupling is exact in tangent space (Jacobian \(B\) has rank 4), and DDP/iLQR leverages this to plan feasible trajectories.
24.4 Comparative Analysis
Case Study
Nonlinearity Source
Tangent Space Variation
DDP Benefit
Spacecraft
Coriolis (linear CW)
None (constant)
Reduces to LQR
Robot Arm
Configuration-dependent inertia
High (varies with \(q\))
Essential for convergence
Quadrotor
Rotation matrix, underactuation
Moderate (orientation-dependent)
Enables coupled control
Common thread: All exploit exact local linearization + integral superposition. Success depends on how much tangent spaces vary, not on “degree of nonlinearity.”
24.5 Reference Implementation Notes
The case studies above are illustrative sketches rather than a released codebase; there is no separate companion repository. A complete implementation would include:
Python/JAX implementations of all three case studies
This thesis has developed a unified geometric framework for understanding nonlinear dynamics through the lens of exact infinitesimal superposition. The central insights are:
26.1 Part I: Local Exactness
Tangent spaces are exactly linear by construction (Chapter 1)
Linearization captures the exact infinitesimal structure of smooth dynamics (not an approximation)
Time evolution passes through a continuous family of exact linear snapshots (Chapter 2)
Local optimal control (LQR, Lyapunov) succeeds by exploiting this exactness (Chapter 3)
Residuals measure geometric curvature, not modeling error (Chapter 4)
26.2 Part II: Integral Accumulation
Integration preserves linear superposition of variations (Chapter 5)
State transition operators transport perturbations linearly through moving tangent spaces (Chapter 6)
Physical conservation laws (impulse, work) are integral consequences (Chapter 7)
Global residuals scale quadratically in perturbation size and linearly in time (Chapter 8)
DDP/iLQR/MPC succeed by solving exact local problems and integrating results (Chapter 11)
27 Conceptual Contributions
27.1 Reframing Linearization
The most significant conceptual contribution is reframing linearization from “first-order approximation” to “exact infinitesimal structure.” This shift in perspective:
Explains why linear techniques work so well for nonlinear systems
Provides rigorous justification for stability analysis via eigenvalues
Clarifies the role of residuals as geometric features vs. errors
Unifies diverse control methods (LQR, gain scheduling, MPC) under a common framework
27.2 The Superposition Spectrum
Nonlinearity does not eliminate superposition—it relocates it along a spectrum:
Domain
Superposition
Source
Single tangent space
Exact
Linearity of derivatives
Infinitesimal time
Exact
Limit definition of \(\dot{x}\)
Integrated variations
Exact (first-order)
Linearity of integration
Finite perturbations
Approximate
Quadratic residuals from curvature
Global state space
Generally fails
Tangent spaces vary significantly
27.3 Geometric vs. Analytic Perspectives
This thesis bridges two perspectives:
Analytic: Taylor expansions, \(o(\cdot)\) notation, limit arguments
Geometric: Manifolds, tangent bundles, curvature
showing they are two languages for the same mathematical structure. This connection enriches both control theory (typically analytic) and differential geometry (typically abstract) by grounding geometric concepts in engineering applications.
28 Practical Impact
The framework has immediate practical implications:
28.1 Design Methodology
Linearize locally with confidence (it’s exact at each point)
Design linear controllers (LQR, pole placement) for tangent space
Schedule gains as operating point changes (following tangent spaces)
Monitor residuals to detect when nonlinear effects dominate
Switch to nonlinear methods (DDP, MPC) when residuals grow large
28.2 Analysis Tools
Lyapunov stability via linearization with rigorous local conclusions
Sensitivity analysis via state transition operators
Performance bounds via residual quantification
Model validation via geometric residual vs. modeling error decomposition
28.3 Software Implementation
The discrete-time perspective (Part III) directly informs numerical implementations: - Per-step exact linearization - Accumulated variations via recursive state transition - Real-time residual monitoring - Adaptive timestep selection based on curvature
29 Limitations and Open Questions
29.1 Scope Limitations
As discussed in Chapter 1, the framework requires \(C^1\) smoothness. Extensions needed for:
Discontinuous systems (switching, impacts, friction) via differential inclusions
Hybrid systems (discrete + continuous) via hybrid automata
Constrained systems (inequality constraints) via complementarity theory
Stochastic systems (noise-driven) via stochastic differential geometry
29.2 Computational Challenges
Part III (control synthesis) needs substantial development:
Convergence rates of DDP/iLQR not fully characterized
Stability guarantees for MPC under model mismatch unclear
Computational complexity for high-dimensional systems prohibitive
Can residual bounds be improved? Current \(O(\epsilon^2)\) scaling is crude
What is the role of higher-order geometry? (Torsion, connection, etc.)
How to extend to infinite-dimensional systems? (PDEs, delay equations)
What about non-smooth optimization? (Bang-bang control, L1 costs)
30 Future Work
30.1 Theoretical Extensions
Short-term (1-2 years): - Sharper residual bounds via curvature tensor analysis - Explicit connection to Riemannian geometry (metric structure, geodesics) - Stochastic extension via tangent spaces to Wiener processes - Hybrid system formulation with jump maps
Long-term (3-5 years): - Infinite-dimensional extension to PDE-constrained optimization - Information-geometric perspective (Fisher information on manifolds) - Quantum-classical correspondence in phase space
30.2 Practical Applications
Immediate (0-1 year): - Software library implementing state transition operators and residual monitoring - Worked examples: quadrotor, robot arm, spacecraft (complete Part III) - Comparison studies: proposed framework vs. standard methods
Near-term (1-3 years): - Industrial case studies (automotive, aerospace, manufacturing) - Integration with commercial MPC packages - Real-time embedded implementations
30.3 Pedagogical Development
Interactive visualizations of moving tangent spaces
Open-source course materials based on this framework
Textbook expanding this thesis with exercises and solutions
31 Final Reflection
This thesis has argued that nonlinear systems are exactly linear in the instantaneous tangent-space mapping. This is not mere wordplay—it is a precise mathematical statement with profound implications for control theory, numerical analysis, and geometric mechanics.
The tangent space is not an approximation we impose for convenience. It is the exact local structure of smooth dynamics, inherited from the definition of derivatives and limits. Superposition does not disappear in nonlinear systems—it relocates from global state space to local tangent spaces, where it holds exactly at each fixed state. The tangent space moves with the state, so trajectories are not additive—but the instantaneous structure is exact.
This perspective transforms how we think about nonlinearity: - Not as a failure of linear methods - But as a call to think geometrically: trajectories winding through manifolds, tangent spaces providing exact local snapshots, curvature encoding nonlinear coupling
The practical payoff is substantial: it explains why LQR works, why Lyapunov analysis is rigorous, why MPC succeeds—all by recognizing that these methods exploit the exact local linearity that exists everywhere in smooth nonlinear systems.
The tangent space framework is not an approximation theory. It is the geometric foundation of nonlinear mechanics.
A unified framework showing how complex nonlinear systems can be understood as a sequence of simple linear systems, allowing us to control them precisely.
The Winding Road (Tangent Spaces)
A mountain road is curvy overall, but at any single step, the ground under your feet is flat. Similarly, nonlinear systems are complex overall, but at any single instant, they follow simple linear rules.
Think of it like: Hiking a mountain. The whole mountain is curved, but every step you take is on a flat patch of ground. By taking many small steps on flat patches, you can climb the curved mountain.
Freeze Frame (Exact Superposition)
In a "freeze frame" of time, we can separate different forces (like gravity vs. engine power) and add them up simply. This allows us to calculate the exact effect of each force at that specific moment.
Think of it like: A grocery bill. The total cost is complex, but it's just the sum of individual items (apples + bread + milk). At the checkout counter (the specific instant), you can add them up linearly.
Accumulating Change (Integration)
We can predict the future by stringing together these simple linear snapshots. Even though the system changes constantly, calculating the path one small step at a time gives the correct result.
Think of it like: Compound interest. The growth rate is constant for a day, but over a year, the total amount grows exponentially. We calculate the year's growth by adding up 365 days of simple interest.
Key Takeaway: Complex systems are just a series of simple linear moments strung together; understanding this allows us to control robots and spacecraft with precise, simple math.
References
Abraham, R., & Marsden, J. E. (1978). Foundations of Mechanics (2nd ed.). Benjamin/Cummings.
Arnold, V. I. (1989). Mathematical Methods of Classical Mechanics (2nd ed.). Springer.
Bryson, A. E., & Ho, Y.-C. (1975). Applied Optimal Control. Hemisphere Publishing.
Jacobson, D. H., & Mayne, D. Q. (1970). Differential Dynamic Programming. Elsevier.
Khalil, H. K. (2002). Nonlinear Systems (3rd ed.). Prentice Hall.
Lee, J. M. (2012). Introduction to Smooth Manifolds (2nd ed.). Springer.
Liberzon, D. (2003). Switching in Systems and Control. Birkhäuser.
Mayne, D. Q., et al. (2000). Constrained model predictive control: Stability and optimality. Automatica, 36(6), 789-814.
Sastry, S. (1999). Nonlinear Systems: Analysis, Stability, and Control. Springer.
Tassa, Y., et al. (2012). Synthesis and stabilization of complex behaviors through online trajectory optimization. In Proc. IEEE/RSJ IROS (pp. 4906-4913).
Appendices
32 Appendix a: Mathematical Prerequisites
[To be added: Brief review of manifolds, tangent bundles, differential forms, Lie groups for readers unfamiliar with differential geometry]
33 Appendix B: Proofs and Derivations
This appendix supplies explicit bounds and convergence statements used in the main text.
34 B.1 Residual Scaling Bound
Consider a control-affine nonlinear system \[
\dot{x} = f(x) + G(x)u,
\] with \(f, G \in C^2\) on a compact tubular neighborhood \(\Omega\) around a nominal trajectory \((x^\star(t), u^\star(t))\) on \([t_0, t_1]\).
Let \(\delta x_1, \delta x_2\) be perturbation trajectories generated by two infinitesimal input perturbations \(\delta u_1, \delta u_2\), and define the superposition residual \[
r(t) := \delta x_{12}(t) - \delta x_1(t) - \delta x_2(t),
\] where \(\delta x_{12}\) is the perturbation generated by \(\delta u_1 + \delta u_2\).
Since \(f\) and \(G\) are \(C^2\) on compact \(\Omega\), there exists \(M_H > 0\) such that for all \(z \in \Omega\): \[
\|D^2 f(z)\| \le M_H, \qquad \|D^2(G(z)u^\star(t))\| \le M_H.
\]
By second-order Taylor expansion around \(x^\star(t)\), the non-canceling part of the additivity defect is quadratic: \[
\dot{r}(t) = A(t)r(t) + \eta(t), \qquad \|\eta(t)\| \le \frac{M_H}{2}\|\delta x_1(t)+\delta x_2(t)\|^2,
\] with \(r(t_0)=0\) and \(A(t)=\partial_x(f+Gu^\star)|_{x^\star(t)}\).
Using variation of constants, \[
r(t_1)=\int_{t_0}^{t_1}\Phi(t_1,\tau)\eta(\tau)\,d\tau,
\] and a uniform transition bound \(\|\Phi(t_1,\tau)\|\le C_\Phi\), we obtain \[
\|r(t_1)\|
\le \frac{C_\Phi M_H}{2}\int_{t_0}^{t_1}\|\delta x_1(\tau)+\delta x_2(\tau)\|^2\,d\tau.
\] Therefore, \[
\|r(t_1)\| \le C_r \int_{t_0}^{t_1}\|\delta x(\tau)\|^2\,d\tau
\quad\text{for } C_r := \frac{C_\Phi M_H}{2},
\] which is the required quadratic scaling.
35 B.2 Curvature-Tensor Interpretation
Let \(\mathcal{M}\) be a Riemannian state manifold with Levi-Civita connection \(\nabla\) and curvature tensor \(R\). For nearby displacement fields \(v,w \in T_x\mathcal{M}\) transported along the nominal flow, the first non-commuting term in second-order transport is governed by curvature: \[
\nabla_v\nabla_w \xi - \nabla_w\nabla_v \xi
= R(v,w)\xi.
\] In coordinates, this contributes second-order coupling of the form \[
r^i \approx \frac{1}{2}R^i_{\;jkl}\,v^j w^k \,\Xi^l + O(\|v\|^3+\|w\|^3).
\] Hence residual magnitude is controlled by sectional curvature: \[
\|r\| \le C_K \|v\|\,\|w\|\,\|\Xi\| + O(3\text{rd order}),
\] which formalizes “residuals as curvature.”
36 B.3 Discrete-Time Convergence Rate
For Euler discretization with step \(\Delta t\): \[
x_{k+1}=x_k+\Delta t\,f(x_k,u_k),
\] the local truncation error is \(O(\Delta t^2)\) under \(C^2\) smoothness. Over \(N=(t_1-t_0)/\Delta t\) steps, global integration error is \[
\|e_N\| = O(\Delta t).
\] For linearized-transport accumulation, each step introduces second-order residual \(O(\|\delta x_k\|^2\Delta t)\) and first-order discretization defect \(O(\Delta t^2)\). Under bounded perturbations \(\|\delta x_k\|\le \bar{\delta}\): \[
\sum_{k=0}^{N-1} O(\|\delta x_k\|^2\Delta t) = O(\bar{\delta}^2 (t_1-t_0)),
\] and \[
\sum_{k=0}^{N-1} O(\Delta t^2) = O(\Delta t).
\] Therefore the discrete approximation converges with order one in \(\Delta t\), with an additive curvature-controlled term: \[
\|e_N\| \le C_1 \Delta t + C_2\int_{t_0}^{t_1}\|\delta x(t)\|^2dt.
\] This gives an explicit practical criterion: refine timestep to reduce integration error, and reduce perturbation magnitude (or operate in lower-curvature regions) to reduce geometric residuals.