Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Lorenz '96 Two-Tier — Slow-Fast Multiscale Chaos

Authors
Affiliations
University of Valencia

The two-tier Lorenz '96 system couples a set of slow large-scale variables XkX_k to a set of fast small-scale variables Yj,kY_{j,k}. It is the standard test bed for studying multiscale dynamics, sub-grid parameterization, and stochastic closures.

What you’ll learn:

  1. How slow and fast variables interact through coupling

  2. How the Hovmoller diagram reveals the multiscale structure

  3. How coupling strength controls fast-variable activity

  4. How to differentiate through a multiscale simulation

Background

The system consists of DxD_x slow variables XkX_k and Dy×DxD_y \times D_x fast variables Yj,kY_{j,k}:

dXkdt=(Xk+1−Xk−2) Xk−1−Xk+F−hcb∑j=1DyYj,k\frac{dX_k}{dt} = (X_{k+1} - X_{k-2})\, X_{k-1} - X_k + F - \frac{hc}{b} \sum_{j=1}^{D_y} Y_{j,k}
dYj,kdt=cb (Yj+1,k−Yj−2,k) Yj−1,k−c Yj,k+hcb Xk\frac{dY_{j,k}}{dt} = cb\,(Y_{j+1,k} - Y_{j-2,k})\, Y_{j-1,k} - c\, Y_{j,k} + \frac{hc}{b}\, X_k

The key parameters are:

ParameterRoleStandard value
FFExternal forcing on slow variables18
hhCoupling coefficient1
bbAmplitude ratio (fast / slow)10
ccTime-scale ratio (fast / slow)10

The fast variables evolve c=10c = 10 times faster than the slow variables and have b=10b = 10 times smaller amplitude. The coupling term hc/bhc/b transfers energy between scales.

1. Create the model

We use Dx=36D_x = 36 slow variables and Dy=10D_y = 10 fast variables per slow variable (360 fast variables total).

Lorenz96t(
  params=L96TParams(F=weak_f32[], h=weak_f32[], b=weak_f32[], c=weak_f32[])
)

2. Forward simulation

The fast variables require a small time step. We integrate for 2 time units (roughly 10 days in the L96 convention) and save at a moderate output rate.

Slow variables: (400, 36)
Fast variables: (400, 360)

3. Slow and fast time series

The slow variables look like noisy oscillations (similar to standard L96), while the fast variables show rapid fluctuations modulated by the slow-variable envelope.

<Figure size 1300x600 with 2 Axes>

4. Hovmoller diagrams — slow vs fast

The slow-variable Hovmoller shows the familiar wave propagation. The fast-variable Hovmoller (averaged over the jj-index) reveals the small-scale activity modulated by the large-scale pattern.

<Figure size 1400x500 with 4 Axes>

5. Energy partition

The total energy splits between slow and fast components. The fast variables carry less energy per mode but there are DyD_y times as many of them.

<Figure size 1200x350 with 1 Axes>

6. Coupling strength

The coupling coefficient hh controls how strongly the fast variables feed back onto the slow variables. We compare three regimes: uncoupled (h=0h = 0), weakly coupled (h=0.5h = 0.5), and strongly coupled (h=1h = 1).

<Figure size 1500x400 with 3 Axes>

7. Adjoint methods and differentiation

Differentiating through a stiff multiscale ODE is more challenging than a single-scale system because the fast variables demand many time steps. diffrax provides three adjoint methods:

AdjointMemoryGradientsUse case
RecursiveCheckpointAdjointO(N)O(\sqrt{N})ExactDefault
DirectAdjointO(N)O(N)ExactShort windows
BacksolveAdjointO(1)O(1)ApproxLong windows
ImplicitAdjointO(1)O(1)ExactSteady states
  • RecursiveCheckpointAdjoint (the default) uses optimal checkpointing: re-computes forward steps from saved checkpoints, giving exact gradients with O(N)O(\sqrt{N}) memory.

  • DirectAdjoint stores the entire forward trajectory — exact but O(N)O(N) memory.

  • BacksolveAdjoint solves the continuous adjoint ODE backwards with O(1)O(1) memory. Gradients are approximate.

  • ImplicitAdjoint uses the implicit function theorem: du∗/dθ=−(dg/du)−1 dg/dθdu^*/d\theta = -(dg/du)^{-1}\, dg/d\theta. Exact and O(1)O(1) memory, but only for steady-state / fixed-point solves.

For this system with dt=5×10−4dt = 5 \times 10^{-4}, even a short window of T=0.5T = 0.5 requires 1000 steps. RecursiveCheckpointAdjoint is the best default because it avoids storing all 1000 forward states.

7a. Gradient w.r.t. parameters

Use case: parameter estimation. All four coupling parameters (F,h,b,c)(F, h, b, c) are differentiable. The adjoint equations are:

∂L∂θ=∫0Tλ(t)⊤∂f∂θ dt,θ=(F,h,b,c)\frac{\partial \mathcal{L}}{\partial \theta} = \int_0^T \lambda(t)^\top \frac{\partial f}{\partial \theta}\, dt, \qquad \theta = (F, h, b, c)
--- Gradient w.r.t. parameters ---
  dL/dF = 193.4378
  dL/dh = -128.4092
  dL/db = 39.7463
  dL/dc = 0.3340

7b. Gradient w.r.t. initial state

Use case: 4D-Var. Find the initial condition (X0,Y0)(\mathbf{X}_0, \mathbf{Y}_0) that best fits observations. The gradient is the adjoint state at t=0t = 0:

∂L∂u0=λ(0),u0=(X0,Y0)\frac{\partial \mathcal{L}}{\partial \mathbf{u}_0} = \lambda(0), \qquad \mathbf{u}_0 = (\mathbf{X}_0, \mathbf{Y}_0)

In practice one often assimilates only the slow variables and lets the fast variables spin up.

--- Gradient w.r.t. initial state ---
  dL/d(X0) shape: (36,)
  dL/d(X0) norm:  52.9541
  dL/d(Y0) shape: (360,)
  dL/d(Y0) norm:  1425.0918

7c. Joint gradient — parameters and state

Use case: weak-constraint 4D-Var. Simultaneously optimize the initial condition and all coupling parameters.

--- Joint gradient ---
  dL/dF         = 193.4378
  dL/dh         = -128.4092
  dL/d(X0) norm = 52.9541
  dL/d(Y0) norm = 1425.0918

8. Ensemble divergence

We launch 20 ensemble members and track how the slow-variable spread grows over time.

<Figure size 1300x400 with 2 Axes>

Summary

Conceptsomax API
Create a modelLorenz96t.create(F=18, h=1, b=10, c=10)
Initial conditionL96TState.init_state(ndims=(Dx, Dy))
Forward simulationmodel.integrate(state0, t0, t1, dt)
Grad w.r.t. paramseqx.filter_grad(loss)(model) — all 4 params
Grad w.r.t. statejax.grad(loss)(state0) — dL/d(X0, Y0)
Joint gradjax.grad(loss, argnums=(0, 1))(state0, model)
Ensembleeqx.filter_vmap(integrate_one)(batch_states)

Key takeaways:

  • The fast variables evolve c=10c = 10 times faster than the slow variables, requiring a small time step (dt≈5×10−4dt \approx 5 \times 10^{-4})

  • Coupling (h>0h > 0) feeds fast-variable energy back into the slow dynamics, modifying the large-scale attractor

  • This system is the standard test bed for learned sub-grid parameterizations (replacing the fast variables with a closure)