Climate engineering megaproject as a sequential decision problem

Compare five small decision policies on the same illustrative model: a direct rule, modified-cost search, approximate-value search, a short lookahead and a hybrid. Explore one path, many possible paths, scope-action ribbons and a framing tool for your own project. Coefficients are invented teaching assumptions; results are not climate forecasts or policy recommendations.

1. Powell-style model specification

Decision plan for the climate-megaproject

Operational actuation The project carries a deployed intervention scope u_t that actually affects climate dynamics over the next quarter.

Management decisions Each quarter management chooses a scope adjustment Δu_t ∈ {−δ,0,+δ}, a monitoring intensity m_t, and a governance / permitting push e_t.

Uncertainty reduction Monitoring is an explicit information action. Higher m_t costs more and reduces measurement noise. Belief uncertainty contracts only when deployed scope exposes the unknown effectiveness; observing zero exposure teaches nothing about it.

Quarterly trigger The management loop compares the observed quarterly temperature trend against the trend expected under the current belief state, then revises scope and information effort.

State, information, transition, objective

{
"State S_t": {
  "Physical R_t": ["estimated temperature anomaly T̂_t", "active scope u_t", "funding balance B_t (soft limit)", "governance readiness G_t"],
  "Information I_t": ["baseline warming trend g", "last observed quarterly trend y_t", "last expected trend ŷ_t"],
  "Belief B_t": ["posterior mean μ_k,t of effectiveness", "posterior sd σ_k,t of effectiveness"]
},
"Decision x_t": ["scope change Δu_t", "monitoring level m_t", "governance effort e_t"],
"Exogenous W_{t+1}": ["natural variability shock", "measurement noise", "governance shock", "cost shock"],
"Transition": "S_{t+1} = S^M(S_t, x_t, W_{t+1})",
"Objective": "minimize cumulative composite cost = temperature deviation + overshoot risk + spend + governance fragility + change friction"
}

The simulator keeps the actual temperature and a fixed effectiveness coefficient separate from the manager’s state. Policies receive only the estimated level, scope, funding balance, observed readiness and posterior belief. New shocks and measurements arrive after the choice. The Gaussian belief update uses known process and measurement variances; its prior can assign weight to ineffective or adverse outcomes.

Cost combines squared target deviation, extra exceedance penalty, spending, readiness shortfall, scope-change friction and funding deficit, discounted by 0.985 per quarter. Funding is a soft balance: negative values mean an unfunded model path, not approval to spend. The score uses actual simulated temperature after the outcome; decisions use estimates.

The seed, quarters and run count below apply to every comparison. Each refresh rebuilds the benchmark and both path views. Powell’s policy taxonomy · Model framework · This model and repair record

2. Implemented policy classes

Computed benchmark across policy representatives

Every row is recomputed from the current assumptions, horizon and seed batch. Policies share identical underlying shocks per seed; outcomes differ because their actions differ. Lower composite cost is better within this toy model.

3. Single-path simulation

Quarterly log

4. Many-path simulation and branching view

Branching paths (scope-action ribbons)

Each polyline is one sample path. Vertical position encodes the scope decision that quarter: decrease / hold / increase. Bundles show how often this sampled policy chooses the same scope change; they do not establish robustness or optimality.

5. Frame your own project as Powell + RL

Powell frame


      

Equivalent RL / MDP frame


      

Information-attention checklist

6. Notes, limits, and interpretation

What this is

A stylized quarterly programme with a fixed hidden effectiveness and a conjugate Gaussian belief update. PFA is a declared rule; CFA modifies one-step cost; VFA adds a hand-written downstream value; DLA compares first actions using a four-quarter deterministic rollout and a PFA tail; the hybrid adds a terminal value.

What this is not

Not a physical climate model, not a welfare analysis, not a claim about the advisability of any real intervention, and not a trained deep-RL controller.

Why it is still useful

It separates state, choices, incoming information and outcomes, and lets you inspect a small representative of each class. The lookahead uses the current belief mean and expected observations; it does not value contingent future learning exactly, search all policies or establish which class is superior.