1. Powell-style model specification
Decision plan for the climate-megaproject
Operational actuation The project carries a deployed intervention scope u_t that actually affects climate dynamics over the next quarter.
Management decisions Each quarter management chooses a scope adjustment Δu_t ∈ {−δ,0,+δ}, a monitoring intensity m_t, and a governance / permitting push e_t.
Uncertainty reduction Monitoring is an explicit information action. Higher m_t costs more and reduces measurement noise. Belief uncertainty contracts only when deployed scope exposes the unknown effectiveness; observing zero exposure teaches nothing about it.
Quarterly trigger The management loop compares the observed quarterly temperature trend against the trend expected under the current belief state, then revises scope and information effort.
State, information, transition, objective
{
"State S_t": {
"Physical R_t": ["estimated temperature anomaly T̂_t", "active scope u_t", "funding balance B_t (soft limit)", "governance readiness G_t"],
"Information I_t": ["baseline warming trend g", "last observed quarterly trend y_t", "last expected trend ŷ_t"],
"Belief B_t": ["posterior mean μ_k,t of effectiveness", "posterior sd σ_k,t of effectiveness"]
},
"Decision x_t": ["scope change Δu_t", "monitoring level m_t", "governance effort e_t"],
"Exogenous W_{t+1}": ["natural variability shock", "measurement noise", "governance shock", "cost shock"],
"Transition": "S_{t+1} = S^M(S_t, x_t, W_{t+1})",
"Objective": "minimize cumulative composite cost = temperature deviation + overshoot risk + spend + governance fragility + change friction"
}
The simulator keeps the actual temperature and a fixed effectiveness coefficient separate from the manager’s state. Policies receive only the estimated level, scope, funding balance, observed readiness and posterior belief. New shocks and measurements arrive after the choice. The Gaussian belief update uses known process and measurement variances; its prior can assign weight to ineffective or adverse outcomes.
Cost combines squared target deviation, extra exceedance penalty, spending, readiness shortfall, scope-change friction and funding deficit, discounted by 0.985 per quarter. Funding is a soft balance: negative values mean an unfunded model path, not approval to spend. The score uses actual simulated temperature after the outcome; decisions use estimates.
The seed, quarters and run count below apply to every comparison. Each refresh rebuilds the benchmark and both path views. Powell’s policy taxonomy · Model framework · This model and repair record
2. Implemented policy classes
Computed benchmark across policy representatives
3. Single-path simulation
Quarterly log
4. Many-path simulation and branching view
Branching paths (scope-action ribbons)
5. Frame your own project as Powell + RL
Powell frame
Equivalent RL / MDP frame
Information-attention checklist
6. Notes, limits, and interpretation
What this is
A stylized quarterly programme with a fixed hidden effectiveness and a conjugate Gaussian belief update. PFA is a declared rule; CFA modifies one-step cost; VFA adds a hand-written downstream value; DLA compares first actions using a four-quarter deterministic rollout and a PFA tail; the hybrid adds a terminal value.
What this is not
Not a physical climate model, not a welfare analysis, not a claim about the advisability of any real intervention, and not a trained deep-RL controller.
Why it is still useful
It separates state, choices, incoming information and outcomes, and lets you inspect a small representative of each class. The lookahead uses the current belief mean and expected observations; it does not value contingent future learning exactly, search all policies or establish which class is superior.