Start with evidence before commitment
Follow the answer into the next choice.
The diagram branches on what the organiser learns. It lets you inspect the next decision under either answer.
- Choose “1. Worth buying”. Follow the high-demand signal, return to the first choice, then follow the low-demand signal. The two answers lead to different bus reservations.
- Choose “3. Too late”. The booking deadline is now day 1 and the survey takes two days. The survey is unavailable for this commitment, so the model reserves one bus now.
- Return to “1. Worth buying”. Compare the value bridge: 17.5 expected service-value points after survey and waiting costs, versus 15 for committing immediately. The gain is 2.5 points under these assumptions.
Open the six worked examples → The other examples explore unhelpful clarity, routine evidence, stopping and successive tests.
Why this small decision matters
This foray asks how management decisions, evidence gathering and scheduled work can be understood together. The shuttle isolates one part: keeping a useful choice open while information arrives. Delivery itself is represented by fixed payoffs here; the model does not build an operating timetable.
The source, the construction, the limit
Powell’s framework separates state, decision, new information, transition and objective. Our independently written finite decision tree applies that structure to the shuttle: beliefs change after evidence, and the next choice uses only what has been learned.
It enumerates the declared small model, subject to floating-point arithmetic. It does not train a reinforcement-learning policy or establish a real transport recommendation.
Sources and what the essay adds · Model assumptions · Method and verification · Original foray aim
Browse all eight experiments → Framing worksheets, policy comparisons, paths, an IT game and a clearly marked historical report.
Earlier model · same foray
Sequence, belief & policy lab
Both original views remain here. Follow one climate-programme path, then compare illustrative policy representatives and frame a recurring management decision.

Toy scenario
Choose now; learn more before the next decision.
In this earlier toy programme, a manager adjusts intervention and information-gathering effort each quarter as noisy temperature observations arrive. Other experiments include committing to transport capacity and sequencing IT work. Decisions can use what has been observed, never tomorrow’s result.
When is it worth paying or waiting for evidence before committing?
A small laboratory for separating management decisions, arriving information and state change—then testing policies without pretending a stylised simulation is a forecast.
Purpose
Make the governance loop explicit
The manager does not choose an action after seeing the future shock. Information received since the previous choice is already represented in St. The current choice xt is followed by new exogenous information Wt+1, which updates the next state.
This corrects the common slide-deck blur between “observe, decide, evolve” and Powell’s compact repeating sequence. The distinction matters when testing whether a policy is using only information available at decision time.
Corrected taxonomy
Four policy classes, not four rival fields
These are small representatives of Powell’s classes, not claims of optimality. Reinforcement learning can supply methods inside or across these classes; it is not substituted for the fourth class.
PFA · Policy function approximation
A direct mapping from state to action. Here: a transparent set-point rule with an uncertainty-sensitive information choice.
CFA · Cost function approximation
Choose an action by optimising a deliberately modified immediate cost. Here: penalties add target margin and uncertainty reserve.
VFA · Value function approximation
Choose using immediate contribution plus an approximation of downstream value. Here: remaining effort, target gap and belief variance.
DLA · Direct lookahead approximation
Illustrative four-quarter rollouts—not MCTS or a proof of optimality. The tail policy uses sampled hidden parameters without learning inside the rollout, so this representative cannot establish the value of evidence. See the exact information-respecting essay.
P90 objective · lower is better
The objective is an explicit modelling choice: programme cost + schedule pressure + guardrail exceedance, with a terminal penalty. It is not a real climate-policy welfare function.
| Class | Median final | P(above guardrail) | Median cost | P90 objective |
|---|
Selected policy · path distribution
Scenario tree · not MCTS
The tree branches on possible information arrivals and displays the resulting policy decisions. It does not maintain visit counts, explore actions or back up values, so calling it MCTS would be incorrect.
Management framing tool
Match decision cadence to information cadence
Start with a minimum viable decision model. Use RL only when repeated interaction, a stable state/action representation and dependable feedback justify it.