Decide, learn, decide again — working trial 8.1
Job and theoretical bridge
Compare acting, buying information and staging a fictional renewal while preserving the information available at each decision. The formulation uses Warren Powell’s five-element framework: observable state, decision, new information, transition and objective. The constraints, equations, numbers and three policies here are newly authored, not supplied or validated by Powell.
Primary references checked during construction:
- Warren B. Powell, A Universal Framework for Sequential Decision Problems, ORMS Today, 6 February 2023: source for the five-element formulation, not validation of this toy.
- Powell, Supervising theses on sequential decision problems: reinforces that new information arrives after the decision.
Model
Observable state contains period, remaining capital, unchanged/staged/renewed phase, poor-condition belief, whether inspection occurred, its revealed signal and the current known window. Policies receive only that state. The simulator separately owns a fixed hidden condition, one inspection random draw per time and one window draw per time. Mixed seeded draws are generated before actions; unused draws do not shift later randomness. Seeds are reproducible examples, not secure secrets. Local source/exports are inspectable; the claim is an algorithmic information boundary, not secrecy from a browser developer.
| Inspect costs capital and reveals a symmetric noisy signal afterward. Bayes updates belief: p×likelihood(signal | poor) divided by the total signal likelihood. Stage halves exposure; renew removes it. Renewal/staging need a window and sufficient capital. Inspection is once only and before the last period. An invented belief bound blocks unmitigated inspection/deferral; it is not an operating rule. Capital plus hidden-condition period loss plus terminal residual is the declared objective in synthetic £k. No ticket revenue appears. |
A no-admissible-action trajectory stops and has no full-horizon objective: cost is null, excluded from completed means, and counted explicitly. Summary denominators and jointly completed paired differences are shown. Unequal feasible subsets are not a fair ranking. The app does not reward infeasibility with a truncated low cost.
Policies: early renewal acts at its first affordable window; inspect-then-decide inspects once and mitigates when posterior belief is at least 50%; stage-then-finish stages first and completes at its next opportunity. Each has a declared legal-action fallback. None is trained or claimed optimal.
Walkthrough and independent reference
Inspect reveals a signal only after recording the action. Read the next state, choose again, and inspect the completed objective. Restart same scenario reuses its tape. Update assumptions and restart changes parameters and clears the manual path. Comparison tapes begin at the next seed (wrapping the 32-bit seed range) and exclude the active manual seed; their displayed paths cannot reveal that manual trajectory. The comparison uses the same tapes for every policy and displays action paths and distributions.
Two-period exact enumeration: prior poor probability 0.5, perfect inspection, both windows available, costs inspection=2, renew=12, stage=6, finish=8, poor period loss=9, good loss=1, terminal poor=8/good=2. Early renewal costs 12 in both states. Inspect-then-decide costs 2+9+12=23 if poor and 2+1+1+2=6 if good, expectation 14.5. Stage-then-finish costs 6+4.5+8=18.5 if poor and 6+0.5+8=14.5 if good, expectation 16.5. Information can cost more than it saves.
Checks and full ambition
node --test apps/sequential-decisions/app.test.mjs runs nine tests: exact enumeration, future-information isolation, common-tape determinism, adjacent-seed mixing, infeasible-cost handling, atomic action/import failures, capital/admissibility reconciliation and exclusion of the manual seed from comparison samples and null full-horizon objectives for partial paths.
The simpler comparator is a small decision tree with explicit information timing. The local numerical demonstration is complete. It is not a learned/optimal policy, forecast, engineering model or permission to operate, defer real maintenance or take access. Practical usefulness has not been measured.