← Return to the example · Supporting files

The monthly portfolio review

An unfinished exploration by Lawrence Rowland, December 2020. Reading guide added 1 October 2026.

What should a portfolio review decide this month, and what should it learn before deciding again? The original note considers advancing, suspending or cancelling projects alongside investigating, reviewing and assuring them. It asks whether the information arriving between meetings is good enough, and whether the decision cadence fits that information.

At this review, state S zero informs decision x zero. Progress and new information W one lead into next month's updated state S one and decision x one.

This 2026 reading diagram interprets the note’s sequence. The notebooks do not implement the incoming-information or assurance parts of that idea.

Read the original December 2020 note · Methods guide · Library

What is worth recovering?

The note starts with two useful alternatives: what decisions can we make? and what are we trying to achieve? It then connects today’s decision with tomorrow’s information and the next decision. Three distinctions matter:

The original note also asks whether information is gathered effectively, decisions are made promptly, and reviews occur at a useful frequency. Its OODA comparison and possible outputs—a meeting procedure, decision checklist and decision tables—are ideas to inspect, not delivered products or a prescribed next research agenda.

What the notebooks actually do

Both use the same broad lifecycle:

Proposal → Business case → Planned → Started → Complete, with cancellation as an alternative terminal state.

The chosen actions are promote, maintain and cancel. Their effects and rewards are supplied by hand. An action is selected randomly at each step; neither notebook learns a policy, represents an assurance action, consumes new evidence, or chooses between projects under shared resource constraints. A simulation step is not calibrated to a calendar month.

Main notebook: repeated random trials

Open Portfolio_lifecycle_statemachine.ipynb.

This version sets up 99 trials, each with two projects starting at Proposal and nine decisions per project. It advances the stored state after each decision. Complete and Cancelled are absorbing states with zero subsequent reward. Promoting a Started project yields a supplied reward of 50; maintaining a non-terminal project costs 2. These are illustrative scores, not measured money or benefits.

Known counting defect: success increases on every subsequent step spent in Complete. It is not a count of completed projects. For example, completing on step 4 adds to the counter at steps 4 through 9. The saved output is evidence of an earlier random run, not an estimate of a learned strategy’s success rate.

Copy1: a different earlier attempt

Open Portfolio_lifecycle_statemachine-Copy1.ipynb.

This version starts one project at Proposal and one at Business Case, uses five decisions each, and changes two costs: maintaining a Started project costs 15 rather than 2; cancelling at Business Case costs 5 rather than 2. It is not an exact duplicate.

Known transition defect: the loop copies state = p.state before its decisions but never updates that local variable. Although it writes the result back to the project, the next call still uses the starting state. The saved trace consequently shows a cancelled project returning to Proposal. It cannot be read as a valid lifecycle simulation.

The two notebook files and their saved outputs are unchanged and have not been rerun for this review. Their contrasting attempts remain available for inspection; neither is presented as an executable solution ready for portfolio use.

How to read the original claims

The note’s title is Reinforcement learning for project portfolios. Its broader question is sequential decision-making; the surviving code implements random action selection and fixed transitions, not a learning algorithm. Claims about selecting the best design, optimising long-term benefits or choosing the most informative action describe ambitions, not demonstrated results.

The brief assertions about stationarity and ergodicity are not established in the note. They require a defined process and, for relevant long-run behaviour, a policy. The note also conflates changing review frequency with the learning step-size alpha; those are different choices. These clarifications are included at the original note’s entrance so it can be read in context.

Sources and limits

The useful result is an explicit set of questions for a portfolio review. There is no learned policy, calibrated uncertainty model, validated investment recommendation or evidence of operational benefit here.

Return to the Library’s methods guide