The Cash-Flow Game — learning, live

Overvloed & Fabel B.V., groothandel in feestartikelen. Every week the network sees the books and decides, for each open supplier invoice: pay now, or wait. It has been told the rules and the goal — the end-of-year net position — and nothing else. Three human rules of thumb are already on the board. Nothing here is a recording: the network below is untrained, and starts learning when you press the button.

0.0s
elapsed
0
training rounds
0
company years lived
0
invoice decisions
0
decisions / second

Net position at the end of the test year — a year it never trains on

the network — 2,929 numbers, starting from nothing
Discount Hunter — take the 2% discount, otherwise pay on the due date
Stretcher — pay only when due (€93 below the Hunter, so the two lines sit on top of each other)
Prompt Payer — pay everything the moment it arrives

Press start. It begins knowing nothing, and it does get worse before it gets better.

What it actually did with the year, week by week

paid early, took the 2% discount paid on time paid late held past due bottom row: credit line drawn · gold: btw quarters

A 2,929-parameter network, reinforcement learning from the rules alone. Built by Karim & Claude. The full trained version, and the story of how it got there, is at the Cash-Flow Game.