← Rivering · beecork.com/rivering
A method, and what it found in 262 tables
Two hundred tables, and nobody can say what half of them are for. Rivering is deciding — once, on purpose — what each one is: every table turns out to be a noun or a river, taking one of a few named shapes, and that tells you what you may do to it.
A table that has been rivered answers for itself. An unrivered one can only be guessed at — and guessing is where the damage comes from, every time.
A shop writes every sale in one notebook. Then someone also writes today's total on a whiteboard.
Now two things answer "how much did we sell today?" For a while they agree, and the whiteboard is genuinely useful — reading one number beats adding up a hundred lines.
Then one day they disagree. A sale went in the notebook while the whiteboard was being wiped. And here is the part that matters: nobody in the shop can tell which one is wrong. Both were written by hand. Both look equally official. The disagreement isn't the problem — the problem is that there's no way to settle it.
Keep the notebook. Throw away the whiteboard. Add up the notebook when you need a total — and if you must keep a whiteboard for speed, it may never be written on by hand, only recopied from the notebook.
There is a second cure, and it is worth knowing because sometimes it is the only one available. The shop is stuck because nothing recorded the writing. If every write — to the notebook and to the whiteboard — were itself logged, the disagreement stops being symmetric: you can see which write landed and which didn't, and decide. That is what an audit trail buys, and it is why a system with a complete one survives a shape that would otherwise be unrecoverable. It is the more expensive cure. Prefer one original where you can have one.
These get confused constantly — including by us, while working this out. They are unrelated ideas and both matter.
Meaning B is why the arrow matters. If copies only ever flow one way, disagreement has a cure: throw the copy away and recompute it. No investigation, no judgement, no meeting. If data can flow both ways, a disagreement has no cure — that's the shop with two answers, and it is the single most expensive shape in a database.
Meaning A is what tells you a new river is needed. When you find facts about some subject smeared across five tables with no log of its own, that subject wants a river.
Worth saying plainly: the contact diary is not a sea that other rivers empty into. It is itself a river, pointed at a person. If the sea is anything, it's the whole system — the body of water the rivers run through. Rivers don't join. They run in parallel, each aimed at its own subject.
Meaning B tells you copies flow one way. It doesn't tell you which of two tables is the original — and standing in front of two things holding the same fact, that is the only question that matters.
You don't choose. What the rows are decides it.
“Levan entered the welcome sequence on Tuesday” has a time and a person. It's an event, so it's a river line. “Levan is currently in the welcome sequence” has no time and describes no occurrence — it's a state, and a state is what you get by reading events.
The arrow always runs from the event to the state. An event is never derived from a state, because a state cannot say when, or how many times.
And the arrow holds for one reason: the derivation throws something away. “Is currently enrolled” is the same answer for someone who joined once and for someone who joined, left and rejoined — two different histories, one value, so nothing can run the arrow backwards. Which has a consequence worth stating outright: if a copy loses nothing, it is not derived at all. A table that retains everything the history contained is the history in other clothes, and it inherits a river's rules rather than a derived table's. Derivation presupposes loss.
This matters because the mistake is easy and expensive. If you decide the state table is the original, then the events look like a copy and get deleted — and once they're gone, nothing can rebuild them. The state can always be rebuilt from the events. Never the reverse. That asymmetry is what makes it a rule rather than a preference.
One thing the arrow does not promise: that a state is worked out from its own river. “Levan's last email open” is a fact about a person, but it comes from the river of email events — whose subject is a message, not a person. The chain runs: this message was opened, this message was sent to Levan, therefore Levan's last open. State is always worked out from events; not always from the events about itself.
When two tables hold overlapping answers, it feels like a judgement call about which one to trust. It is not. The direction is already fixed by what the rows say.
“Levan entered the welcome sequence on Tuesday” — has a time, has a person, describes an occurrence. An event.
“Levan is currently in the welcome sequence” — no time, no occurrence. A state.
The arrow always runs from the event to the state. An event is never worked out from a state, because a state cannot say when, or how many times. The state can always be rebuilt from the events; never the reverse. That asymmetry is what makes this a rule rather than a preference.
“Is currently enrolled” is the same answer for somebody who joined once and for somebody who joined, left, and rejoined. Two different histories, one value. Nothing can run the arrow backwards, because the information is simply gone.
That consequence catches a whole class of mislabelled tables. A copy retaining everything the history contained is the history in other clothes, and it inherits a river's rules — never edited, kept for its retention — not a derived table's. Derivation presupposes loss. A table that has lost nothing has not derived anything; it has duplicated.
That a state is worked out from its own river. “Levan's last email open” is a fact about a person, but it comes from the river of email events, whose subject is a message.
State is always worked out from events — not always from the events about itself. Looking for the missing river next to the state, and concluding there isn't one, is how a perfectly ordinary derived table gets promoted to an original.
A river is not just "a table you only add to." Five decisions have to be made, and skipping any of them produces a log that looks healthy and cannot be used.
What every line is about. A person. An account. One message. This is the direction, in meaning A — and it is what decides whether you need a new river or not.
How big one line is. A conversation, or every keystroke in it? This is a real decision with no default.
A fixed list of what a line may be called — and the identity of the thing kept in its own column, never inside the name.
What may and may not be written in it. Below.
How long it is kept. Three archetypes depend on this answer, and nobody writes it down.
More detail goes into the river you already have, as more kinds of line. A new river is only ever for a different subject.
This one saves a lot of wasted work. Wanting to know more about a person — which lesson they opened, not just which course — is never a reason for a second river about them. It is a reason for more kinds of line in the one you have. You only reach for a new river when the subject itself changes: from what happened to a person, to what happened to an account, or to a message.
Which raises the obvious problem: “a different subject” depends entirely on what you call things. Call the subject “a message” and email and SMS are one. Call them “an email” and “an SMS” and they are two. A rule you can steer by choosing a word is not a rule.
A river's subject is the noun its rows point at. Nothing is named, so nothing can be argued.
“Noun” here means a thing that exists in its own right — a person, a message, a course. It gets a full treatment as the first of the two kinds, in the next section; for now, a thing rather than an event.
Both of those message logs point at rows in the same notifications table. So the system already decided the subject, and it decided “a notification” — the moment somebody built one notifications table covering both channels. No word had to be chosen.
And it settles the harder cases the same way. Two campaign tables are two separate nouns, so campaigns are two subjects — and whether they should be one noun is a different question, about the nouns, which has to be answered first.
If two logs point at the same noun, they are recording the same thing twice, in two places, with two vocabularies. Anyone asking “what happened to this?” has to remember to look in both — and one day somebody won't.
The events follow the noun. One noun, one river. To justify two rivers you must first argue the noun should be two tables — and win that on the noun's own terms.
Which is why the two message logs merge and the two campaign tables don't: there is one notifications table covering both channels, and there are two campaign tables, deliberately. Same rule, opposite answers, no special pleading needed.
The reverse case, and it's subtler, because a single line legitimately mentions several nouns. “Levan paid £240 for the coaching package” touches a person and a product. That is not two subjects.
Take the noun away and read the line again. Still means something? It was a reference. Means nothing at all? That was the subject.
Remove the product and “Levan paid £240” still says something. Remove Levan and “paid £240” says nothing — nobody paid. So the line has one subject and one reference, not two subjects.
It is genuinely two rivers when that test gives different answers for different lines — when some rows collapse without a person and others collapse without a domain. The tell is a subject column filled in for some rows and empty for the rest, where the empty ones are about something else entirely.
In a multi-tenant system every row points at the tenant. “Rows per tenant” is multi-tenancy, not a unit.
Derive it rather than listing it: count how many tables each id column appears on. Here tenant_id sits on more than half of 270 tables, and no real subject comes close. This is not pedantry — measuring the unit against the scope read 595,089 rows per subject on the contact diary and classified a 1.3-million-row log as a heartbeat. Against its real subject it reads 83.
Take a support conversation. Someone sends forty messages back and forth. Does the log get forty lines, or one?
The system this came from had already decided this, and decided well. It writes three — started, handed to a human, resolved — and each carries a count of how many messages there were.
Both extremes are wrong in ways that are easy to miss. Forty lines and one person's history is flooded; you can't see anything else that ever happened to them. No lines at all and you lose that they ever spoke to you. Three lines with a count keeps the shape of the conversation — you can still tell "asked once and left" from "went back and forth ten times" — without drowning everything around it.
The unit is the smallest thing a person would actually want to see on a timeline. For a conversation, that's the conversation — not the keystroke.
And it does not have to stay a matter of taste. A unit is too fine when one subject's rows crowd out everything else on that subject's timeline — and that is a number, not an argument.
Here is the scale, measured rather than guessed. Across 77 rivers and heartbeats in one real system, the 99th percentile of rows-per-subject ran: median 19, 75th percentile 72, 90th 202, 95th 953, worst 31,372. 83% sit under 100. Only three of the seventy-seven cleared 1,000 — and all three turned out to be heartbeats rather than rivers.
So: under 100 is ordinary. 100–1,000 is worth a look. Over 1,000 is almost certainly a heartbeat wearing a river's clothes. The contact diary runs at 83. Measure against the subject, never a scope.
Nothing else in this document answers the question for you. Every river needs its own answer, and it should be written down.
The contact diary has a fixed list of names — payment_recorded, zoom_joined, community_post_created — around 125 of them, each deliberately added. The audit log has no list at all. Its name column takes free text, and this is what got written into it:
Rule five is a fence around rule three. Rule three says that if a fact never reached the river, the fix is to write it — and that is exactly how stage_arrived got in. So write the missing fact only if it is a fact about the person. A conclusion the system drew is not a missing fact, and the river is the wrong place to keep it.
The cost of getting this wrong is not subtle. With those lines included, the thickest thread on the funnel map was arrived → left at 33,917 people, and its reverse right behind it at 31,169 — the engine talking to itself, drawn as though it were customer behaviour.
And the property that makes a river trustworthy at all: rows are never edited. You don't correct history — you add a line underneath. That's why withdrawing consent writes a new record rather than changing the old one. A row that can be overwritten cannot prove what was true last March.
How long the river is kept. This is the decision the archetypes quietly depend on, and the one almost nobody makes on purpose.
A summary inherits it: a rebuild reaches only as far back as the river still goes, so everything older than the window is permanent rather than derived. A heartbeat needs its own, never shared with a river. And a rule that appears nowhere else on this page: anyone proposing to prune a log has to ask what rebuilds depend on that window. Shortening retention looks like a storage decision and is secretly a correctness one.
Some heartbeats are binned within days; typing indicators are nobody's loss. Others are kept forever on purpose — the graph of who watched a webinar and when they left is the analytics. Same archetype, opposite retention, both correct. Which is why it has to be a decision rather than a property.
And it needs a default, or it will be skipped. A decision nobody is forced to make is a decision nobody makes. So: a river with no stated retention is kept forever, and counts as unrivered until somebody states it. Forever is the safe default — it destroys nothing — and calling it unrivered is what stops “forever” becoming a decision by accident.
A summary is allowed to outlive its river, and the industry does this deliberately: keep thirty days of raw rows and years of rollups. But past the river's window that summary is an original, not a copy, and nothing can recompute it. A fine design and a terrible accident — and the only difference between them is whether anyone wrote it down. Which is why the two retentions are always set together.
This page is arranged for looking things up — every kind with its rules, the diagrams, the checklists. If you would rather be walked through the whole idea from nothing, in small steps that build on each other, start with the companion booklet: Rivering, step by step. Same material, forty-two steps, no prior knowledge assumed.
There are two kinds. Everything else is an archetype — a named shape that one of the two takes, carrying extra rules of its own. Knowing which turns a judgement call into a checklist, and a table that is two things at once is where nearly every real problem turns out to live.
The river — what happened. A verb; ordered; never edited.
The noun — what is. A thing; current; may be replaced.
Of the river: the branch and the heartbeat. Of the noun: state, summary and the lock. Every rule an archetype carries is the rule its old kind carried — nothing changed but where the rules live.
Why this replaced a list (2026-09-10). The list had been arguing with itself in two places, both visible in its own text: it said “a heartbeat is a river” and then listed it separately, and it needed a footnote explaining that state and summary did not contradict each other, because their rules were identical but one. Two kinds and a set of archetypes settles both.
If any of this feels abstract, it has been practised since 1494. Accounting is this model, worked out over five centuries by people who could not afford to get it wrong.
Every entry, in order, never erased. Sub-ledgers — receivables, payables — split off when the volume demands it.
Exactly that. A branch is a sub-ledger: the same water, in a channel of its own.
The chart of accounts is the list of things. A balance is never typed — it is what the entries add up to.
The account is a noun; the balance is state. “No noun is ever written directly”, stated as professional practice.
The rules match one for one, in the same order of importance. Never erase an entry — a mistake is corrected by a new, opposite entry, so the history of the mistake survives. A balance is derived, never written. A closed period is frozen even though it could still be recomputed — which is our sealed mark, arrived at independently.
An accountant who typed a balance straight into an account instead of deriving it from the ledger would be committing fraud — the number would no longer be answerable to the entries beneath it. In software the same act is called keeping a counter, and it ships without comment. Three counters were checked in one production system on 2026-09-10; all three had drifted, one by a factor of eight.
Which also settles the “is this general?” question. Accounting is not a special domain that happens to suit this shape. It is the one domain where getting it wrong was expensive enough, early enough, that the shape was found and then enforced by law. Everything else has the same structure and merely tolerates the drift.
Kind is one question. Original or derived is a different one, and they are independent: contacts is an original noun, a state row is a derived noun. Asking them as one question is how second originals get mistaken for caches, and caches for originals.
No noun is ever written directly. Every write goes to a river; every noun is derived. State, summary, the plain noun, contacts itself. Writes go to a river — the main one, a branch, or a heartbeat — and a branch is written to directly, because writing to a branch is writing to a river.
The lock is the archetype where this breaks, and that break is what earns it a name. A lock must REFUSE; refusing is atomic, at the moment of writing; and a derived thing can never refuse. So a lock is claimed by a direct write — the one noun-shaped thing that is written rather than derived.
Archetypes are open; the two kinds are not expected to grow. That asymmetry is the difference between a model and a list. Candidates already visible and recorded rather than adopted: the queue (rows arrive, are claimed, are deleted — not a river, since rows are removed), the link (a noun whose identity is a pair), and the draft (a thing that exists but has not happened yet).
Each kind and archetype below carries the detail that is genuinely its own. Some carry several rules like the river; some have two, because two is all there is. Nothing here is invented for symmetry.
The archetypes read like independent things. Three of these shapes aren't. A noun, its river and its state are one subject described three times, and seeing that makes the rest easier.
A person. Their name, their phone number. The things that are simply true of them.
They registered, opened an email, paid, cancelled. In order, never edited.
They are a customer, on lesson four, in two sequences. Worked out from the river.
Read left to right and the arrow from section 02 is doing the work: the river produces the state; the state never produces the river. And the noun holds only what it is — never what happened to it, and never a state that was written by hand instead of worked out.
They are arranged so that each one can be defined using only what came before it. Read straight down and nothing refers forward.
Which is also a useful thing to notice about the lock. It is the only archetype that cannot be explained on its own — and that is exactly why it is the one people put in the wrong place.
A thing that exists.
A noun has an identity that survives change. Rename a course and it's still the same course, with the same students and the same history. That's what makes it a noun and not an event.
The grammar is not decoration, and it pairs: a noun is a thing; every line in a river is a verb. Paid. Registered. Cancelled. If a row reads as a thing, it belongs here; if it reads as something done, it belongs in a river. That one test settles most tables before you have looked at a single column.
pausedAt and rejectedAt stacked inside a noun are a diary crushed into one row — each slot holds only the most recent occurrence, and everything earlier is gone. The same defect has a larger form in document stores: an array of events embedded inside the parent document is a river living inside a noun. It is recommended practice there when the set is small, bounded and always read with its parent — and it becomes the crushed diary the moment it grows. The tell never changes: can you count and filter those events without loading every parent? But overwriting is not itself the defect — losing history nobody decided to lose is. Someone changes their phone number and you overwrite it; the old one is gone and nobody minds. Dimensional modelling names that deliberately — a “type 1” dimension — and treats discarding history as correct where it has no value. The sending-domain case was a defect because nobody decided: the second rejection reason vanished before anyone knew they wanted it. Overwrite on purpose and it is a decision; overwrite by default and it is a loss you discover much later.A record of something that happened.
Rivers are the only genuinely original tables. Everything else on this page is either something a river is about, something computed from one, or something that isn't history at all.
A river needs the five things in the section before this one: a subject, a unit, a vocabulary, the rules, and a retention decision.
The same water, in a separate channel.
“Levan viewed /pricing at 14:32” is a complete fact. It passes every test a river line passes; it is not a sample of anything. And there are 188 of them per person, which is why it cannot live in the timeline — one afternoon's browsing buries a year of somebody's history.
A branch is still the river. Same subject, same kind of content, ordered, never edited. It has been split off for volume and nothing else — which is exactly what makes it a branch rather than an extraction. The water is the same; only the channel is separate.
Branch or heartbeat — one question settles it. Does a single row state a fact? Yes, and only the quantity is the problem → branch. No, only the pile means anything → heartbeat. A branch is the river's own water in another channel; a heartbeat is a different substance taken out of it.
The same thing, measured over and over.
Someone watching a webinar sends a signal every few seconds. That is not fifty events. It is one event — they watched — sampled fifty times.
“Levan paid £240 on Tuesday” is a complete fact standing alone. “Levan still watching at 14:32:05” means almost nothing alone — only the pile says “he watched for forty minutes.” A river line is a fact; a heartbeat row is a dot on a graph. Prefer that test to “is it the same situation sampled over and over”, which needs you to already understand the table. This one you can answer by reading a single row cold.
A heartbeat is a river. It has a subject, a unit, its rows are never edited, they arrive in order — it passes every structural test for one. And its subject is the same noun the river has: remove Levan from “Levan still watching at 14:32:05” and the line collapses, so the contact is the subject; remove the viewing session and the line still means something, so that was only a mention. It is extracted for volume, not subject — its unit is finer than a timeline wants, which is decision two, not a different river.
Choose the interval from the questions, not from convenience. The interval you pick is the finest question anyone will ever be able to ask. Report at 25/50/75/100 and “who watched more than 90%” is unanswerable forever — not slow, unanswerable, because the samples that could have answered it were binned on the heartbeat's own schedule. Write the questions down first and the interval falls out of them.
The heartbeat had been recording since 27 February. The conclusion was first written to the river on 13 May. Two and a half months in between, in which the samples were collected, aged out, and never summarised. The shape of the event was right; the history was simply gone.
Backfill reaches only as far back as the heartbeat's own retention. Past that the answer exists nowhere, and the honest move is to say so rather than to infer it.
What conclusion does it write to the river · at what interval · was that interval chosen from the questions or from convenience · does its coverage match the heartbeat's own age, counted per person · and if not, how far back can it still be replayed. Four of those five are counts. Do not accept prose for any of them.
Where something stands right now — worked out, not decided.
In one line: where something stands now — worked out from a river or a noun, written only as a consequence of the thing it follows, and always rebuildable. That is the whole archetype. Everything below is detail: where it lives, how it gets written, and the two ways it goes wrong. This is the longest section on the page because it is the one that has absorbed every correction; you can stop after the first line and be right about most tables.
This archetype was called “settings” in an earlier draft, and that name was wrong twice over. A setting is one example of state, not the category — and most state is nothing like a setting.
State and summary are the only two shapes derived data comes in, and the line between them is time.
A football scoreboard during the match reads 2–1. True right now, and about to be replaced — someone scores and it becomes 2–2. That is state: one answer, always current, always replaceable.
The referee blows the whistle at 3–1. The number hasn't changed, but something about it has: it can never be replaced again. That match ended 3–1, forever. That is a summary.
Which means the two are exhaustive, not a list somebody collected. Every derived answer describes some moment, and a moment is either finished or it isn't. There is no third kind of moment.
And “not finished” includes moments that haven't started. A forecast — projected revenue for next month — is state: it will be replaced as data arrives. Reading state as “about now” is narrower than the rule allows. The line is closed or not closed, and the future is not closed.
And the case that proves the line is real: “views this month, so far.” It looks like a summary, but the month isn't over, so the number still moves — it is state. The instant the month closes, the same row becomes a summary and will never change again. Same data, different kind, because time passed.
Can a thing have only ONE of these? Then it is a column on the noun. Levan has one “last email opened,” the way he has one phone number. Same shelf, no separate table.
Which is the normal case, and worth saying plainly: a noun ordinarily carries its own state, as columns. Giving state a table of its own is the exception, not the rule — reached for only when the noun cannot hold it. Everything on this page still applies to those columns: worked out from a river, never written by hand, rebuildable. They simply live on the noun rather than beside it.
Can it have MANY? Then it cannot fit on the noun and needs its own table. Levan is in five courses, at a different point in each — there is no way to write five answers in one column.
That is a question about shape, not about kind. Both are state; both follow every rule below. The library version: a member's address goes on their card, but a borrowing needs its own slip, because a slip is about a member and a book and neither one can hold it alone.
State arrives two ways, and the difference tells you where history gets destroyed.
Produced state is safe by nature. Nobody decided “Levan is on lesson four” — finishing lesson three put him there. The events exist whether anyone planned for them or not, so the value can always be rebuilt.
Chosen state has no such guarantee. Somebody picked 500 from a dropdown. Unless somebody deliberately wrote that choice down, nothing recorded it — and the previous value is simply gone.
Chosen state is where the log tends to be missing. Produced state is where it cannot be.
But it is a warning label, not a verdict. Chosen state tells you where to look; it does not decide what is lost. Auditing a real system, produced state turned out to go missing just as easily — a media file's uploads and repairs, a live event's every go-live. Nobody chose those, and nothing recorded them either.
What actually decides it is simpler, and it is the rule worth carrying:
History survives exactly when a river exists whose subject is that thing. Chosen or produced only changes how likely it is that nobody built one.
State is read constantly, and when a person changes something they must see it change immediately. So it cannot wait for a rebuild — which raises the obvious worry: if it is written at the same instant, isn't that the whiteboard all over again?
No — and the difference is small enough to miss, which is why it needs a name.
Shapes two and three are both consequence, not copy. What separates them is who is waiting, not the mechanism. Only the first is ever wrong — and the summary already grants the same permission when it says it “is allowed to lag”, so the two rules contradicted each other until this was written down.
If the river records “the cap changed from 200 to 500 on Tuesday”, then the current value isn't a fact needing storage. It is simply the most recent change. Settings are state — and you get “who changed this?” for free, instead of as a separate feature somebody has to remember to build.
“Can't that be derived? We can have a system that in parallel saves the settings as well — but shouldn't that be derived from the log we already talked about?”
Yes — and that question is what collapsed “settings” into state, and state into something that mostly isn't a table at all.
A later one narrowed it further: “That case is not settings; it's something else.” Which is right. A setting is chosen. Where someone stands in a course is produced. Both are state; only one of them gets its history for free.
One honest wrinkle. A setting nobody has ever changed has no change to replay, so the very first value has to come from somewhere — the defaults the system ships with. That's a small, real exception, and it should stay small. If your settings tables are large, most of what's in them is not defaults. It's the current state of things people changed, and every one of those changes was an event that should have been written down.
What happened during a period that has finished.
Summaries are legitimate and necessary — nobody can replay a hundred thousand events on every page load. The danger isn't that they exist. It's that they look identical to originals right up until the day they disagree.
What makes it a summary rather than state: it describes a period that has closed. “This page got 412 views on 3 March” became true the moment March 3rd ended, and will be true in fifty years. Nothing will ever replace that row.
That last one is the sharp edge where this archetype meets the heartbeat, and it is easy to miss because both halves look correct on their own. A retention policy looks like a storage decision. It is also a correctness decision: shortening a river's retention silently freezes every copy built from it. Anyone proposing to prune a log has to ask what rebuilds depend on that window.
A row whose existence is the point, not its contents.
Think of a “reserved” card on a restaurant table. It says nothing and holds no facts. Remove it and you haven't forgotten anything — you've caused two parties to be seated in the same seat.
And the name is not borrowed from software. A lock on a river is a real thing: a chamber with gates at both ends that lets one vessel through at a time, then closes behind it. That is precisely this kind's job — admit one claim, refuse the second — and it is worth picturing, because a physical lock makes the rule obvious in a way the software word never does.
“Two people cannot share an email address” looks like a lock and isn't. That is identity — a statement about what a contact is, permanent, part of what makes it that thing. It belongs on the noun and moving it elsewhere would break it.
A lock is a statement about what may happen. One active enrolment. One pending card per step. It is temporary, and it is released when the situation ends.
Does it prevent a second thing, or a second action? A second thing is identity, and it lives on the noun. A second action is a lock, and it lives alone.
Earlier we said a lock must not share rows with something rebuildable, because a rebuild deletes first. The same danger reaches further, and the reason is worth following.
Rivers get pruned — and the water word is the truer one: a river loses its old water at the far end, continuously, and nobody decides to keep the water from last March. Logs are the same. Old lines are deleted to save space, or because keeping personal data forever is a liability.
Both are correct and responsible decisions, and both are decisions about the records. Nobody in that conversation is thinking about guarantees.
So a lock sitting in a pruned river dies of a decision that was never about it. The guarantee expires as a side effect of somebody else's storage judgement, and nothing announces it.
Never on anything that is routinely deleted or rebuilt — a river or a derived table. A noun is neither, which is why identity is safe there.
And when one thing is both, that is two rows in two tables. “We scheduled a call for Tuesday” genuinely happened — someone decided it, and it belongs in the river where it will stay. The guarantee that it cannot be scheduled twice is a separate row in its own table, where nothing will prune it by accident. Not a contradiction; just two different things that were being asked of one row.
Be honest about what kind of rule this is. A rebuild serialised against everything the guarantee protects — an exclusive lock held across the whole operation — leaves no observable gap, and a selective rebuild that spares the lock rows avoids it too. So this is not a correctness theorem. It is a robustness rule: both safe variants work, and both fail silently the first time somebody empties the table the ordinary way — no error, no failing test, double-booking back with nothing to announce it. That is a sufficient reason for the rule. It is simply a different reason than “impossible”, and worth saying which one you mean.
Why a lock must live alone. Two of these archetypes give opposite instructions about deletion. Anything derived must be deletable — its test is literally “delete it and rebuild it.” A lock must not be — its existence is the guarantee.
So a row that is both cannot be rebuilt at all. A rebuild is delete-then-recompute, and for those seconds the lock is guarding nothing. Anyone rebuilding the state would quietly hold the door open while they did it.
And separate rows are not enough — they need separate tables. A rebuild empties a whole table, so claims sharing a table with state get tipped out with it. You could delete selectively and spare the claims, but that holds only until someone empties the table the obvious way, and then double-booking returns with no warning at all.
A table you are allowed to delete must not contain anything you are not allowed to delete.
A lock carries no cargo — because anything worth carrying is eventually worth rebuilding, and rebuilding a lock releases it.
“Cargo duplicates truth” is also true, and it's the shallower reason. This is the one that explains why the collision is dangerous rather than merely untidy.
A mark is a sticker on a table that already has a kind. It is still a noun, or state, or a summary — one rule about it simply reads differently.
A red ball and a blue ball are both balls; you do not invent “red ball” as a new shape. Shape is one question, colour is another. Make a mark into a kind and you have to copy every rule of the underlying kind into it — and the day somebody edits one copy and not the other, they drift.
Another company's system, your own repository's seed data, or the physical world. One rule changes: rebuild becomes refetch — and a refetch is not guaranteed to agree with what you had.
A signature, an audit, a filing. Once an accounting period closes its figures may not be recomputed even though they still could be. One rule changes: rebuild becomes forbidden, and corrections go forward as new entries.
A table may carry both marks, or neither, and it is still exactly one kind. That is the tell that they are a different sort of thing altogether.
If you find yourself wanting a third, check first that it is not a kind you have mislabelled. Two marks each changing exactly one rule is a pattern with two data points, and it is thin. A candidate has already turned up — a counter that hands out order numbers, where the value being claimed is the cargo the lock rule forbids — and it is recorded rather than adopted, for that reason.
Every rule so far applies to a table once somebody has decided what it is. The damage lives in the stretches nobody decided about — and a method that only rules on what it is shown will keep walking past them.
This is not hypothetical. A production table was found incrementing counters from six separate places, drifting eightfold, unnoticed for months. The rule that forbids exactly that — no front door, never edited, never incremented — was already written, word for word. It never fired, because nobody had ever classified the table. It looks like a noun: one row per person per webinar, its own id, proper foreign keys.
The rules were never the missing part. A way to go looking was. Three searches, in order.
Names ending _count, _score, _seconds, _progress, or beginning total_. Those are answers somebody worked out. A table holding them is a summary however it is shaped — own id, foreign keys and all. Shape is not evidence. The columns are.
Search the application for every update, upsert and increment against that table. More than one writer is the finding. A summary has exactly one writer, and it is the rebuild. Six writers means six places that can drift apart, none of which ever recounts.
Count the same fact twice — once from the counter, once from the river — and put the two numbers next to each other. Do not argue the principle. Show the gap.
Watch records held poll answers, CTA clicks, chat messages and an engagement score as running totals, incremented by hand from six places and never recomputed. Counted against the river that holds the same events, the totals had drifted — silently, for months. Every filter asking "did they answer a poll" was returning the wrong people.
Nothing errored. No test failed. That is the ordinary way this fails, and it is why the search matters more than the rule.
A table, a column, an index, a cache. The simplest thing that works is the right thing, and anything extra carries the burden of proof.
A derived table is never built in advance. It is built when something measurably became too slow, and its existence is the record of that measurement. "We will need this later" is how a schema fills up with second originals — each one true on the day it was written and drifting quietly ever after.
Three consequences worth stating outright.
If it holds a fact that lives nowhere else, it is not derived and the razor does not apply to it. What you have is a river or a noun that has been mislabelled — go back to the kinds.
An index, then a view, then a materialised view, then a table. A materialised view is worth reaching for before a table, because nobody can write to one: "derived, never written directly" stops being a promise anybody can break and becomes a property of the thing.
A derived table with no working rebuild is a second original wearing a faster coat. Ship them together or ship neither — and "the rebuild could be written" is not a rebuild.
The commonest way to misapply the razor is to read it as fewest tables win and start removing things. It says something narrower: every table must earn its place — and being measurably needed for speed is a real reason.
So the usual outcome of applying it to a working system is not deletion. It is demotion.
Before: an original — the only place the fact lives.
After: derived — rebuilt from the river, never written directly, kept for speed.
A worked example. An access table answers “can this person watch this?” on every page load. It can be folded from the river — grants minus revokes — so it is not an original. But it sits on the hot path, and replaying somebody's history thousands of times a day is exactly the case a derived table exists for. Delete it and it returns within a week wearing a cache's name.
Rewire before you delete, always. Rewiring is reversible; deleting a table that turned out to be the only copy is not — and “it looked derived” is precisely how a second original gets destroyed. Demotion is most of the work and none of the drama: stop the direct writes, add the rebuild, prove the rebuild agrees with what is there, and only then argue about whether it should exist at all.
Locks are the least intuitive archetype, and obvious once you see the problem they solve.
You send a campaign to 10,000 people. Halfway through the server restarts — a deploy, a crash, anything. The job wakes up and starts again, not knowing it already emailed 5,000 people. Those 5,000 get it twice.
The obvious fix is to check first: before emailing someone, look whether we already did. That works with one worker. You have several, for speed.
Push the idea to its end — everything is derived from the river, including the nouns — and it very nearly holds. It fails in exactly two places, and both are worth knowing, because the argument is good enough to be tempting.
A river narrates what happened. By the time anything is derived from it, the thing has already happened. But refusing has to occur at the moment of writing, and the obvious approach — look first, then write — has a gap in it.
Two people register with the same address in the same instant. Request A looks: nothing there. Request B looks: nothing there. Both write. Both checks were correct; they simply looked before either had written. Rebuild from that river afterwards and you will faithfully rebuild two of them, because two is what happened.
This does not mean the river cannot be the single source. Put the constraint on the river and it refuses the second line itself. But notice what just happened: whatever holds the rule that refuses is the noun, whatever it is called.
So river and noun are two jobs, not necessarily two tables. One narrates; one refuses. You may put both jobs in one table. You may not have zero things doing the second one.
Some columns are not facts that happened — they are values the system invented. An access token. A generated identifier. Anything already printed inside a link that was emailed to somebody.
Rebuild that row and a different token is minted. Every link already sent stops working. The rebuild was faithful to the events and still broke the world, because the value was never in the events.
The fix belongs inside the principle rather than beside it: the event carries the minted value at the moment it is minted. Then the river holds it too, and the row is rebuildable after all. This is a requirement of deriving everything, not an exception to it.
"Everything is in the river" is a claim to measure, not to assume. Count the rows in the noun with no matching line in the river.
Done on a real system, the answer was 24,710 imports and 1,181 migrations — and exactly three ordinary registrations. Bulk history poured in from outside, whose events had happened in somebody else's system or never happened at all. In normal running the river was complete; the hole was the importer.
That number cuts both ways, which is why it is worth having. It refutes the lazy objection ("the river is full of holes") and it names the real work ("close the side door") in the same measurement.
Everything above, used once on a real question — including the two places it nearly went wrong.
Automations put people through a sequence of emails. Two lines go into the contact diary: entered the welcome sequence, and exited the welcome sequence. There are 41,765 of them.
Separately, an enrolment table records who is currently in which automation.
Same fact in two places — so by rule 1, one has to go. All 41,765 diary lines were marked for deletion.
35,541 of the 41,765 had nothing left to be a copy of. The automations had been deleted at some point, and the enrolment records went with them. So for 85% of these lines, the “original” we were deferring to no longer existed.
The arrow had been drawn backwards. And once you ask which one is the original? — section 02 — it was never a judgement call: “entered on Tuesday” is an event, so it's a river line. “Is currently enrolled” is a state. The events are the original, always.
This is the part worth generalising. What looks like one simple table — a list of who's in which automation — is doing three separate jobs:
Who is currently enrolled. Entirely derivable — it's just everyone who entered and hasn't exited.
Which step of the sequence they're on. Not derivable — nothing writes step-moves into the diary.
Stopping one person being enrolled twice by the same trigger. Enforced in the database, and invisible in the code.
Notice only job one is a mirror. Someone seeing that and saying “this is derivable, we can rebuild it” would be right about job one and would silently destroy the other two. People would get enrolled twice, and nothing would explain why.
Before you simplify a table, count its jobs. A table that looks like one thing is the normal case, not the exception.
And there is a procedure for it — it is not a matter of noticing. List three things about the table in front of you. Every write site: two write paths that never overlap are two jobs until proven otherwise. Every read site, and what it filters on: a query filtering on a column most rows leave empty is reading a different job from the one that fills it. Every constraint, unique indexes especially — a uniqueness rule is always a job, and it is the job most often invisible in the code. The enrolment table gave three writers, two readers filtering on different columns, and one unique index. That is exactly the three jobs it turned out to have.
“Which step are they on” genuinely cannot be worked out from the diary. The tempting conclusion is so we have to keep this table as an original.
Rule 3 says that's the wrong conclusion. It isn't a verdict, it's a symptom — and the question is why is the diary missing this? The answer is unremarkable: nothing writes a line when someone moves from step 3 to step 4. Not impossible, just never done.
So the fix is to record step-moves, after which job two becomes derivable like job one, and the whole table can be a mirror.
The goals side of the same system used to have the identical column — a box holding “which step is this person on.” It's gone. Both it and its companion table survive only as comments explaining why they were removed. Position is now worked out by reading the diary.
Same shape, same fix, already shipped — just never applied to automations. Worth checking, when you find one of these: has some other part of the system already solved it?
The lock and the state live on the same rows. By the summary's rule, that means the table cannot be rebuilt — emptying it to recompute it would leave the double-enrolment guard guarding nothing.
So moving the lock out is not a detail of the tidy-up. It's the first step of it, and nothing else can happen until it's done.
Keep all 41,765 lines — they were miscategorised, not wrongly judged. Everything else is a project, and nothing is broken while it waits.
None of these are illustrations. Each was found by asking “which kind is this?” and measured against a live database on 5 September 2026.
The record of who did what to the business is a river. It was being filled automatically with every web request — including typing indicators and background token refreshes, which are heartbeats. Two kinds in one table, and the real entries inherited the heartbeats' 90-day deletion.
Webinar registrations are locks. Six columns were hung on them, duplicating a record they already held a pointer to — including the access token people log in with. Measured where the copies must agree, they don't.
Fourteen tables store counts a river could produce. They are accurate today — but only ever nudged up by one, and not one has a rebuild. So the handful already wrong will stay wrong permanently.
The audit log's name column takes free text, with the identity of the thing glued inside the sentence. So every line is unique — nothing can be counted, and “everything that happened to this ticket” has no answer.
The life of a sending domain lives as fourteen date columns across five nouns. Each column holds one occurrence.
A block of 169,425 alarm rows was going to be moved as one group. 197 of them are still pending — which puts them inside a uniqueness rule, making them locks. The other 169,228 are finished and hold nothing. Same table, same category, and no single operation is right for both.
The finished ones also turned out to carry cargo: 71,783 notifications point back at them through a link the database doesn't know about — so nothing would have refused the delete.
Counted properly, six subjects keep a log: the person (1.5M lines), the message, the conversation, the workflow run, who did what, and the media file — that last one only for its deletion.
Every one of those is about people, or about talking to people. Not one is about the business's own objects — an account, a sending domain, a campaign, a page, a course. Twenty-two tables hold that kind of history in columns instead, each slot keeping only the most recent occurrence.
Sorting all twenty-two by whose history it was produced a total split, not an approximate one. Person, or a person paired with something: safe. The thing itself: overwritten.
Email deliveries and SMS deliveries are recorded in two separate tables. Both hold the same columns — which message, which person, what happened, when — and both point at rows in the same notifications table. Same subject, recorded twice, with two vocabularies.
A column-counting scan missed this entirely: the tables are small, so they share only six meaningful columns — under the threshold. Asking what they point at found it immediately.
Thirteen rivers carry facts about a person. Twelve reach the contact diary. Two more are correctly kept out of it, because they're heartbeats. The funnel map is textbook derived state — deleted and rebuilt on demand.
This is what rivering actually consists of, table by table. Six questions, in order — the first “yes” tells you the kind and archetype, and that tells you the rules.
These are deliberately not in the same order as the kinds and archetypes above. That order teaches — each defined using only what came before. This one identifies — the sharpest question first, so most tables are settled in one or two steps. Same kinds and archetypes, different job.
A kind is not a property of a thing in the world. It is a property of a table in a system. A copy of somebody else's original is derived and external, whatever shape it has.
The same customer is a noun in the operational database and derived state in the warehouse that copies it. So a warehouse dimension carrying valid_from and valid_to is not a noun breaking the no-history rule — it was never a noun here. Same for a read replica, a search index, a cache. Settle the scope first, or you will apply a noun's rules to something that is really a mirror.
0. Name the database you are rivering. A copy of somebody else's original is derived and external.
1. Did something happen? → river — then keep going to 2.
2. Same situation repeated; one row meaningless alone? → heartbeat. No — every row is a fact, but there are far too many per subject for a timeline? → branch (a person's call).
3. Would deleting a row cause something to happen twice? → lock.
4. Can you name the line that would have to be written? → derived (no rebuild yet = derived with a gap). Can't name one? Ask why not — outside this database means external.
5. Has the moment finished? → summary, else state.
6. A thing in its own right? → noun.
None of them? It is unrivered — not a kind, just work nobody has done. Put it on the list.
The full version, with what each answer commits you to:
It asked could you rebuild this. When the river was missing it answered no and walked on — precisely what rule 3 forbids. The first repair asked should you be able to, which fixed that and introduced a worse fault: a question you can steer by how ambitious you feel is not a question. Almost anything could be derived if you were willing to write a river for it — the same flaw this page identifies in “one subject, one river”, arriving by a different door.
So it points instead of asking. Name the line. team_memberships — two columns, a user and a team, nothing recording who joined when — has an obvious line (added to team), so it is state with a missing river, not a mystery. A surname has none, so it is simply what a noun holds. And knowledge_chunks, a document's text chunked and embedded, is state built from a noun: delete every row, run the chunker again, get it back.
Reached for a new log? These three, in order, usually stop you.
The counter-intuitive result of asking these on a real system: the answer was fewer logs, not more. Five subjects looked like they each wanted one — the account, the sending domain, the email stream, the campaign, the plan. They are all the same subject wearing different hats: things inside the business that a person on the team acts upon. One log covers all five, and its shape is already sitting there in the audit trail — who, what, which thing, when.
These are separate from identifying the kind, and they are the two that cost real damage when skipped. Both were learned the hard way, in that order.
If this row stopped existing, what would break — not who reads it, but what stops working?
“Nobody reads it” is a weaker answer than it sounds. The right question is what depends on it. In the audit log, 85,835 rows turned out to be read by nothing at all — every feature that queries that table filters on a column those rows leave empty.
If this can't be worked out from a river — why not?
“It can't be derived” is a symptom, not a verdict. Usually the river is missing something and the fix is to start recording it — after which the copy becomes ordinary derived data.
Where this came from. The general principle — store what happened, compute the rest — is well established and goes by event sourcing. The formal name for state and summary is a materialized view; the formal name for a river's subject is an aggregate root. Your bank balance works this way: it isn't stored anywhere, it's the sum of your transactions — which is exactly why a statement can always show you why the balance is what it is.
Three things here are not in the textbook. Heartbeats — the rule that repeated measurement is not history — are usually ignored, and standard event sourcing happily mixes metrics into the stream. “It can't be derived is a symptom” is an auditing rule; event sourcing assumes the log is complete and offers nothing for the moment you find it isn't. And the unit rule is missing entirely — the literature has almost nothing to say about how big one line should be, which is the first question anyone actually hits.
How these rules were arrived at, and why that matters for trusting them. None was designed. Each came out of a real audit of a real system, usually from something going wrong — and several replaced an earlier rule of mine that turned out to be worse. “Same subject, one river” started as a column-counting scan, which found two merges that were both already deliberately decided and missed the one that was real. “A lock lives alone” came from noticing that two of these archetypes give opposite instructions about deletion. “State is a column, not a table, unless a thing can have many” replaced a category called “settings” that never made sense. A rule that has not yet been wrong about something has not been tested.
On the name. Rivering is coined, and it is coined for one word rather than for itself: unrivered. Most databases have large stretches nobody has decided about, and there was no word for that state — so it never got pointed at, argued about, or scheduled. A problem with no name stays invisible. River · rivered · unrivered; the log itself is a river, and the work of deciding is rivering.
The case that was open, now closed: “external” is a mark, not a kind. Some data is copied from systems the business doesn't own — counts pulled back from Facebook — and equally, seed defaults that ship in its own repository. None of it can be worked out from a river in this database, and question two's answer is “because the source belongs to someone else.” That does not earn a third kind. The source of truth is outside this database attaches to kinds and archetypes that already exist: a borrowed thing is a noun marked external; a borrowed count for a closed month is a summary marked external; a borrowed current number is state marked external. The mark changes exactly one rule — rebuild becomes refetch, and a refetch is not guaranteed to agree with what you had. Everything else about the kind still applies.
And a second mark: “sealed”. A summary can be frozen by an act outside the database — a signature, an audit, a filing, a legal deadline. Once an accounting period is closed its figures may not be recomputed even though they still could be: a correction is posted as an adjusting entry in the next period, never as a rebuild of the closed one. A sealed summary has stopped being derived and become an original. The mark changes one rule — rebuild becomes forbidden, and corrections go forward as new entries — and everything else stays true.
Two marks, both orthogonal to kind, both changing exactly one rule. If you find yourself wanting a third, check first that it isn't a kind you have mislabelled.
When reality disagrees with the derivation, write a line — don't edit the copy. The shelf says 1,150 and the computed stock says 1,200. The shelf wins, and that is not the arrow running backwards. The fix is a new river line — counted 1,150 on the 31st, shrinkage −50 — after which the derivation produces the right answer on its own. An observation of the world is an event like any other, and it belongs where events belong.
For the reader who wants the arguments checked. A third companion, What Rivering Proves, states five of this page's claims as propositions with proofs, and labels the rest — empirical, heuristic, definition, policy — with what would falsify each. It is deliberately not the place to start. It exists because writing it changed four things on this page, including downgrading one rule from a theorem to a robustness argument.
How this stands against the established methods. Five of them do the same move — a small fixed vocabulary every table must fit. Kimball (fact, dimension), Data Vault (hub, link, satellite), Anchor Modeling (anchor, attribute, tie, knot), event sourcing (event store, read model), and the ERP split of master / transaction / reference data. This is the sixth attempt, not the first.
The convergence is the evidence. Data Vault's satellite and Anchor's attribute table both exist to keep a thing's history out of the thing itself — which is the noun's rule 1, reached independently by people solving auditability and schema evolution rather than this. A rule that three unrelated communities land on is load-bearing, not taste.
What differs is the job, not the ideas. All five are prescriptive — build a new warehouse this way. This is diagnostic: point it at a database somebody already grew by accident. Four of the five are warehouse-layer; this targets the operational database, which almost nothing classifies. And none of them has a heartbeat, a lock as a named shape rather than a constraint, a unit rule, or the auditing posture that “it can't be derived” is a symptom.
Where they are better, and the position on each. Formality — Anchor is 6NF with proofs; the answer is not to chase proofs but falsifiability, every rule checkable by someone who is not an expert. (The unit rule became a number; question four became “name the line”.) Temporality — Anchor handles it by construction, so borrow it: time is structural, not a rule you remember, and two times belongs in decision one. Tooling — theirs generate schemas, ours diagnoses one; that is a different fight and the lead worth extending. Performance — out of scope, said out loud. This method has nothing to say about query speed; use a dimensional model for that. The refusal is what makes the other three credible.
And the real competition is not those five. The one thing aimed at the same reader is a folk taxonomy from beginner tutorials — entity, lookup, junction, transaction, audit/history. It maps cleanly onto these kinds and archetypes (entity → noun; transaction and audit → river; lookup → a noun marked external; junction → state, usually plus a lock) and has no heartbeat, summary or lock at all. More to the point, it sorts by shape — “a junction table has two foreign keys” — where these sort by job. A shape carries no rules. That is why nothing hangs off their list and everything hangs off this one.
How far this has been tested — the denominator. Say this plainly whenever the kinds are called complete, because the sample is small. The list was derived from, and stress-tested against, one relational database of 270 tables, in one company, built by one team's habits. It has since been run against seven domains nobody here designed: double-entry accounting (which produced the sealed mark), perpetual inventory (which confirmed it — the physical count is an event, so you write a line), bitemporal records in banking, insurance and health (which produced the two-times rule), time-series rollups (which confirmed that a summary outliving its river must be a decision), graph databases (independent confirmation — Neo4j's own guidance splits “a thing that exists independently” from “an event between two things”, which is noun and river in other words), warehouse star schemas (which produced the scoping rule above), and document stores (where the concept holds but the mechanism degrades — nothing enforces a reference, and events can be embedded inside their parent). No domain has yet required a third kind. The archetypes are open and have grown once — the branch, after a production river was found carrying 317,000 rows of page views that fitted no existing shape.
What that does not license: it has not been run against anything shaped for a regulator rather than an application, nor against a system where the same fact lives in two stores at once — Postgres and Redis, say. The “one original” rule is silent about which store holds it, and the scoping rule is only half an answer. Anyone calling the list universal is claiming more than has been measured. The honest sentence is: two kinds held across one relational system of 270 tables and seven outside domains; the archetypes are open and grew once, with the branch; everything else that did not fit became a mark, a rule, or a scoping question — never a kind.