Cassandra UPDATE is an upsert, and it cost us a phantom partition per user
Users reported cards in their dating deck with no name and no photo. Not broken images — cards for people who did not appear to exist. The deck was being served correctly from the table it read. The table was wrong.
The bug in one line
The ranking pipeline iterates regions and updates candidate scores. The write looked roughly like this:
// for each region being processed:
session.Query(`UPDATE deck_cards SET score = ? WHERE region = ? AND user_id = ? AND card_id = ?`,
score, region, userID, cardID).Exec()
region there is the region the loop is currently processing. The card's
own region — the one it was written under — is a different value. In
PostgreSQL this is a WHERE clause that matches nothing: zero rows updated,
possibly a metric somewhere, no damage.
Cassandra has no such thing as an UPDATE that matches nothing. UPDATE is
an upsert. The partition key here is (region, user_id), so an update naming
a region the row does not live in does not fail and does not no-op. It
creates the partition. A new row, in a new partition, containing only the
primary key columns and the one column you set. Every other column is null.
That is the nameless card. It was a row whose entire content was the fact that something had once written a score into a partition that should never have existed.
Why it survived review
Nothing about the statement is wrong-looking. It is a parameterised update against a table that exists, with a full primary key, and it returns no error. The test suite passed, because the tests seeded a single region and the loop variable and the row's region were therefore always equal. The bug requires two regions to exist to be visible at all, and the local fixtures had one.
It is also invisible to the type system, because both values are string.
Nothing distinguishes "the region I am iterating" from "the region this row
belongs to" except intent.
The fixes, in order of how much they help
- Read the region from the card, not the loop. The immediate fix, one line. It stops new phantoms.
- Delete the existing ones. A phantom partition is identifiable — every non-key column is null. A scan-and-delete pass cleared them. This is unglamorous and unavoidable; the bug had already written its garbage.
- Make the two regions different types. A
CardRegionand anIterationRegion, both wrapping a string, both zero-cost at runtime, and the compiler now rejects the original statement. This is the one that stops it recurring. The one-line fix does not survive the next person who writes a similar loop. - Assert in the test fixtures that more than one region exists. The tests could not have caught this. That is a property of the fixtures, not of the tests, and it was worth fixing at the fixture layer.
The general lesson
Every database has an operation whose failure mode is silence, and the ones
that bite are the operations you have used a thousand times somewhere else.
UPDATE ... WHERE returning "nothing happened" is such a deep assumption that
it does not present itself as an assumption at all.
The way to find these is not to be more careful. It is to ask, for each write: if the key is wrong, what happens? If the answer is "an error", you are fine. If the answer is "a new row", you have a class of bug that will not announce itself.
The full case study is the match service.