Designs by Duhart All writing

·5 min read·SystemDesign · DistributedSystems · GraphDatabases · DataEngineering

Highlander Rules: Can there only be one? Identity management in distributed systems

Why I stopped treating a user as a row and started treating them as a node with edges into every product.

One box labelled Account in the centre, with four lines out to boxes labelled Dating, Billboard, BoxOffice and Chat.
TL;DR. One person signed in to my platform exists in three identity stores across Dating, Billboard, BoxOffice and chat, and the stores did not agree on who that person was. Model the person as a node and each product relationship as an edge, and the disagreements become visible instead of silently empty screens.

Series: The graph you already have

  1. This post
  2. Groundhog Day: Why did my retry only happen once? Kafka consumer redelivery semantics (coming 6 Oct)
  3. Seinfeld: A read about nothing? Successful queries that quietly return empty results (coming 7 Oct)
  4. Two-Face: One status code, two meanings? Query APIs that separate NotFound from Unimplemented (coming 8 Oct)
  5. Fellowship of the Ring: One diagram to rule them all? Four product failures mapped as a graph (coming 9 Oct)

The short version

The first reason to reach for a graph is not scale, it is identity. When one person acts across several products, every product keeps its own idea of who they are, and the joins between those ideas are where data goes missing without an error. Draw the person as one node and the products as edges and you can see which edge is broken.

Background

I built a music streaming platform (Billboard), a video service (BoxOffice), a dating app with real-time chat, and the services between them, all behind one login. It is a one-person platform with a small real user base, running on one server plus Cloudflare.

The login is shared. The idea of "the user" is not. Postgres has a Django table called public.accounts_identity. The Go services target a different table, public.identities. The dating gateway reads a third thing at request time, a Cassandra table called users_root, in a different database from the other two.

Scope

What this covers

  • How one person became three identity records, and how that showed up as empty screens
  • The entity and edge model that replaced the joins
  • What Netflix's Part 1 journey (one member, streaming then games on another device) has in common with it

What it does not

  • Real-time ingestion, storage and query. Posts 2, 3 and 4 of this series cover those
  • Any claim about Netflix's internals beyond what the post itself says
  • Scale. Netflix writes millions of records a second. I have a handful of real users

The challenge

The failure never looked like an identity bug. It looked like an empty profile. A real account opened its dating profile on the web and saw a blank page. The profile existed: display name, bio, prompts and preferences were all sitting in the match_cards table in Cassandra. The web app read localStorage and nothing else, and fell back to a blank profile on a miss. The obvious tables, dating_profiles and dating_photos, had no row for that user at all. And every photo URL on the card was a file:// path on the phone that uploaded it, so the media could never cross devices. Nothing threw. Every read returned a valid, empty answer.

A box labelled One login on the left with three arrows to three tables on the right: public.accounts_identity (Postgres, Django) read by the Django backend, ticked; public.identities (Postgres, Go) read by the Go services and users_root (Cassandra) read by the dating gateway, both on red dashed arrows and marked with a cross.
One login, three tables that each think they are the user.

Solutions and process

  1. Attempt one: find the right table. I looked in dating_profiles first because the name said profile. Wrong. The match card IS the dating profile, and about 21,000 candidate cards live in match_cards while about 40,000 rows sit in dating_profiles. A user can be in one population and not the other.
  2. Attempt two: look the card up by user id. match_cards is partitioned by (service_region, geohash4), and nothing indexed the user id, so a profile screen could not find its own owner's card. A lookup table, match_card_by_user, fixed that, backfilled with 20,998 rows. It is a hint, never the source of truth, because four write paths across two services insert into match_cards.
  3. The reframe: stop asking which table holds the user. Draw one account node, then an edge per product: owns-card (Dating), owns-feed-profile (Billboard), watches (BoxOffice), member-of (chat). Every screen that showed nothing was an edge that nobody had written, or an edge pointing at the wrong node.
  4. The rule that fell out of it, which I set for the feed service: one profile per mode under one login. The account is the node. Each product's profile is its own node, joined by an edge, never merged into one record, so a visitor to your music profile cannot walk across to your dating profile.
  5. What I did not do: build a graph database. At this size the edges are rows in the stores that already exist. What changed is the model I reason with, and the checks I write against it.

The model, as types. One account node, one profile node per product, and edges that say who wrote them.

TypeScript
type Mode = "dating" | "billboard" | "boxoffice" | "chat";

type Node =
  | { kind: "account"; id: string }
  | { kind: "profile"; mode: Mode; id: string };

type Edge = {
  from: Node;
  to: Node;
  rel: "owns" | "follows" | "matched" | "messaged";
  writer: string; // the one service allowed to write it
  at: number;     // ms since epoch
};

const owns = (
  accountId: string,
  mode: Mode,
  profileId: string,
  writer: string,
): Edge => ({
  from: { kind: "account", id: accountId },
  to: { kind: "profile", mode, id: profileId },
  rel: "owns",
  writer,
  at: Date.now(),
});
type Mode = "dating" | "billboard" | "boxoffice" | "chat";

type Node =
  | { kind: "account"; id: string }
  | { kind: "profile"; mode: Mode; id: string };

type Edge = {
  from: Node;
  to: Node;
  rel: "owns" | "follows" | "matched" | "messaged";
  writer: string; // the one service allowed to write it
  at: number;     // ms since epoch
};

const owns = (
  accountId: string,
  mode: Mode,
  profileId: string,
  writer: string,
): Edge => ({
  from: { kind: "account", id: accountId },
  to: { kind: "profile", mode, id: profileId },
  rel: "owns",
  writer,
  at: Date.now(),
});
An Account node in the centre with four edges: owns-card to Dating profile (writer MatchService), drawn red and dashed and marked never written; owns-feed-profile to Billboard profile (writer GlobalFeedGateway); watches to BoxOffice (writer BoxOfficeGateway); member-of to Chat (writer GlobalChatService).
A blank screen is an edge nobody wrote.

What it gets wrong

Takeaway

Where this meets Netflix's engineering

How and Why Netflix Built a Real-Time Distributed Graph: Part 1, Ingesting and Processing Data Streams at Internet Scale: Its example journey is a member who watches Stranger Things and then plays Stranger Things: 1984 on a tablet, which is why the graph spans streaming, games and devices.

The shape is the same, one person acting across several products and devices; the difference is everything else, since Netflix ingests millions of records a second into a purpose-built graph and I have a handful of real users whose edges are still rows in the stores that already existed.

Over to you

Identity and platform engineers: when one person uses several of your products, where does the canonical 'who is this' live, and who is allowed to write the edges?

References

  1. How and Why Netflix Built a Real-Time Distributed Graph: Part 1
  2. ByteByteGo: How Netflix built a real-time distributed graph
  3. Designs by Duhart, live demos

Next: the Kafka consumer that skipped every error it was supposed to retry, and why 'error means redeliver' is an assumption.

More: LinkedIn · Instagram. Portfolio and case studies: designsbyduhart.org.

If any of this saved you an afternoon, Buy me a coffee.