Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
55 commits
Select commit Hold shift + click to select a range
93d1174
Reading-guide depth rules, and reading-criterion.md rewritten against…
DvirDukhan Aug 3, 2026
0a83d06
Review feedback: no length cap, collapsed answers, line-anchored snip…
DvirDukhan Aug 4, 2026
2e0d65b
Merge remote-tracking branch 'origin/master' into dvirdukhan-turbo-eu…
DvirDukhan Aug 5, 2026
a4b9f3e
Make the reading-guide depth rules checkable, and anchors checkable o…
DvirDukhan Aug 5, 2026
3c11daa
Convert topic 0's four remaining reading guides to the depth standard
DvirDukhan Aug 5, 2026
3bfad18
Convert topic 9's four reading guides to the depth standard
DvirDukhan Aug 5, 2026
ed92009
Convert topic 6's six reading guides to the depth standard
DvirDukhan Aug 5, 2026
9e3e778
Convert topic 1's eight reading guides to the depth standard
DvirDukhan Aug 5, 2026
ebc666b
Convert topic 5's five reading guides to the depth standard
DvirDukhan Aug 5, 2026
8768c95
Convert topic 3's five reading guides to the depth standard
DvirDukhan Aug 5, 2026
71f0e9d
Convert topic 11's six reading guides to the depth standard
DvirDukhan Aug 5, 2026
b0864e5
Convert topic 10's five reading guides to the depth standard
DvirDukhan Aug 5, 2026
b11d07c
Convert topic 8's six reading guides to the depth standard
DvirDukhan Aug 5, 2026
9a9f27b
Convert topic 4's six reading guides to the depth standard
DvirDukhan Aug 5, 2026
bcde8b5
Convert topic 2's seven reading guides to the depth standard
DvirDukhan Aug 5, 2026
17826da
Convert topic 7's five reading guides to the depth standard
DvirDukhan Aug 5, 2026
ff95e36
topic 12: six reading guides to the depth standard
DvirDukhan Aug 5, 2026
9795d20
SESSION-LOG: the topics 1-12 depth rollout
DvirDukhan Aug 5, 2026
6790cd7
topic 14: six reading guides to the depth standard
DvirDukhan Aug 5, 2026
48dd209
topic 13: seven reading guides to the depth standard, and a headline …
DvirDukhan Aug 5, 2026
6f52b5f
topic 19: six reading guides to the depth standard
DvirDukhan Aug 5, 2026
7459cf9
topic 16: six reading guides to the depth standard
DvirDukhan Aug 5, 2026
d96483e
topic 21: five reading guides to the depth standard
DvirDukhan Aug 5, 2026
4eaedbd
topic 15: six reading guides to the depth standard
DvirDukhan Aug 5, 2026
7daaa1c
topic 22: four reading guides to the depth standard
DvirDukhan Aug 5, 2026
38cf112
topic 20: seven reading guides to the depth standard
DvirDukhan Aug 5, 2026
0e05b8b
topic 17: seven reading guides to the depth standard
DvirDukhan Aug 5, 2026
65eb8b9
topic 18: six reading guides to the depth standard
DvirDukhan Aug 5, 2026
2f0514d
experiments: stop ignoring Cargo.lock
DvirDukhan Aug 5, 2026
38ac3e0
SESSION-LOG: batch 2 of the reading-guide depth rollout (topics 13-22)
DvirDukhan Aug 5, 2026
5d74e45
topic 30: four reading guides to the depth standard
DvirDukhan Aug 5, 2026
0978cbe
topic 25: seven reading guides to the depth standard
DvirDukhan Aug 5, 2026
a0451c5
topic 27: five reading guides to the depth standard
DvirDukhan Aug 5, 2026
bf71ce0
topic 24: six reading guides to the depth standard
DvirDukhan Aug 5, 2026
72c5dc9
topic 29: four reading guides to the depth standard
DvirDukhan Aug 5, 2026
0c22728
topic 28: five reading guides to the depth standard
DvirDukhan Aug 5, 2026
8401fd6
topic 33: four reading guides to the depth standard
DvirDukhan Aug 5, 2026
cf0eb04
topic 32: four reading guides to the depth standard
DvirDukhan Aug 5, 2026
6d7bf76
topic 31: four reading guides to the depth standard
DvirDukhan Aug 5, 2026
f8c034f
topic 26: seven reading guides to the depth standard
DvirDukhan Aug 5, 2026
387c10f
topic 23: six reading guides to the depth standard
DvirDukhan Aug 5, 2026
95d1bbe
SESSION-LOG: batch 3 of the reading-guide depth rollout (topics 23-33)
DvirDukhan Aug 5, 2026
72beea7
topic 34: four reading guides to the depth standard
DvirDukhan Aug 5, 2026
691b597
topic 36: four reading guides to the depth standard
DvirDukhan Aug 5, 2026
7567bc9
topic 35: four reading guides to the depth standard
DvirDukhan Aug 5, 2026
b65bf04
topic 39: four reading guides to the depth standard
DvirDukhan Aug 5, 2026
191c1b4
topic 38: four reading guides to the depth standard
DvirDukhan Aug 5, 2026
f18eaa7
topic 37: four reading guides to the depth standard
DvirDukhan Aug 5, 2026
fc5ea3d
topic 41: four reading guides to the depth standard
DvirDukhan Aug 5, 2026
08e76d6
topic 43: four reading guides to the depth standard
DvirDukhan Aug 5, 2026
cbd4a23
topic 42: four reading guides to the depth standard
DvirDukhan Aug 5, 2026
5f02d8d
topic 40: four reading guides to the depth standard
DvirDukhan Aug 5, 2026
1b51550
reading guides: make the collapsed answers actually render
DvirDukhan Aug 5, 2026
863208b
depth rules: batch-4 log entry, and the gate turned on for all 230 gu…
DvirDukhan Aug 5, 2026
67c53a2
book workflow: the depth job's comment described the ratchet it no lo…
DvirDukhan Aug 5, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
16 changes: 16 additions & 0 deletions .github/workflows/book.yml
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,22 @@ concurrency:
cancel-in-progress: true

jobs:
# The reading guides' depth rules (CLAUDE.md § Reading-guide depth) have a
# mechanical part — each step declaring its input and output, a collapsed
# answer under every checklist item, a line-number gutter on every quoted
# snippet. Across 230 guides those survive only if a script enforces them.
# `--all` drops the ratchet the rollout ran behind: every guide is converted,
# so a file that does not follow the rules is a new one that skipped them
# rather than one the rollout has not reached. Nothing here needs a
# toolchain, so it runs first and fast.
depth:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7

- name: Reading guides follow the depth rules
run: python3 tools/check-reading-depth.py --check --all

build:
runs-on: ubuntu-latest
steps:
Expand Down
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -3,3 +3,4 @@ target/
/book/
/mermaid.min.js
/mermaid-init.js
/.cache/
17 changes: 17 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,23 @@ A self-paced database-internals learning path, rendered as an mdBook (`book.toml
- **Generators are seeded**, so every figure reproduces exactly apart from timings. Lockfiles are committed for the same reason.
- **Exercise lanes must degrade, not crash.** A bench binary on a fresh clone prints its provided lanes and a `[stub — ...]` note for the rest, and exits 0. Never let a `todo!()` panic hide a measurement above it.

## Reading-guide depth

The `reading-*.md` chapters teach from zero: a reader who knows systems but not the chapter's theory must be able to finish without leaving the page. `topics/00-performance-toolbox/reading-criterion.md` is the reference implementation of these rules; match it.

- **Define every term at first use.** A term of art (t-test, p-value, quartile, IQR, MAD, standard error, null hypothesis, arithmetic intensity, coordinated omission) gets a **bold** name and a one-sentence plain-language definition at the point it first appears, *before* any argument leans on it. A step may use only terms defined in an earlier step or defined on the spot. Borrowed jargon — using a word the guide never defined because the source material used it — is the failure this rule exists to stop.
- **Every step declares its input and output.** Each `### Step N` opens with a `> **In:** … **Out:** …` blockquote naming the dataset it consumes, *which earlier step produced it*, and what it emits. When one stage forks into two datasets used by different downstream steps, the fork gets its own numbered step. "Is this the same data as the previous section?" must never be left to the reader to infer.
- **A formula gets its symbols named and one worked example.** Quote it as the source actually computes it, name every symbol, then run it once on 3–5 concrete numbers so a real answer comes out. Arithmetic printed in a guide is verified like any other number in this repo — compute it, don't estimate it.
- **Anchors are verified against the pinned clone, file *and* line.** Citing the right line of the wrong file is the same error class as inventing a number. Re-grep every anchor before committing; state the version the line numbers belong to.
- **A quoted snippet carries the line numbers it actually occupies, and names the one to look at.** Put the real number in the gutter of every line, mark elided ranges (`// ... 131–139: bookkeeping ...`) rather than silently closing a gap, and say in the prose which line carries the argument ("the line to focus on is 277, its only `return`"). A snippet anchored to the function signature while quoting code forty lines below it leaves the reader unable to find anything. Pseudocode gets a `// ILLUSTRATION — not quoted from the crate` header and a pointer to the real code.
- **Describe what the code does, not what the technique usually does.** criterion 0.5.1's `Slope::fit` is a one-field struct fitting through the origin, so the textbook "the intercept absorbs the overhead" account of least squares is simply false there. Read the implementation before writing the explanation, and prefer the honest, weaker claim over the tidy, wrong one.
- **Every `## Done when` box carries its answer in a collapsed `<details>` block**, introduced by "Answer each before unfolding it." The checklist is a self-test, so the answer must be reachable without leaving the page but never visible by accident. Answers restate the reasoning rather than pointing back at a step number, and are held to the same standard as the body: real anchors, real numbers, the honest claim.
- **Never trade a definition, a worked example or an answer for brevity.** These chapters have no length target — a guide that assumes vocabulary is not shorter, it is unfinished. Cut redundancy instead.

The mechanical half of these rules is enforced by `python3 tools/check-reading-depth.py` (step In/Out blockquotes, collapsed answers under every `Done when` item, line-number gutters on quoted snippets, the section spine); run it on a guide before committing it. `--check` is a ratchet — a guide that has started following the rules must follow all of them, and the guides the rollout has not reached yet are reported without failing. The other half — definitions, worked arithmetic, honest claims — is judgement and stays the writer's job.

Anchors are checked with `python3 tools/pinned-source.py`, which opens a file at the revision the pin table records (`show`, `grep`, `check`, `list`). It uses a real clone under `~/repos` when one is present and otherwise fetches that same commit into a gitignored `.cache/`, so an anchor can be verified on a machine that has not cloned 85 upstream repos. A repo that is not in the pin table — a crate read from the cargo registry — needs `--ref` and a version stated in the guide.

## Topic package shape

Each `topics/NN-name/` contains: `README.md` (study guide, opening with *the problem, measured* — the provided benchmark lane's real output), four to seven `reading-*.md` guides in the concept-first format (framing lead → "the problem in one sentence" → numbered `### Step N` sections → how to read the source → questions → `## Done when` checklist → references), `notes.md` (a `## Baseline (provided lane, <machine>, measured <date>)` section recording the real output, *then* the reader's prediction worksheet — leave those cells empty, they are the exercise), and `experiments/` — a Rust crate with **lane 1 implemented and two lanes stubbed**, where the stub tests are the specification and the reference numbers live in `notes.md`.
Expand Down
32 changes: 30 additions & 2 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,16 @@ Every topic follows the same shape, because the shape is what makes it checkable
format: an H1 with the idea in it, a framing lead, **the problem in one sentence**,
then `### Step N` sections that build each concept using only terms defined in
earlier steps, then how to read the source material with the concepts in hand,
questions to answer, a "done when" checklist, and references.
questions to answer, a "done when" checklist, and references. The depth rules
that make "using only terms defined in earlier steps" enforceable — define every
term at first use, declare each step's input and output, work every formula on
concrete numbers — are in
[CLAUDE.md](https://github.com/AviAvni/database-learning-path/blob/master/CLAUDE.md#reading-guide-depth)
(an absolute link because `CLAUDE.md` is not a book chapter), with
[reading-criterion.md](topics/00-performance-toolbox/reading-criterion.md) as the
reference chapter. Their mechanical half is checked by
`python3 tools/check-reading-depth.py` — run it on a guide before committing, and
see [Building the book](#building-the-book) for the CI gate.
- **`notes.md`** — a `## Baseline (provided lane, <machine>, measured <date>)` section
recording the provided lane's real output with the analysis, then a
predictions-vs-measurements worksheet. **The worksheet's cells are meant to be
Expand Down Expand Up @@ -80,6 +89,18 @@ legitimate shape for a topic; inventing a number to fill the slot is not.
`file:line` anchors in the guides mean something. Regenerate it with
`python3 tools/pin-table.py` after cloning or updating a reference repo — putting
a SHA in each guide instead would mean thousands of them drifting separately.
To read a file at that pinned commit — to write an anchor, or to check one that
is already there — use `python3 tools/pinned-source.py`:

```bash
tools/pinned-source.py show lmdb mdb.c -r 1350:1365 # with real line numbers
tools/pinned-source.py grep lmdb 'mdb_env_pick_meta' --path mdb.c
tools/pinned-source.py check lmdb mdb.c:1356 --contains 'meta page'
```

It reads your clone when you have one and otherwise fetches the same commit into
a gitignored `.cache/`, so anchors stay checkable without cloning every upstream
repo the guides cite.
- **Generators are seeded.** Anyone must be able to reproduce a figure exactly.
- **Notes capture *why* a design wins** and what it trades away — not summaries.

Expand All @@ -94,7 +115,14 @@ mdbook serve # or: mdbook build
CI ([.github/workflows/book.yml](.github/workflows/book.yml)) builds HTML and PDF on
every push to `master` and deploys to GitHub Pages. Before committing content, build
locally and check that mermaid diagrams render and internal links resolve — a broken
link is invisible in markdown and obvious in the book.
link is invisible in markdown and obvious in the book. The same workflow runs
`tools/check-reading-depth.py --check`, which holds every reading guide that has
started following the depth rules to all of them:

```bash
python3 tools/check-reading-depth.py topics/03-btree-internals/ # one topic
python3 tools/check-reading-depth.py --stats # rollout progress
```

A second workflow ([verify.yml](.github/workflows/verify.yml)) runs
`./verify.sh --summary` and a `-D warnings` build of all 45 crates on every push and
Expand Down
8 changes: 4 additions & 4 deletions FINDINGS.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,7 +28,7 @@ instead.
| 9 | [Concurrency](topics/09-concurrency/README.md) | A global mutex gets **2.9× slower** from 1 to 16 threads (8.65 → 2.96 Mops/s). Padding "independent" counters to 128 B is worth **17.8×**; 64 B only half-fixes it on M-series. | `./verify.sh 09` |
| 11 | [Execution Models](topics/11-execution-models/README.md) | Volcano tops out at **103 M rows/s**, and gets *slower* as selectivity rises (74.7 M at 95%) — surviving the filter is what costs, not the filter. | `./verify.sh 11` |
| 12 | [Columnar Analytics](topics/12-columnar-analytics/README.md) | The scan floor is **24–57 GB/s** on a 150 GB/s machine. This lane previously printed **19,047,619 GB/s** — a hoisted loop, caught by its own implausibility. | `./verify.sh 12` |
| 13 | [Graph Engines](topics/13-graph-engines/README.md) | The same two-hop query is **101× slower** from supernodes than from random nodes (4.9 µs → 495 µs) — and reaches *fewer* distinct nodes. | `./verify.sh 13` |
| 13 | [Graph Engines](topics/13-graph-engines/README.md) | The same two-hop query is **101× slower** from supernodes than from random nodes (4.9 µs → 495 µs), because it reaches **77× more** nodes per query — 78,907 against 1022. | `./verify.sh 13` |
| 14 | [Vector Search](topics/14-vector-search/README.md) | Brute force: **117 QPS** at recall 1.000. That single point is what every ANN index is betting against. | `./verify.sh 14` |
| 15 | [Replication & Consensus](topics/15-replication-consensus/README.md) | Follower fsync policy alone spans **59×** (341 → 20,174 entries/s). Batching fixes the median and leaves the p99 at 2980 µs. | `./verify.sh 15` |
| 16 | [Testing & Correctness](topics/16-testing-correctness/README.md) | Seeded crash testing catches planted bugs at **48.8% to 99.6%** per seed — same harness, four wildly different odds of ever finding out. | `./verify.sh 16` |
Expand All @@ -38,7 +38,7 @@ instead.
| 20 | [GraphBLAS](topics/20-graphblas/README.md) | SpMV bandwidth decays **20.7 → 12.3 GB/s** as the graph grows. Hypersparse indexing is **50× smaller** (80.4 MB → 1.59 MB) and sweeps rows **175× faster**. | `./verify.sh 20` |
| 21 | [Formal Methods](topics/21-formal/README.md) | The hand-ordered rewriter answers `(a*2)/2` with `(a << 1) / 2` and stops. One locally-excellent rewrite destroys the cancellation — the phase-ordering trap, in four lines. | `./verify.sh 21` |
| 22 | [Standard Benchmarks](topics/22-benchmarks/README.md) | TPC-H Q1 and Q6 measured at **5.2–5.7** and **9.0–14.4 GB/s** effective; YCSB-E's p999 is **12.9 µs** against read-only's 4.0 µs. | `./verify.sh 22` |
| 23 | [Full-Text Search](topics/23-fulltext/README.md) | Exhaustive BM25 spans **0.009 ms to 10.378 ms** across four two-term queries — 272,310 postings against 159. Term rarity, not query complexity. | `./verify.sh 23` |
| 23 | [Full-Text Search](topics/23-fulltext/README.md) | Exhaustive BM25 spans **0.009 ms to 10.378 ms** across four queries — 272,310 postings against 159. Term rarity, not query complexity. | `./verify.sh 23` |
| 24 | [Graph Algorithms](topics/24-graph-algorithms/README.md) | Same node and edge count, RMAT vs uniform: **15.6 M triangles vs 5428**, and 447 ms vs 195 ms. Degree skew is the workload. | `./verify.sh 24` |
| 25 | [Graph ML](topics/25-graph-ml/README.md) | The message-passing kernel *is* an SpMM: **4.31 ms at 16.82 GFLOP/s**, against 5.65 ms for the dense transform beside it. | `./verify.sh 25` |
| 26 | [Probabilistic Structures](topics/26-probabilistic/README.md) | A point miss costs **246 ns** (binary search) or **299 ns** (BTreeMap); a 224 MB HashSet does it in **28 ns**. That gap is what a filter is bidding for. | `./verify.sh 26` |
Expand All @@ -47,7 +47,7 @@ instead.
| 29 | [Distributed Transactions](topics/29-distributed-txn/README.md) | The workload's own conflict rate goes **0.3% → 99.6%** as Zipf θ moves 0.5 → 1.3. Contention is a property of the data, before any protocol. | `./verify.sh 29` |
| 30 | [Time-Series](topics/30-timeseries/README.md) | delta+varint gives **11.00 B/sample for all four shapes** — a constant series compresses exactly as well as random noise, because only the timestamp is being compressed. | `./verify.sh 30` |
| 31 | [CRDTs](topics/31-crdts/README.md) | Last-write-wins on 10 keys with per-write sync loses **94.98%** of writes — 37,991 of 40,000 acknowledged writes that no replica remembers. | `./verify.sh 31` |
| 32 | [HTAP](topics/32-htap/README.md) | One copy, one coarse lock: adding full scans takes writes from **10.5 M per 2 s to 94**, and p99 from 334 ns to **2.7 s**. Every scan is a write outage. | `./verify.sh 32` |
| 32 | [HTAP](topics/32-htap/README.md) | One copy, one coarse lock: adding full scans takes writes from **11.4 M per 2 s to 69**, and p99 from 333 ns to **7.49 s**. Every scan is a write outage. | `./verify.sh 32` |
| 33 | [Temporal Graphs](topics/33-temporal-graphs/README.md) | Static reachability reports 25,031 reachable pairs where time-respecting paths number **137** — **99.5% false positives** on the sparse contact graph. | `./verify.sh 33` |
| 34 | [Debugging & Diagnosis](topics/34-debugging/README.md) | A closed-loop benchmark reports **p99 = 1.0 µs** where an open-loop one reports **90 ms** on identical work — coordinated omission, a 90,000× lie. | `./verify.sh 34` |
| 35 | [Overload Control](topics/35-overload/README.md) | A 10-second outage ends at t=40 s. At 140 QPS (of 300 capacity) goodput stays at **zero until t=161 s**; at 280 QPS it **never recovers** — the outage outlives its own trigger. | `./verify.sh 35` |
Expand All @@ -57,7 +57,7 @@ instead.
| 39 | [Fraud & Identity Graphs](topics/39-fraud-identity-graphs/README.md) | Two row-based rankers fail in *opposite* regimes: degree ranking scores **0.00** precision without camouflage, obscurity ranking **0.00** with it. | `./verify.sh 39` |
| 40 | [Security & Attack Graphs](topics/40-security-attack-graphs/README.md) | A directory reporting **8 privileged accounts, forever** has **1969 of 2000 users** holding a path to Domain Admin — and your exposure number depends on how long the collector ran. | `./verify.sh 40` |
| 41 | [On-Chain Analytics](topics/41-onchain-analytics/README.md) | The industry-default haircut rule marks **98% of addresses** tainted from one theft; 658 of them are under 0.1% tainted. An 1816 court case does better. | `./verify.sh 41` |
| 42 | [Recommendations & Social](topics/42-recommendations-social/README.md) | Recommending bestsellers to everyone gets **35.3% hit-rate@50** with **92.2% overlap** between users' lists. Popularity is not a weak baseline. | `./verify.sh 42` |
| 42 | [Recommendations & Social](topics/42-recommendations-social/README.md) | Recommending bestsellers to everyone gets **34.0% hit-rate@50** with **92.3% overlap** with the global bestseller list. Popularity is not a weak baseline. | `./verify.sh 42` |
| 43 | [Ops Dependency Graphs](topics/43-ops-dependency-graphs/README.md) | One gray failure: **34 of 55 services alert** and the broken one is not among them — it ranks 35th by failure count, 41st by error rate, at exactly the baseline. | `./verify.sh 43` |

## How to read this table
Expand Down
Loading