perf(db): index the foreign keys and list sort columns - #317
Conversation
Every index in the schema was a primary key or a UNIQUE constraint. None
covered a foreign key, so lookups by project_id, date-ordered list pages,
and ON DELETE CASCADE checks were all sequential scans of the child table.
Measured on a local Postgres 16 loaded with 500k expenditures, 50k
memberships, 50k donations, 20k reports, 5k users, 2k projects:
project-filtered expenditure page 6846 buffers, 48.4ms -> 28 buffers, 0.8ms
unfiltered expenditure page 6846 buffers, 23.8ms -> 28 buffers, 1.2ms
GET /projects for a non-admin 406 buffers, 6.2ms -> 34 buffers, 1.1ms
project-filtered report page 230 buffers, 7.8ms -> 15 buffers, 2.3ms
GET /projects/{id}/donors 423 buffers, 7.4ms -> 82 buffers
DELETE one project (FK cascade) 179.2ms -> 42.0ms
DELETE one user (FK cascade) 127.4ms -> 38.8ms
Seven indexes, ~11MB total at that row count. All additive, so they are
safe for the currently deployed code.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This PR contains a database migration
It will be applied to the production database automatically when this PR merges, before the new lambda code is deployed. Please confirm before requesting review:
|
|
Database Types Check Complete The database schema files were modified, but the regenerated TypeScript types are identical to the existing ones. No changes were needed and the type definitions are already up to date. |
Both lessons from the index audit: aggregating in the lambda costs a full scan no index can fix, and new filter/join/sort columns need an index because Postgres does not create them for foreign keys. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
Database Types Check Complete The database schema files were modified, but the regenerated TypeScript types are identical to the existing ones. No changes were needed and the type definitions are already up to date. |
Why
The schema had eleven indexes, and every one of them was a primary key or a UNIQUE constraint. Not one covered a foreign key — Postgres doesn't create those for you.
That means every lookup by
project_id, every date-ordered list page, and everyON DELETE CASCADEcheck was a sequential scan of the entire child table. This migration adds the seven indexes the code actually asks for.How the access patterns were determined
Full inventory of every query in
apps/backend/lambdas/*andshared/lambda-auth/, plus the two open PRs (#310, #315) so this doesn't get invalidated the moment they merge.Findings that shaped the index set:
project_membershipsandproject_donationsboth have a UNIQUE constraint whose leading column is the wrong one for how they're queried. UNIQUE is(project_id, user_id)butGET /projectsfor a non-admin looks up byuser_idalone — and that's the landing page of the app. Same story forproject_donations: UNIQUE is(donor_id, project_id), but everything filters onproject_id.project_membershipsfilter on bothproject_idanduser_id, so the existing UNIQUE already serves them. No index needed there.authenticate.ts:55(users WHERE cognito_sub = $1) runs on every authenticated request, butcognito_subis already UNIQUE. Nothing to do.project_idfilter withORDER BY spent_on DESC, hence the composite rather than a bare FK index — the second column removes the sort node entirely.Measurements
Local Postgres 16, schema built from the migrations in this repo, loaded with 500k expenditures / 50k memberships / 50k donations / 20k reports / 5k users / 2k projects.
EXPLAIN (ANALYZE, BUFFERS), warm cache. Buffer counts are the honest metric here — wall-clock on a warm local container flatters everything.GET /projectsfor a non-adminGET /projects/{id}/donorsDELETEone project (FK cascade)DELETEone user (FK cascade)The expenditure list went from reading ~53 MB per request to ~220 KB.
For the cascades, the time was almost entirely in one FK trigger:
expenditures_project_id_fkeyalone was 146 ms of the 179 ms project delete, andexpenditures_entered_by_fkeywas 108 ms of the 127 ms user delete.Cost
Seven indexes, ~11 MB total at the row counts above.
expenditurescarries three of them, so its write path takes three extra index maintenance ops per insert — acceptable for a table that is read on nearly every page of the app and appended to a few times a day.Notes on what is deliberately not here
expenditures.status. The new admin review queue in feat(expenses): admin approve/deny review flow with receipt upload #315 filters by status client-side — it fetches the list and narrows topendingin the browser. There is no SQL predicate onstatus, so an index (even a partial one) would never be used. If that filter moves server-side,CREATE INDEX ... ON expenditures (status, spent_on) WHERE status = 'pending'becomes worth adding;pendingis a ~15% minority of rows, so a partial index would be well-targeted.expenditures.category. The dashboard's two aggregates (projects/handler.ts:75and:83) are inherently full scans — one is aGROUP BY project_idover the whole table, and the other'scategory IS NOT NULLpredicate matches ~100% of rows. Measured 86 ms and 51 ms, and no index changes that.projects/handler.ts:83streams every expenditure row into the Lambda and buckets it by month in a JS loop; that endpoint needs a server-side aggregate, not an index. Worth a follow-up issue.NOT NULL, so Postgres reads the index backwards (Index Scan Backward) and gets the ordering for free — confirmed in the plans above. ADESCindex would buy nothing.Safety
Adding an index is pure expand: no column changes, no live
INSERTbreaks, and the currently deployed code only gets faster. Satisfies the "safe for the code that is live right now" rule inapps/backend/db/README.md.Plain
CREATE INDEX, notCONCURRENTLY— the migrator wraps the run in a single transaction, and CI rejectsCONCURRENTLYfor that reason.Verification
migrations-freshdoes it.seed.sqlapplies on top of the migrated schema.migrations-guardfilename regex and the destructive-SQL/CONCURRENTLYscan locally; both pass.shared/types/db-types.d.tschange — indexes don't alter generated column types.🤖 Generated with Claude Code