From d10e2b4b2cd04c5fb27b312ff1f7b304f421ab7a Mon Sep 17 00:00:00 2001 From: Greg Burd Date: Wed, 29 Jul 2026 12:56:01 -0400 Subject: [PATCH] test(coverage): wire recd group into COV_RECD block for *_rec.c branches The per-operation recovery handlers (__xxx_recover: redo/undo/abort/ do-nothing branches) in db/db_rec.c (3189 br), hash/hash_rec.c (2069), btree/bt_rec.c (2366) and qam/qam_rec.c (574) were the biggest cold BRANCH surface. The subset ran only recd001/002/016 on btree, leaving the hash and queue handlers at 0% and most btree/db redo/undo branches cold. Add a COV_RECD block (default on) that drives the recd group driver-per-test (each in its own tclsh, per-test timeout, recdscript.tcl orphan cleanup) -- recd tests use conflicting globals (recd006 kvals scalar vs recd010 array) so they cannot share one interpreter. The set is curated for coverage-per-minute across every access method each test supports; whole block ~10 min, every test PASSes well inside a 300s timeout. Recd-driven branch % (measured, gcc --coverage, -O0): db/db_rec.c 14.6 -> 16.4 btree/bt_rec.c 12.3 -> 21.6 hash/hash_rec.c 0 -> 24.2 qam/qam_rec.c 0 -> 45.5 recd001/002/016 removed from COV_TESTS (superseded by the driver-per-test block; recd001 was 6.5 min and added ~1 branch over the rest). recd003/ 007/015 deliberately excluded: too slow under -O0 for the marginal branches they add (recd003 does not unlock the still-cold __bam_root/ irep/rcuradj or __db_ovref handlers). No recovery-correctness bug surfaced: all forward+backward replays verified clean. --- test/coverage/README.md | 52 +++++++++++++++++++++- test/coverage/run_coverage.sh | 81 ++++++++++++++++++++++++++++++++++- 2 files changed, 130 insertions(+), 3 deletions(-) diff --git a/test/coverage/README.md b/test/coverage/README.md index eba636bf3..10f76a8ad 100644 --- a/test/coverage/README.md +++ b/test/coverage/README.md @@ -37,6 +37,7 @@ Knobs (all optional env vars): | `COV_XA_UPG` | `1` | run the XA + on-disk-upgrade drivers (fast, non-hanging) | | `COV_BACKUP` | `1` | run the hot-backup-API + compaction-recovery drivers (fast, non-hanging) | | `COV_DEAD_REG` | `1` | run the deadlock-detector + DB_REGISTER multi-process tests | +| `COV_RECD` | `1` | run the recovery-record handler tests (`recd` group, driver-per-test) | Set `COV_REP=1` to add the replication suite (biggest cold surface: rep/ + repmgr/ ~= 12.4k lines at 0.8%). It moves them to ~56% line / ~42% branch. Those @@ -210,6 +211,52 @@ worker must not wedge the whole run): (`src/env/env_register.c`, `__envreg_register`/`__envreg_isalive`) and runs recovery. Lift: `env_register.c` **0%→~55%**. +## Recovery-record handlers (`COV_RECD`) + +The per-operation recovery handlers — `__xxx_recover` in `db/db_rec.c` +(3189 branches), `hash/hash_rec.c` (2069), `btree/bt_rec.c` (2366), +`qam/qam_rec.c` (574) — are the biggest cold **branch** surface in the tree. +Each handler has redo / undo / abort / do-nothing branches, and the main +`COV_TESTS` loop used to run only `recd001/002/016` on **btree**, so the hash +and queue handlers sat at **0%** and most redo/undo branches of the btree/db +handlers were cold. + +`COV_RECD` (on by default) drives the `recd` group. `op_recover` +(`testutils.tcl`) crashes the env at every log-record boundary and replays +recovery forward (redo) **and** backward (undo/abort) — exactly the cold +branches. Like `COV_REP`/`COV_DEAD_REG` these run **driver-per-test** (each in +its own `tclsh`, per-test timeout, `recdscript.tcl` orphan cleanup): several +`recd` tests use conflicting globals (e.g. `recd006` sets `kvals` as a scalar, +`recd010` uses it as an array) so they cannot share one interpreter. + +The set is curated for **coverage-per-minute** across every access method each +test supports — splits (`recd002` on btree/hash/queue/recno), file-id reuse +(`recd005` ×4 AMs), nested txns (`recd006` btree/hash), deep/many-child txns +(`recd008`), recnum stability (`recd009`), off-page-dup splits (`recd010`), +cursor adjust (`recd013` btree/hash), queue-extent create/delete (`recd014` +queueext, the only thing that reaches `__qam_delext_recover`), checksum-error +(`recd016`), crypto (`recd017`), checkpoint/commit (`recd018`), txn-id wrap +(`recd019`), intermediate dirs (`recd020`), aborted-prepared page-alloc +(`recd022`), reverse split (`recd023`), streaming partial-put overflow +(`recd024`), TXN_BULK (`recd025`), big-key-to-internal (`recd004`). Whole +block ~10 min. Lift (recd-driven branch %): + +| file | before | after | +|------|--------|-------| +| `db/db_rec.c` | 14.6% | **16.4%** | +| `btree/bt_rec.c` | 12.3% | **21.6%** | +| `hash/hash_rec.c`| 0% | **24.2%** | +| `qam/qam_rec.c` | 0% | **45.5%** | + +**Deliberately excluded** (too slow for the marginal branches they add under +`-O0`): `recd001` (~6.5 min/method — the other tests already cover its per-op +redo/undo branches; dropping it entirely moved the totals by ~1 branch); +`recd003` (~7.3 min, dup — does **not** unlock the still-cold +`__bam_root/irep/rcuradj` or `__db_ovref` handlers, which need scenarios no +bounded `recd` test reaches); `recd015` (~4.6 min — prepared-txn branches +already hit by `recd002`/`recd022`); `recd007` (~4.3 min — db_rec +create/rename branches already covered by the compaction + upgrade drivers). + ## MVCC freeze/thaw + cache resize (`mvcc001`) `mp/mp_mvcc.c` and `mp/mp_resize.c` sat near-cold (functional baseline ~9%) @@ -447,13 +494,16 @@ why they lag the single-process access-method tests: | 0.0 | 0.0 | 1017 | 644 | `log/log_verify_util.c` | log verification | | 0.0 | 0.0 | 923 | 942 | `rep/rep_util.c` | replication | | 0.0 | 0.0 | 864 | 540 | `repmgr/repmgr_sel.c` | replication manager | -| 0.0 | 0.0 | 836 | 2069 | `hash/hash_rec.c` | hash recovery records | | 0.0 | 0.0 | 704 | 556 | `repmgr/repmgr_net.c` | replication manager | | 0.0 | 0.0 | 701 | 556 | `repmgr/repmgr_msg.c` | replication manager | | 0.0 | 0.0 | 556 | 656 | `rep/rep_elect.c` | replication elections | | 0.0 | 0.0 | 541 | 460 | `rep/rep_automsg.c` | replication | | 0.0 | 0.0 | 431 | 470 | `sequence/sequence.c` | sequences | | 5.4 | 2.0 | 445 | 306 | `qam/qam_files.c` | queue extent files | +| 40.1 | 24.2 | 836 | 2069 | `hash/hash_rec.c` | hash recovery records (was 0%, COV_RECD) | +| 41.6 | 21.6 | 1015 | 2366 | `btree/bt_rec.c` | btree recovery records (12.3%→, COV_RECD) | +| 34.1 | 16.4 | 1342 | 3189 | `db/db_rec.c` | db recovery records (14.6%→, COV_RECD) | +| 66.7 | 45.5 | 291 | 574 | `qam/qam_rec.c` | queue recovery records (was 0%, COV_RECD) | | 30.1 | 17.4 | 1397 | 1702 | `btree/bt_compact.c` | btree compaction (was 0%) | | 54.7 | 43.1 | 223 | 174 | `env/env_register.c` | DB_REGISTER / failchk (was 0%, COV_DEAD_REG) | | 57.1 | ~ | 394 | 398 | `xa/xa.c` | XA transactions (was 0%, COV_XA_UPG) | diff --git a/test/coverage/run_coverage.sh b/test/coverage/run_coverage.sh index 2fedb8950..2deff7add 100755 --- a/test/coverage/run_coverage.sh +++ b/test/coverage/run_coverage.sh @@ -104,8 +104,11 @@ bld="$root/build_unix" # compression only ever marshals 32-bit lengths); they are exhaustively # property-tested by test/pbt/pbt_compint.c, a separate tier not in this # subset -- so db_compint's ceiling here is ~24%, not 100%. -: "${COV_TESTS:=lock001: txn001: ssi001: ssi002: recd001:btree \ - recd002:btree recd016:btree env007: \ +# +# recd001/002/016 are NOT listed here: the recd recovery-record group is run +# by the dedicated COV_RECD block below (driver-per-test, timeout-guarded), +# which covers *_rec.c branches across all access methods. +: "${COV_TESTS:=lock001: txn001: ssi001: ssi002: env007: \ lock007: \ btree/test001 btree/test111 test143:btree \ hash/test001 hash/test006 hash/test010 hash/test025 hash/test077 \ @@ -381,6 +384,80 @@ if [ "${COV_DEAD_REG:-1}" = 1 ]; then echo " .gcda files after dead/register: $(find . -name '*.gcda' | wc -l)" fi +# --- optional recovery-record handlers (COV_RECD=1, default on) -------------- +# The per-operation recovery handlers in db/db_rec.c, hash/hash_rec.c, +# btree/bt_rec.c and qam/qam_rec.c (__xxx_recover: redo / undo / abort / +# do-nothing branches) are the single biggest cold BRANCH surface +# (db_rec.c 3189 branches, hash_rec.c 2069, bt_rec.c 2366, qam_rec.c 574). +# The main COV_TESTS loop runs only recd001/002/016 on btree, so the hash and +# queue recover handlers sat at 0% branch and most redo/undo branches of the +# btree/db handlers were cold. +# +# The `recd` group drives exactly those branches: op_recover (testutils.tcl) +# crashes the env at every log-record boundary and replays recovery forward +# (redo) AND backward (undo/abort). They CANNOT share one tclsh -- several +# recd tests use conflicting globals (e.g. recd006 sets `kvals` scalar, +# recd010 uses it as an array) and each spawns recdscript.tcl subprocesses -- +# so, like COV_DEAD_REG, they run driver-per-test with a per-test timeout and +# orphan-subprocess cleanup. +# +# The set is curated for coverage-per-minute across the FOUR access methods +# each test supports (splits recd002, fileid-reuse recd005, nested-txn recd006, +# deep/many-child recd008, recnum recd009, off-page-dup splits recd010, cursor +# adjust recd013, queue-extent create/delete recd014, checksum-error recd016, +# crypto recd017, checkpoint/commit recd018, txn-id-wrap recd019, intermediate +# dirs recd020, aborted-prepared page-alloc recd022, reverse-split recd023, +# streaming partial-put overflow recd024, TXN_BULK recd025, big-key-to-internal +# recd004). Lifts (recd-driven branch %): db_rec.c 14.6->16.4, bt_rec.c +# 12.3->21.6, hash_rec.c 0->24.2, qam_rec.c 0->45.5. Whole block ~10 min. +# +# DELIBERATELY EXCLUDED (too slow for the marginal branches they add): +# recd001 (~6.5 min/method): the other tests already cover its per-op +# redo/undo branches -- dropping it entirely changed the totals by ~1 +# branch. recd001:btree/recno each add nothing the set doesn't have. +# recd003 (~7.3 min, dup): does NOT unlock the cold __bam_root/irep/rcuradj +# or __db_ovref handlers (they need scenarios no bounded recd test hits), +# so its 7 minutes buy only a handful of already-adjacent branches. +# recd015 (~4.6 min, many prepared txns): prepare/recover branches are +# already exercised by recd002/recd022's prepare paths. +# recd007 (~4.3 min, file create/delete): db_rec create/rename branches are +# covered by the recd_compact + upgrade drivers already in the subset. +# Set COV_RECD=0 to skip. +if [ "${COV_RECD:-1}" = 1 ]; then + echo "== run recovery-record handler tests (COV_RECD=1) ==" + : "${COV_RECD_TIMEOUT:=300}" + # Each entry is "test:method:extra-arg" (extra blank for tests taking none). + cov_recd_tests=( + "recd002:btree:0" "recd002:hash:0" "recd002:queue:0" "recd002:recno:0" + "recd004:btree:" + "recd005:btree:" "recd005:hash:" "recd005:queue:" "recd005:recno:" + "recd006:btree:" "recd006:hash:" + "recd008:btree:" "recd009:btree:" "recd010:btree:" + "recd013:btree:" "recd013:hash:" + "recd014:queueext:" + "recd016:btree:" "recd017:btree:" "recd018:btree:" "recd019:btree:" + "recd020:btree:" "recd022:btree:" "recd023:btree:" "recd024:btree:" + "recd025:btree:" + ) + recdtcl="$bld/.cov-recd-one.tcl" + for spec in "${cov_recd_tests[@]}"; do + t="${spec%%:*}"; rest="${spec#*:}"; m="${rest%%:*}"; a="${rest#*:}" + printf 'source ../test/tcl/test.tcl\nsource ../test/tcl/%s.tcl\nif {[catch {eval %s %s %s} r]} { puts "FAIL %s %s: $r"; exit 3 }\nputs "PASS %s %s"\n' \ + "$t" "$t" "$m" "$a" "$t" "$m" "$t" "$m" > "$recdtcl" + timeout "$COV_RECD_TIMEOUT" "$TCLBIN" "$recdtcl" >/tmp/cov-recd-$t-$m.log 2>&1 + rc=$? + # kill orphan recdscript.tcl subprocesses a hung/killed test may leave + pkill -f "$recdtcl" 2>/dev/null || true + pkill -f 'recdscript' 2>/dev/null || true + if [ $rc -eq 124 ]; then echo "HANG $t $m" + elif [ $rc -eq 0 ] && grep -q "^PASS $t $m" /tmp/cov-recd-$t-$m.log; then echo "PASS $t $m" + else echo "FAIL $t $m (rc=$rc)"; fi + find TESTDIR -mindepth 1 -delete 2>/dev/null || true + done + rm -f "$recdtcl" + echo " .gcda files after recd: $(find . -name '*.gcda' | wc -l)" +fi + # --- aggregate --------------------------------------------------------------- echo "== capture (lcov) ==" # NOTE: capture from .libs (not .) -- libtool double-compiles, and only the