[SPARK-58281][ML][CONNECT] Avoid parent overcounting in PipelineModel size estimates#57451
[SPARK-58281][ML][CONNECT] Avoid parent overcounting in PipelineModel size estimates#57451zhengruifeng wants to merge 6 commits into
Conversation
| "copy should create an instance with the same parent") | ||
| } | ||
|
|
||
| test("PipelineModel estimated size") { |
There was a problem hiding this comment.
The new regression test covers only a single-stage pipeline whose one stage is a Model (StringIndexer). The case _ => 0L branch; the novel, behavior-changing part of this PR (the old default walk counted non-model stage bytes; the new code skips them) is never exercised. The assertion is also upper-bound-only (estimatedSize < 16 KiB), so a zero-returning implementation would pass it. Consider adding a pipeline that mixes a Model stage with a non-model Transformer stage, plus a lower-bound assertion and a hasParent-preserved check, mirroring the SPARK-57521 tests in ModelSuite (ModelSuite.scala:36-47).
uros-b
left a comment
There was a problem hiding this comment.
Thank you @zhengruifeng, I left just one comment - otherwise LGTM
f01d197 to
88e5126
Compare
What changes were proposed in this pull request?
Adds a
PipelineModel.estimatedSizeoverride that sums model-stage estimates individually and directly estimates non-model transformer stages. This prevents a reflection walk of the complete pipeline object graph.Adds a regression test that fits a pipeline containing
StringIndexerand asserts its size estimate remains below 16 KiB.Why are the changes needed?
SPARK-57521 clears the top-level model parent before the default size walk. A copied
PipelineModelstill contains copied model stages whose estimator parents are retained, so a reflection walk can still reach shared Spark-session state. Estimating nested model stages independently applies their parentless-copy behavior at every model boundary.Does this PR introduce any user-facing change?
Yes. Pipeline model cache-size estimates no longer include shared state reachable through nested stage parents, preventing unnecessary cache overcounting.
How was this patch tested?
StringIndexerpipeline size-estimation regression test.JAVA_HOME=/usr/lib/jvm/java-17-openjdk-amd64 build/sbt 'mllib/Test/compile'\n\n### Was this patch authored or co-authored using generative AI tooling?\n\nGenerated-by: Codex GPT-5