diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/470415967.md b/tools/import_validation/troubleshooting_guides/bugs_summary/470415967.md new file mode 100644 index 0000000000..f8131191e3 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/470415967.md @@ -0,0 +1,7 @@ +# Bug 470415967 Summary + +## Problem Description +Discrepancies in GDP data were reported between Data Commons and the World Bank website, specifically for Nigeria, Kenya, and Hungary. Investigations revealed that the source data at the World Bank had been updated recently, causing a mismatch with the existing data in Data Commons. A refresh run failed to download the updated data, and subsequent runs encountered validation failures, including issues with deletion checks and missing reference errors. + +## Resolution +The issue was resolved by correcting DCID mappings for the Channel Islands and Kosovo (changing XKX to XKS), which resolved validation and reference errors in the import pipeline. Following these corrections, the data import pipelines were successfully re-executed, which synchronized the data with the updated World Bank records. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/472258775.md b/tools/import_validation/troubleshooting_guides/bugs_summary/472258775.md new file mode 100644 index 0000000000..949146cbe7 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/472258775.md @@ -0,0 +1,7 @@ +# Bug 472258775 Summary + +## Problem Description +An automated refresh of World Bank datasets led to validation failures due to missing observations for multiple countries. Root cause analysis revealed issues related to country DCID naming and name discrepancies at the source. + +## Resolution +The team implemented a workaround to retain historical data for indicators that were deleted from the source run. A PR was submitted to apply the changes, resolving the validation blocks. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/472605922.md b/tools/import_validation/troubleshooting_guides/bugs_summary/472605922.md new file mode 100644 index 0000000000..0c05d851f7 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/472605922.md @@ -0,0 +1,7 @@ +# Bug 472605922 Summary + +## Problem Description +The auto-refresh for the `USCensusPEP_Sex` import failed. The root cause was identified as intermittent unavailability of data at the source for programmatic download. + +## Resolution +The issue was resolved by manually triggering the pipeline execution once the source data became available. The manual trigger was successful and the import completed without further issues. The bug was marked as fixed after verification. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/472606851.md b/tools/import_validation/troubleshooting_guides/bugs_summary/472606851.md new file mode 100644 index 0000000000..517afdb50a --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/472606851.md @@ -0,0 +1,7 @@ +# Bug 472606851 Summary + +## Problem Description +The automated refresh for the `EurostatData_Education_Attainment` import failed, leading to the deletion of several data points. + +## Resolution +A Root Cause Analysis (RCA) was conducted to investigate the deletions. Following the review of the RCA and implementation of the recommended changes, the job was re-run and confirmed to be working correctly with a satisfactory validation report. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/474326901.md b/tools/import_validation/troubleshooting_guides/bugs_summary/474326901.md new file mode 100644 index 0000000000..d009fc928f --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/474326901.md @@ -0,0 +1,7 @@ +# Bug 474326901 Summary + +## Problem Description +The auto-refresh job for Eurostat Education Enrollment failed. Analysis in the differ tool was performed to investigate the discrepancies and data series deletions. + +## Resolution +It was determined that no schema or code modifications were required. The issue was resolved by updating the `latest_version.txt` file (following approval) and successfully re-running the job, yielding a clean differ report. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/478186511.md b/tools/import_validation/troubleshooting_guides/bugs_summary/478186511.md new file mode 100644 index 0000000000..99a21ccf9f --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/478186511.md @@ -0,0 +1,7 @@ +# Bug 478186511 Summary + +## Problem Description +A failed data import was reported for `USCensusPEP_Sex`. Investigation revealed that while local runs were clean, automated job runs showed inconsistent results, including deletions of data series that were not reproducible locally. This anomaly suggested a discrepancy between the local and automated environments or intermittent source data issues. + +## Resolution +The discrepancies were documented in an anomaly tracking spreadsheet for further analysis. After reviewing the diff outputs and confirming they met acceptable criteria despite the inconsistencies, it was decided to proceed with the current versions. The issue was marked as fixed after verification. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/479399481.md b/tools/import_validation/troubleshooting_guides/bugs_summary/479399481.md new file mode 100644 index 0000000000..263170d553 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/479399481.md @@ -0,0 +1,7 @@ +# Bug 479399481 Summary + +## Problem Description +The automated import for `USCensusPEP_Annual_Population` failed during download, which cascade-failed the subsequent data processing stages with deletion errors. + +## Resolution +The import manifest was modified to isolate the download and processing steps, and the code was updated to explicitly fail and log errors if a download fails rather than continuing with incomplete data. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/481243356.md b/tools/import_validation/troubleshooting_guides/bugs_summary/481243356.md new file mode 100644 index 0000000000..ae9eff5673 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/481243356.md @@ -0,0 +1,7 @@ +# Bug 481243356 Summary + +## Problem Description +The import failed due to Eurostat's transition from country/region grouping classification EA20 to EA21. The new `nuts/EA21` ID was missing from the Knowledge Graph (KG), causing linting errors and threshold violations for deleted records (since EA20 was removed). + +## Resolution +The new place ID `nuts/EA21` was added to `ProvisionalNodePlaces.mcf` to address the lint errors. Additionally, `nuts/EA21` was temporarily added to `skip_places.csv`, and a historical data file was created for the deleted `nuts/EA20` records. A CL with these changes was approved and merged. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/481245546.md b/tools/import_validation/troubleshooting_guides/bugs_summary/481245546.md new file mode 100644 index 0000000000..b57681bee3 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/481245546.md @@ -0,0 +1,7 @@ +# Bug 481245546 Summary + +## Problem Description +A recurring failure of the `EurostatData_Education_Attainment` import occurred, linked to a regression from a previous fix (Bug 481243356). + +## Resolution +A new fix was applied, which included the creation of a historical data file and a corresponding code change. The approach was reviewed and approved, and the issue was resolved with clean differ and validation reports provided. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/482902862.md b/tools/import_validation/troubleshooting_guides/bugs_summary/482902862.md new file mode 100644 index 0000000000..7a7a3dfec7 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/482902862.md @@ -0,0 +1,7 @@ +# Bug 482902862 Summary + +## Problem Description +The auto-refresh job for the World Development Indicators import failed due to the deletion of 4 Statistical Variable (SV) observations. Investigation of the source data confirmed that these observations had been deleted from the origin by the World Bank. + +## Resolution +After reviewing the source-side data deletions, it was approved to accept the change and update to the latest version of the data. The import pipeline was updated to reference the latest version, and the job was completed successfully. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/482946661.md b/tools/import_validation/troubleshooting_guides/bugs_summary/482946661.md new file mode 100644 index 0000000000..25ee581d70 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/482946661.md @@ -0,0 +1,7 @@ +# Bug 482946661 Summary + +## Problem Description +There was a failure in the auto-refresh process for imports related to state-level employment data. The issue was specifically caused by a missing reference or incorrect mapping for a statistical variable representing person counts in the transportation sector. + +## Resolution +The problem was resolved by updating the mapping for the statistical variable to a more specific industry classification (NAICS-based). This involved modifying the configuration that maps source data fields to the internal statistical variables. The fix was verified through local testing and the changes were integrated into the main data pipeline. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/483219293.md b/tools/import_validation/troubleshooting_guides/bugs_summary/483219293.md new file mode 100644 index 0000000000..07410768fb --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/483219293.md @@ -0,0 +1,7 @@ +# Bug 483219293 Summary + +## Problem Description +The import job for `USCensusPEP_Sex` encountered a read timeout when attempting to access a source URL. This was a recurring issue caused by the source server taking longer to respond than the default timeout settings in the import script. + +## Resolution +The root cause analysis recommended increasing the timeout duration within the script code to handle slower responses from the source. The issue was resolved by implementing this timeout adjustment, and the bug was marked as fixed. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/485260648.md b/tools/import_validation/troubleshooting_guides/bugs_summary/485260648.md new file mode 100644 index 0000000000..a6ccf85c15 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/485260648.md @@ -0,0 +1,7 @@ +# Bug 485260648 Summary + +## Problem Description +The auto-refresh process for the `USCensusPEP_By_Sex_Race` import failed because the source data URL was intermittently unavailable, causing the import to error out during the scheduled run. + +## Resolution +The issue resolved itself once the source URL became accessible again. Subsequent runs of the data import were successful, and the bug was marked as fixed. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/486801970.md b/tools/import_validation/troubleshooting_guides/bugs_summary/486801970.md new file mode 100644 index 0000000000..47a5642b90 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/486801970.md @@ -0,0 +1,7 @@ +# Bug 486801970 Summary + +## Problem Description +The automated refresh for the `USCensusPEP_PopulationEstimatebyRace` import failed during its scheduled execution. + +## Resolution +The import job was manually re-run and completed successfully. Validation outputs, including a differ log and verification reports, were checked to confirm data integrity. The bug is verified and fixed. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/489948342.md b/tools/import_validation/troubleshooting_guides/bugs_summary/489948342.md new file mode 100644 index 0000000000..ddb96f771b --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/489948342.md @@ -0,0 +1,7 @@ +# Bug 489948342 Summary + +## Problem Description +The import for World Development Indicators failed because a required file (`API_IND_DS28_en_csv_v2_10671`) was missing from the source World Bank data. + +## Resolution +The issue was resolved by updating the `latest_version.txt` file and addressing the deleted flags for India (IND) in the differ report. The pipeline was successfully re-run, and a validation report was generated. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/493190090.md b/tools/import_validation/troubleshooting_guides/bugs_summary/493190090.md new file mode 100644 index 0000000000..d450c6fa3d --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/493190090.md @@ -0,0 +1,7 @@ +# Bug 493190090 Summary + +## Problem Description +The auto-refresh for the `USCensusPEP_PopulationEstimatebyRace` data import failed during its scheduled run. + +## Resolution +The import issue was addressed and fixed. Successful data import was confirmed via validation reports. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/4935600377.md b/tools/import_validation/troubleshooting_guides/bugs_summary/4935600377.md new file mode 100644 index 0000000000..b17e90a383 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/4935600377.md @@ -0,0 +1,7 @@ +# Bug 4935600377 / 493560037 Summary + +## Problem Description +The auto-refresh for the `USCensusPEP_Annual_Population` data import failed due to an error in the execution pipeline. (Note: The bug ID `4935600377` in documentation corresponds to Bug `493560037` in Buganizer). + +## Resolution +This issue was identified as a duplicate of a canonical, systemic failure within the core import pipeline rather than a dataset-specific error. The pipeline issue is currently assigned and being tracked centrally. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/496059688.md b/tools/import_validation/troubleshooting_guides/bugs_summary/496059688.md new file mode 100644 index 0000000000..ed41907b07 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/496059688.md @@ -0,0 +1,7 @@ +# Bug 496059688 Summary + +## Problem Description +A validation failure was encountered during the Eurostat Education Enrollment import process, blocking the pipeline completion. + +## Resolution +The issue was resolved by generating a new validation report and verifying the data consistency. The fix was verified, and the validation checks passed. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/497802532.md b/tools/import_validation/troubleshooting_guides/bugs_summary/497802532.md new file mode 100644 index 0000000000..f3c655ea9a --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/497802532.md @@ -0,0 +1,7 @@ +# Bug 497802532 Summary + +## Problem Description +Intermittent "403 Forbidden" errors occurred when attempting to download data files from census.gov. Specifically, CSV URLs were inaccessible while XLS URLs remained working, causing batch job failures and unintended data deletions. + +## Resolution +The team addressed the intermittent URL accessibility issues. The fix was verified through successful re-runs of the national, state, and county-level population estimate imports. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/498154643.md b/tools/import_validation/troubleshooting_guides/bugs_summary/498154643.md new file mode 100644 index 0000000000..b627fb3325 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/498154643.md @@ -0,0 +1,7 @@ +# Bug 498154643 Summary + +## Problem Description +A validation failure was encountered during the processing of Eurostat regional statistics (`EurostatData_Fertility`) due to additional data deletions in the source data. + +## Resolution +The deletions were reviewed and approved by the team. The data import pipelines were updated and successfully re-executed, and the validation checks passed. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/500622108.md b/tools/import_validation/troubleshooting_guides/bugs_summary/500622108.md new file mode 100644 index 0000000000..1236d4aa73 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/500622108.md @@ -0,0 +1,7 @@ +# Bug 500622108 Summary + +## Problem Description +The automated refresh for the `USCensusPEP_AgeSexRace` data import was failing, and critical execution logs were missing from the logging outputs, making troubleshooting difficult. + +## Resolution +A code change (PR 1950) was merged to restore logging outputs for the dataset import process. The fix was verified, and the validation report completed successfully. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/500945912.md b/tools/import_validation/troubleshooting_guides/bugs_summary/500945912.md new file mode 100644 index 0000000000..631e5c9d98 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/500945912.md @@ -0,0 +1,7 @@ +# Bug 500945912 Summary + +## Problem Description +A failure was reported in the download script for an import. Additionally, there were unexpected deletions of data series for certain Statistical Variables. It was identified that these deletions occurred because a specific StatVar wasn't properly recognized or was missing from the internal schema, leading the system to treat existing data as obsolete. + +## Resolution +The issue was resolved by adding the `-existing_statvar_mcf` parameter to the `manifest.json` file. This ensures that the system correctly identifies and preserves the existing Statistical Variables during the import process. The fix was verified as successful in subsequent runs. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/502079403.md b/tools/import_validation/troubleshooting_guides/bugs_summary/502079403.md new file mode 100644 index 0000000000..f26f4f7db5 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/502079403.md @@ -0,0 +1,7 @@ +# Bug 502079403 Summary + +## Problem Description +The auto-refresh cron job for the school district statistics import was missing from the scheduler, preventing automated updates. + +## Resolution +The refresh configuration was analyzed. Since newer data for 2024-25 was available at the source, steps were initiated to manually download the files and run a semi-automated refresh. The issue is assigned and is being resolved. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/502090898.md b/tools/import_validation/troubleshooting_guides/bugs_summary/502090898.md new file mode 100644 index 0000000000..2f1d7e8613 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/502090898.md @@ -0,0 +1,7 @@ +# Bug 502090898 Summary + +## Problem Description +A validation failure occurred due to API issues causing unintended deletions during an import. This was related to how historical data was being handled and flagged for deletion when not found in the current source run. + +## Resolution +The fix involved a multi-step process: creating a plan for historical data retention, updating the import configuration to store historical data properly, and finally updating the latest version configurations to remove the deletion flags. This ensured that valid historical data was preserved even if not present in the most recent source run. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/504879314.md b/tools/import_validation/troubleshooting_guides/bugs_summary/504879314.md new file mode 100644 index 0000000000..47e9c424ac --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/504879314.md @@ -0,0 +1,7 @@ +# Bug 504879314 Summary + +## Problem Description +The automated data refresh for the `EurostatData_Education_Attainment` dataset failed due to a validation error, which blocked the data import pipeline. + +## Resolution +A validation fix was implemented to resolve the failure. Following verification, the validation report confirmed that the imported data successfully passes all validation checks. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/506961224.md b/tools/import_validation/troubleshooting_guides/bugs_summary/506961224.md new file mode 100644 index 0000000000..b9928ffcd7 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/506961224.md @@ -0,0 +1,7 @@ +# Bug 506961224 Summary + +## Problem Description +The World Bank dataset refresh failed with a "length mismatch" error in `world_bank_stat_vars.mcf` caused by new indicators and significant row deletions. The task was blocked because the executor run lacked the necessary write permissions to update production files. + +## Resolution +A new deduplication logic was merged (PR 2077) to handle historical data. A workaround was identified to manually update the MCF file and copy historical data files to the production environment by a user with higher permissions. diff --git a/tools/import_validation/troubleshooting_guides/bugs_summary/507394518.md b/tools/import_validation/troubleshooting_guides/bugs_summary/507394518.md new file mode 100644 index 0000000000..6d185db8b6 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/bugs_summary/507394518.md @@ -0,0 +1,7 @@ +# Bug 507394518 Summary + +## Problem Description +The data import for `US_SAT_ACT_Participation` failed during automated refresh validation. + +## Resolution +A code fix was merged and deployed. Validation checks successfully ran and passed, verifying that the SAT/ACT data import is stable. diff --git a/tools/import_validation/troubleshooting_guides/docs_summary/bls_ces_state_deletion_resolution.md b/tools/import_validation/troubleshooting_guides/docs_summary/bls_ces_state_deletion_resolution.md new file mode 100644 index 0000000000..a55b46074d --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/docs_summary/bls_ces_state_deletion_resolution.md @@ -0,0 +1,9 @@ +# Google Doc: BLS_CES_State Deletion Resolution Summary + +* **Original URL:** https://docs.google.com/document/d/1QK9pMoSFR7STb78aFgN_NOupcJzU_IdWp4tyIvzrwis/edit +* **Purpose/Overview:** This document outlines the resolution plan for a validation failure in the `BLS_CES_State` data import on April 20, 2026. +* **Problem Description:** The import process failed because 4.57% (33,120 records) of the data points were deleted, which exceeded the default 0% threshold. +* **Root Causes:** The Bureau of Labor Statistics (BLS) announced official reductions and eliminations of several series data points starting with the release of January 2026 data as part of their annual benchmarking and sample review. +* **Resolutions/Findings:** + * Since these deletions are official source-side changes, the deleted records cannot be recovered. + * To resolve the import failure, the deleted data was preserved by extracting and storing it in a historical file, uploading it to Content Native Storage, and adjusting the validation settings to allow the pipeline to complete. diff --git a/tools/import_validation/troubleshooting_guides/docs_summary/bls_imports_issues_summary.md b/tools/import_validation/troubleshooting_guides/docs_summary/bls_imports_issues_summary.md new file mode 100644 index 0000000000..bb4e91721b --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/docs_summary/bls_imports_issues_summary.md @@ -0,0 +1,9 @@ +# Google Doc: BLS Imports Issues Summary + +* **Original URL:** https://docs.google.com/document/d/1bQbwKMVPtcehUZ3zI3r2kGlidMiE6UjEN9faKlJgtUs/edit +* **Purpose/Overview:** This document provides a general overview of multiple common issues and fixes across various BLS (Bureau of Labor Statistics) data imports (CES, CPI_Category, CES_State, etc.). +* **Problem Description:** Multiple BLS pipelines failed due to outdated files, missing references, and blocked automated downloads. +* **Root Causes & Fixes:** + * **Outdated latest_version.txt:** The Cloud Batch run date was newer than the version indicated in the `latest_version.txt` file, causing false deletion detections. Fixed by manually updating the `latest_version.txt` configuration. + * **Dependency on Production Schema Releases:** New Statistical Variables (SVs) added to the schema caused missing reference errors until the production release was finalized. Resolved by updating the pipeline to query the AutoPush DC schema instance. + * **Source Blocked Automated Downloads:** BLS upgraded their security posture using TLS Fingerprinting, which blocked automated python scripts with a `403 Forbidden` error. Resolved by using the `curl_cffi` library to mimic a standard browser's TLS handshake signature. diff --git a/tools/import_validation/troubleshooting_guides/docs_summary/dc_imports_execution_learnings.md b/tools/import_validation/troubleshooting_guides/docs_summary/dc_imports_execution_learnings.md new file mode 100644 index 0000000000..c8a7ad30e2 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/docs_summary/dc_imports_execution_learnings.md @@ -0,0 +1,9 @@ +# Google Sheet: DC - Imports Execution Learnings Summary + +* **Original URL:** https://docs.google.com/spreadsheets/d/1lAm-EPB6o9U9btQHbdRruGhx_DFerlhQF43Ba6skI4Q/edit +* **Purpose/Overview:** This spreadsheet tracks ongoing execution learnings, issues, and recommendations across multiple Data Commons import pipelines. +* **Key Entries & Actions:** + * **Unexpected Characters:** Source data issues resolved by implementing PV Mapping and processor code changes. + * **Missing Places/DCIDs:** Validation failures resolved by manually adding missing places to production (e.g. `b/472258775`). + * **Pipeline Automation:** Recommendation to convert Semi-Auto Refresh setups into Full Auto Refresh pipelines to minimize manual restarts (e.g. `b/474348464`). + * **Hangs and Performance:** Documented cases of copy service hangs requiring manual restarts, and low parallel performance in the KG Import UI. diff --git a/tools/import_validation/troubleshooting_guides/docs_summary/eia_nuclear_outages_summary.md b/tools/import_validation/troubleshooting_guides/docs_summary/eia_nuclear_outages_summary.md new file mode 100644 index 0000000000..3613f7b8eb --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/docs_summary/eia_nuclear_outages_summary.md @@ -0,0 +1,9 @@ +# Google Doc: EIA_NuclearOutages Summary + +* **Original URL:** https://docs.google.com/document/d/1GfF_sdUCg4d3fmkJoRCarQpFoGtLIPfezoKwoEXXL84/edit +* **Purpose/Overview:** This document details the analysis of an import failure in the `EIA_NuclearOutages` dataset on March 3, 2026. +* **Problem Description:** The import failed validation because it found 273 (0.01%) deleted records, which violated the 0% deletion threshold. +* **Root Causes:** The deletions represented pre-commercial testing phase records for Vogtle Unit 4 (from January to March 2024). Once the unit entered official commercial operation, the EIA cleaned up the dataset, removing these pre-commercial records from their historical data. +* **Resolutions/Findings:** + * It was decided not to preserve the deleted data as historical, since the deleted points only represented testing-phase records and the volume was negligible. + * The issue was resolved by updating the `latest_version.txt` file to acknowledge the deletion as a deliberate source update, which allowed the pipeline execution to proceed. diff --git a/tools/import_validation/troubleshooting_guides/docs_summary/eia_seds_summary.md b/tools/import_validation/troubleshooting_guides/docs_summary/eia_seds_summary.md new file mode 100644 index 0000000000..c71ff86489 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/docs_summary/eia_seds_summary.md @@ -0,0 +1,11 @@ +# Google Doc: EIA_SEDS Summary + +* **Original URL:** https://docs.google.com/document/d/1Kvhz-7AaWpaxSvSN1wHDP75-3CVXTfZT7dffBgZNTaY/edit +* **Purpose/Overview:** This document describes the validation failure of the `EIA_SEDS` data import on April 20, 2026. +* **Problem Description:** The import failed due to 1,248 (0.05%) deleted records and 56 missing reference warnings. +* **Root Causes:** + * **Deletions:** The source data removed zero-value placeholders that were not part of the NL Statistical Variables. + * **Missing References:** The source introduced new Statistical Variables that were not defined in the Data Commons schema. +* **Resolutions/Findings:** + * Acknowledged the deletions by updating `latest_version.txt`. + * Created and merged a Change List (CL) to add the definitions for the 56 new Statistical Variables in Data Commons, then ran a forced update. diff --git a/tools/import_validation/troubleshooting_guides/docs_summary/eurostat_employment_per_sector_summary.md b/tools/import_validation/troubleshooting_guides/docs_summary/eurostat_employment_per_sector_summary.md new file mode 100644 index 0000000000..883073dace --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/docs_summary/eurostat_employment_per_sector_summary.md @@ -0,0 +1,7 @@ +# Google Doc: EurostatData_Employment_Per_Sector Summary + +* **Original URL:** https://docs.google.com/document/d/1CPyqi9DBjv0t4eWanNYMRfJTIGjhubN6Op2Sao5S-9Y/edit +* **Purpose/Overview:** This document analyzes an import failure for the `EurostatData_Employment_Per_Sector` dataset on February 16, 2026. +* **Problem Description:** The process failed because it found 6,645 deleted records (1.39% of the total), which exceeded the allowed 0% threshold for deletions. +* **Root Causes:** The failure was caused by missing source data for the years 1995–1999 across 92 NUTS/IDs. The data was deleted from the official Eurostat source and is not expected to be restored. +* **Resolutions/Findings:** To resolve the validation error, the document recommends preserving the missing data by identifying it through a comparison with the previous production version, storing it in a historical file, and uploading it to Content Native Storage. diff --git a/tools/import_validation/troubleshooting_guides/docs_summary/eurostat_life_expectancy_summary.md b/tools/import_validation/troubleshooting_guides/docs_summary/eurostat_life_expectancy_summary.md new file mode 100644 index 0000000000..d022a55e09 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/docs_summary/eurostat_life_expectancy_summary.md @@ -0,0 +1,9 @@ +# Google Doc: EurostatData_LifeExpectancy Summary + +* **Original URL:** https://docs.google.com/document/d/1h4ETCMddDTWPCCZnE8OaHNYKFAH2WHtUX9tGCxV5caY/edit +* **Purpose/Overview:** This document covers the failure analysis for the `EurostatData_LifeExpectancy` import pipeline on March 19, 2026. +* **Problem Description:** The import pipeline failed due to 53,445 lint errors and 1.15% deleted records. +* **Root Causes:** Eurostat adjusted its data retention policy and stopped publishing historical life expectancy data before the year 2000, resulting in the deletion of all pre-2000 records from the source. +* **Resolutions/Findings:** + * To retain the historical data, the pipeline was updated to run with the deleted rows historical dataset from CNS (Content Native Storage). + * The validation threshold configurations were updated to allow the import to successfully complete with these changes. diff --git a/tools/import_validation/troubleshooting_guides/docs_summary/golden_set_validations_implementation_guide.md b/tools/import_validation/troubleshooting_guides/docs_summary/golden_set_validations_implementation_guide.md new file mode 100644 index 0000000000..eb08d9b1d6 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/docs_summary/golden_set_validations_implementation_guide.md @@ -0,0 +1,7 @@ +# Google Doc: Implementation Guide: Golden Set Validations Summary + +* **Original URL:** https://docs.google.com/document/d/14Fpe5e9jSzzJ5_QTcqQ1AFb2oqjJzXuc1_BmQCX-uns/edit +* **Purpose/Overview:** This is a technical guide designed to help users implement "Golden Set Validations" to protect data imports from regressions and accidental data loss. +* **Problem Description:** The guide addresses recurring data deletion failures by establishing baselines (golden files) that the system can use to verify new data against expected results. +* **Root Causes:** General lack of automated baseline comparisons for data integrity. +* **Resolutions/Findings:** The guide recommends implementing two primary validations: `Check_goldens_output_csv` (for final output data) and `Check_goldens_summary_report` (for structural metrics). It also defines a mandatory directory structure including a `golden_data/` folder and a `validation_config.json` file to manage tolerance thresholds. diff --git a/tools/import_validation/troubleshooting_guides/docs_summary/us_monthly_retail_sales_summary.md b/tools/import_validation/troubleshooting_guides/docs_summary/us_monthly_retail_sales_summary.md new file mode 100644 index 0000000000..4eaec48623 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/docs_summary/us_monthly_retail_sales_summary.md @@ -0,0 +1,9 @@ +# Google Doc: USMonthlyRetailsales Summary + +* **Original URL:** https://docs.google.com/document/d/1mVoeHRkdlnPvTpu-IMee-6d2GQ-htlLDRP2ZYdGW04w/edit +* **Purpose/Overview:** This document outlines a validation failure for the `USMonthlyRetailsales` dataset on April 14, 2026. +* **Problem Description:** The import failed during validation due to 4 (0.01%) deleted records. +* **Root Causes:** The StatVar processor incorrectly treated industry category codes (NAICS) in the source file as financial values and attempted to multiply them by 1,000,000 (standard conversion to USD), which triggered validation inconsistencies. +* **Resolutions/Findings:** + * The StatVar processor logic was updated to ignore NAICS industry codes and avoid converting non-financial records. + * The pipeline was then re-run to confirm clean validation. diff --git a/tools/import_validation/troubleshooting_guides/docs_summary/uscensuspep_agesexrace_summary.md b/tools/import_validation/troubleshooting_guides/docs_summary/uscensuspep_agesexrace_summary.md new file mode 100644 index 0000000000..7e39e59640 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/docs_summary/uscensuspep_agesexrace_summary.md @@ -0,0 +1,7 @@ +# Google Doc: USCensusPEP_AgeSexRace Summary + +* **Original URL:** https://docs.google.com/document/d/10_kFWkpkon9xhVOlOXjIFF0fSyO2IGFKWq6H1pcAcVE/edit +* **Purpose/Overview:** This document provides a detailed analysis of a validation failure that occurred during the `USCensusPEP_AgeSexRace` data import process on April 14, 2026. +* **Problem Description:** The import failed during the validation stage because a data consistency check identified 608 deleted records (representing a negligible percentage of the total data). +* **Root Causes:** The primary cause was an inaccessible source URL from `census.gov`, which prevented the system from retrieving specific data points. However, it was confirmed that these affected statistical variables (SVs) were not part of the NL Statvars. +* **Resolutions/Findings:** The recommended resolution is to update the `latest_version.txt` file to acknowledge the deletions and prevent them from being flagged as errors, followed by a forced re-run of the pipeline. diff --git a/tools/import_validation/troubleshooting_guides/docs_summary/wdi_auto_failure_validation_template.md b/tools/import_validation/troubleshooting_guides/docs_summary/wdi_auto_failure_validation_template.md new file mode 100644 index 0000000000..8f183cecb4 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/docs_summary/wdi_auto_failure_validation_template.md @@ -0,0 +1,12 @@ +# Google Doc: World Development Indicators Auto-Failure Validation Summary + +* **Original URL:** https://docs.google.com/document/d/1PyBmcN-1C_p9y-ML93eaBFyspg1XD5zwkT7EqeTsQsI/edit +* **Purpose/Overview:** Used as a reference template for validation failure analysis, documenting a failure in the WDI pipeline on December 23, 2025. +* **Problem Description:** The pipeline failed due to 1,011 validation errors (deletions) and 908 lint errors. +* **Root Causes:** + * **Syria Deletions:** Data for Syria was missing from the source folder due to a change in the source layout. + * **Kosovo Code Mismatch:** The source data used the 3-letter code `XKX` for Kosovo, which failed validation because Data Commons uses the `XKS` code. + * **Channel Islands Mismatch:** The source data used `CHI` which did not match the Data Commons DCID `ChannelIslands`. +* **Resolutions/Findings:** + * Updated the `latest_version.txt` configuration file to point to the latest source folder structure. + * Updated the `worldbank.py` script to map country code `XKX` to `XKS` and `country/CHI` to `ChannelIslands` to resolve lint validation failures. diff --git a/tools/import_validation/troubleshooting_guides/docs_summary/world_bank_datasets_summary.md b/tools/import_validation/troubleshooting_guides/docs_summary/world_bank_datasets_summary.md new file mode 100644 index 0000000000..63e8df6246 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/docs_summary/world_bank_datasets_summary.md @@ -0,0 +1,11 @@ +# Google Doc: WorldBankDatasets Summary + +* **Original URL:** https://docs.google.com/document/d/1ExMv5_J4vSEY1OVwWLoXbvnH1ltjRsfcBjUqSyw6uXo/edit +* **Purpose/Overview:** This document outlines failures in the `WorldBankDatasets` auto-refresh pipeline on August 25, 2025. +* **Problem Description:** The import pipeline failed due to approximately 1.3 million lint errors and 18,398 deleted data points, exceeding the threshold. +* **Root Causes:** + * **Deletions:** The source data did not contain the observations. + * **Lint Errors:** Aggregated regional groupings and place names in the source data used invalid DCIDs (e.g. Channel Islands as `CHI` and Kosovo as `XKX`) that did not exist in Data Commons. +* **Resolutions/Findings:** + * Resolved the lint errors by mapping invalid place DCIDs using `places.csv` and `skip_places.csv` filters. + * Preserved the deleted data points by writing them to a historical `deleted_rows.csv` file, uploading it to Content Native Storage, and configuring it to skip errors during future pipeline runs. diff --git a/tools/import_validation/troubleshooting_guides/gcp_variables.md b/tools/import_validation/troubleshooting_guides/gcp_variables.md new file mode 100644 index 0000000000..b3455643ca --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/gcp_variables.md @@ -0,0 +1,25 @@ +# GCP Project and Bucket Glossary + +This file defines the Google Cloud Platform (GCP) projects and Cloud Storage (GCS) buckets used across the Data Commons import pipelines, along with descriptions of their purposes. + +## Variables + +### `PROD_PROJECT` +* **Value:** `datcom-204919` +* **Description:** The production Google Cloud project where raw data imports, version logs, and validation outputs are published. + +### `AUTO_REFRESH_PROJECT` +* **Value:** `datcom-import-automation-prod` +* **Description:** The production Google Cloud project dedicated to hosting the automated data refresh cron schedules, Cloud Batch jobs, and pipeline execution logs. + +### `BQ_PROJECT` +* **Value:** `datcom-store` +* **Description:** The project containing Data Commons BigQuery tables (e.g., Knowledge Graph datasets like `datcom-store.dc_kg_latest.NLStatVars`). + +### `PROD_BUCKET` +* **Value:** `datcom-prod-imports` +* **Description:** The primary production GCS storage bucket containing import files, timestamped runs, differ outputs (`obs_diff_log.csv`), and validation results. + +### `BASE_PROJECT` +* **Value:** `datcom` +* **Description:** The base container project referencing general storage and service operations. diff --git a/tools/import_validation/troubleshooting_guides/golden_checks.md b/tools/import_validation/troubleshooting_guides/golden_checks.md new file mode 100644 index 0000000000..c7ad766676 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/golden_checks.md @@ -0,0 +1,209 @@ +# **Implementation Guide: Golden Set Validations** + +## **1\. Overview** + +Due to recurring data deletion failures, we protect our imports using **Golden Set Validations (`GOLDENS_CHECK`)**. This guide helps understand how to generate "golden files" (baselines of expected data) and configure the automated validation system to prevent accidental data regressions for each of the script output csv files listed in the manifest. + +Currently, to mitigate issues related to data loss, we implement two primary validations supported by specific golden baselines: + +1. **`Check_goldens_output_csv`**: Verifies the final output data against an established baseline. +2. **`Check_goldens_summary_report`**: Validates structural metrics (like number of places and dates) against a summary baseline. +3. More regarding golden checks can be seen : [github link](https://github.com/datacommonsorg/data/blob/master/tools/import_validation/Validations.md) + +## **2\. Directory Architecture** + +To implement these checks, your import script folder must contain the following file structure: + +Plaintext + +``` +your_import_folder/ +│ +├── manifest.json # Main import configuration +├── validation_config.json # Custom validation rules +│ +└── golden_data/ # Folder holding your baseline files + ├── golden_summary_report.csv # Generated golden summary + └── golden_observations.csv # Generated golden output data +``` + +## **3\. Step-by-Step: How to Create Golden Files** + +Golden files are created by extracting a snapshot of known-good data using the script `validator_goldens.py`. Run these commands from your terminal inside your import directory. + +### **Step 3.1: Generate the Summary Report Golden File** + +This step tracks metadata properties like Statistical Variables (StatVars), the number of places, and date ranges to ensure future runs don't accidentally drop the entire series. + +Specifically over here, we need to take into consideration the columns that don't change over time in every execution. For eg. + +1. observationPeriods +2. units +3. scalingFactors +4. measurementMethods +5. NumPlaces +6. MinDate + +From the data/tools/import\_validation/validator\_goldens.py directory, execute the following command: + +``` +python3 validator_goldens.py --validate_goldens_input=summary_report.csv --generate_goldens=golden_data/golden_summary_report.csv --generate_goldens_property_sets="StatVar|NumPlaces|MinDate|MeasurementMethods|Units|ScalingFactors|observationPeriods" +``` +Note: A separate rule can also be created in case the values expected for particular stavars are range bounded. For eg. a stavar having the unit "Percent" will mostly be between 0-100. + +### **Step 3.2: Create separate golden outputs for prominent places in observationAbout and statvars (only use if needed)** + +This step targets critical combinations of highly utilized StatVars and top geographical regions to ensure key data points are always preserved. + +From the data/tools/import\_validation/validator\_goldens.py directory, execute the following command: + +``` +python3 validator_goldens.py --validate_goldens_input=output.csv --generate_goldens=golden_data/golden_observations.csv --goldens_must_include="observationAbout:gs://unresolved_mcf/import_validation/top_100k_places.csv" --generate_goldens_property_sets="observationAbout" +``` + +## **4\. Configuring `validation_config.json`** + +Create a file named `validation_config.json` in your import script directory. Paste the configuration below. + +This configuration does two things: + +* Overrides the default deletion tolerance rule (`check_deleted_records_percent`) to a threshold as per history deletions & current deletions should not be more than **0.1%**. +* Activates the golden check rules pointing to the files you created in Section 3. + +JSON + +``` +{ + "schema_version": "1.0", + "rules": [ + { + "rule_id": "check_deleted_records_percent", + "description": "Strictly enforce historical deletion average threshold of 0.1%", + "validator": "DELETED_RECORDS_PERCENT", + "params": { + "threshold": 0.1 + } + }, + { + "rule_id": "check_goldens_summary_report", + "description": "Validates summary_report.csv against the golden summary data", + "validator": "GOLDENS_CHECK", + "params": { + "golden_files": ["golden_data/golden_summary_report.csv"] + } + }, + { + "rule_id": "check_goldens_output_csv", + "description": "Verifies the generated output CSV data matches established critical golden records", + "validator": "GOLDENS_CHECK", + "params": { + "golden_files": ["golden_data/golden_observations.csv"], + "input_files": ["output/observations.csv"] + } + } + ] +} +``` + +## **Below is how the validation\_config.json is structured for imports with multiple output CSVs:** + +JSON + +``` +{ + "schema_version": "1.0", + "rules": [ + { + "rule_id": "check_deleted_records_percent", + "description": "Checks that the percentage of deleted records for the entire import is within threshold.", + "validator": "DELETED_RECORDS_PERCENT", + "params": { "threshold": 0.1 } + }, + { + "rule_id": "check_goldens_national", + "description": "Validates national and state-level 2000+ data against its golden summary report.", + "validator": "GOLDENS_CHECK", + "params": { + "golden_files": ["../../../../golden_data/golden_summary_report_national.csv"], + "input_files": ["../../input0/genmcf/summary_report.csv"] + } + }, + { + "rule_id": "check_goldens_before_2000", + "description": "Validates data before 2000 against its golden summary report.", + "validator": "GOLDENS_CHECK", + "params": { + "golden_files": ["../../../../golden_data/golden_summary_report_before_2000.csv"], + "input_files": ["../../input1/genmcf/summary_report.csv"] + } + }, + { + "rule_id": "check_goldens_after_2000", + "description": "Validates county-level 2000+ data against its golden summary report.", + "validator": "GOLDENS_CHECK", + "params": { + "golden_files": ["../../../../golden_data/golden_summary_report_after_2000.csv"], + "input_files": ["../../input2/genmcf/summary_report.csv"] + } + }, + { + "rule_id": "Check_goldens_output_csv_before_2000", + "description": "Verifies the generated output CSV data matches established critical golden records", + "validator": "GOLDENS_CHECK", + "params": { + "golden_files": ["../../../../golden_data/golden_observations_before_2000.csv"], + "input_files": ["../../../../output/USA_Population_Count_by_Race_before_2000.csv"] + } + }, + { + "rule_id": "Check_goldens_output_csv_after_2000", + "description": "Verifies the generated output CSV data matches established critical golden records", + "validator": "GOLDENS_CHECK", + "params": { + "golden_files": ["../../../../golden_data/golden_observations_after_2000.csv"], + "input_files": ["../../../../output/USA_Population_Count_by_Race_county_after_2000.csv"] + } + }, + { + "rule_id": "Check_goldens_output_csv_state", + "description": "Verifies the generated output CSV data matches established critical golden records", + "validator": "GOLDENS_CHECK", + "params": { + "golden_files": ["../../../../golden_data/golden_observations_national.csv"], + "input_files": ["../../../../output/USA_Population_Count_by_Race_National_state_2000.csv"] + } + } + ] +} +``` + +## **5\. Activating the Config in `manifest.json`** + +The auto-refresh pipeline will only notice your new rules if you explicitly link them in your main `manifest.json` file. + +Add the `validation_config_file` parameter pointing to your file inside the `config_overrides` object: + +JSON + +``` +{ + "import_spec": { + "name": "Your_Import_Name_Here" + }, + "config_overrides": { + "validation_config_file": "validation_config.json" + } +} +``` + +*(Note: Remember that StatVar updates in this environment are driven systematically via the manifest configuration flags rather than manual file renaming.)* + +## **6\. How to Read Validation Failures** + +If a future data refresh breaks these rules, the pipeline will fail, and a report JSON will be generated. + +* **If `Check_goldens_summary_report` fails:** It means a StatVar or a specific geographic series has unexpectedly disappeared from the pipeline. +* **If `Check_goldens_output_csv` fails:** The specific rows of data present in your baseline file but missing in the new run will be explicitly listed in the output log. + +> **When to update golden files:** Only regenerate golden files using the scripts in if a data change is intentional (e.g., source data deprecation, structural schema updates). Always have the changes reviewed by a peer before committing new golden baselines. + diff --git a/tools/import_validation/troubleshooting_guides/major_deletions.md b/tools/import_validation/troubleshooting_guides/major_deletions.md new file mode 100644 index 0000000000..daa34dddd7 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/major_deletions.md @@ -0,0 +1,141 @@ +# **Validating Deletions for the Import** + + The validation process involves : examining the cloud job logs, replicating the import run on a local machine, and identifying the cause of record deletions to implement a permanent fix. + +## **Check the Cloud Logs to See What Failed** + +First, we need to find out why the job failed by looking at the cloud logs and the error messages. + +1. **Check the Cloud Batch Job:** + Go to the Cloud Batch Jobs console and find the latest run for the import. + * Here is the job I looked at for this import: Cloud Batch Job: `worlddevelopmentindicators-1781002802` (Project: `[](gcp_variables.md#auto_refresh_project)`, Region: `us-central1`) +2. **Look for Errors in the Logs:** + Even if the job status says "Succeeded," it might still have errors. Look through the logs (you can use the errors filter to jump straight to them). In my case, the job failed some validation checks because of deletions. + * **Error Message:** Found 0.87% deleted records, which is over the threshold of 0%. +3. **Check the Production Bucket:** + Next, head over to the production cloud bucket to find the specific files for this job run. Navigate down the correct path for the import: + Project: `[](gcp_variables.md#prod_project)`, Bucket: `[](gcp_variables.md#prod_bucket)`, Path: `scripts/world_bank/wdi/WorldDevelopmentIndicators` + * Go to the folder with the latest timestamp as the job you just looked at: 2026\_06\_09T04\_07\_31\_200635\_07\_00 +4. **Review the Validation Files:** + Go into the input0/validation/ folder. Here, you want to look at two specific files: + * **validation\_output.csv:** This tells you the exact reason the checks failed. In my case, 3 checks passed, but the check\_deleted\_records\_percent failed because 3,908 records were deleted. (See table below) + * Location: Project: `[](gcp_variables.md#prod_project)`, Bucket: `[](gcp_variables.md#prod_bucket)`, Path: `scripts/world_bank/wdi/WorldDevelopmentIndicators/2026_06_09T04_07_31_200635_07_00/input0/validation/validation_output.csv` + * **Nodes\_deleted.mcf:** This file shows you the actual list of records that were deleted. + * Location: Project: `[](gcp_variables.md#prod_project)`, Bucket: `[](gcp_variables.md#prod_bucket)`, Path: `scripts/world_bank/wdi/WorldDevelopmentIndicators/2026_06_09T04_07_31_200635_07_00/input0/validation/obs_diff_log.csv` + +## **Run the Import Locally** + +Now, we need to run the import process on our own computer. This helps us confirm if the deletions are really happening, or if it was just a transient failure in the cloud job. + +### **Automated Run using `run_import.sh` (Recommended)** +Instead of executing every step manually (running parser scripts, executing the Java import-tool to build/lint MCFs, downloading previous baselines, and running the differ tool), you can use the unified **`run_import.sh`** utility script. It automates the entire local execution and validation flow. + + #### **Theory of Operation:** + * **Parses the Manifest:** It reads `manifest.json` for the import's resource constraints and import specifications. + * **Prepares the Environment:** It downloads the required `import-tool.jar` and configures execution defaults. + * **Runs User Scripts:** It executes the downloader and generator scripts defined in the manifest. + * **Validates & Runs Diff Checks:** It automatically generates output MCFs, runs lint validations, pulls the previous successful version from GCS to compare, and generates the final validation report (`validation_output.csv` and `obs_diff_log.csv`). + + #### **Execution Commands:** + Navigate to the root directory of your repository and run one of the following commands: + + * **Using Docker (Highly Recommended):** Runs the pipeline in a clean container, avoiding local dependency issues. + ```bash + ./import-automation/executor/run_import.sh -docker scripts/world_bank/wdi/manifest.json + ``` + *If you want to build and use a Docker image with your latest local code changes:* + ```bash + ./import-automation/executor/run_import.sh -d dc-test-executor -docker scripts/world_bank/wdi/manifest.json + ``` + + * **Using your Local Python Host Environment:** + ```bash + ./import-automation/executor/run_import.sh scripts/world_bank/wdi/manifest.json + ``` + The execution outputs, logs, and final validation reports will be stored in `/tmp/WorldDevelopmentIndicators/` (or the folder path specified by `-o `). + +--- + +### **Manual Run (Alternative)** +If you prefer to run the individual steps of the pipeline manually: + +1. **Clone the GitHub Repository:** + Clone the Datacommons data repository to your machine: [https://github.com/datacommonsorg/data.git](https://github.com/datacommonsorg/data.git). Once cloned, navigate to the exact same import path we looked at in the bucket. +2. **Run the Scripts:** + Check the [manifest.json](https://github.com/datacommonsorg/data/blob/master/scripts/world_bank/wdi/manifest.json) file in that folder to understand how the job runs, and manually run the scripts and processes it outlines. + +3. **Generate the MCF File:** + Run the lint and genmcf tests to generate the table.mcf file, which we will need for the differ tool. Run these commands: + * java \-jar java-jar.jar lint output.csv output.tmcf + * java \-jar java-jar.jar genmcf output.csv output.tmcf +4. **Get the Previous Data:** + To run the differ tool, we need to compare the table.mcf file we just created (the current data) with the table.mcf file from the previous successful run. + * Find the timestamp of the last successful run by checking the `latest_version.txt` file in the cloud bucket. GCS Location: Project: `[](gcp_variables.md#prod_project)`, Bucket: `[](gcp_variables.md#prod_bucket)`, Path: `scripts/world_bank/wdi/WorldDevelopmentIndicators/latest_version.txt` + * Go to that timestamp's folder in the bucket and download its table.mcf file. +5. **Run the Differ Tool:** + Now, run this command to compare the two files: + python3 import\_differ.py \--current\_data= \--previous\_data= \--output\_location= \--file\_format=mcf \--runner\_mode=local +6. **Confirm the Results:** + Open the dc\_generated/nodes\_deleted.mcf folder and the validation\_output.csv on your local machine. If they match the results we saw in the cloud bucket during Phase 1, we know the results are accurate\! + +## **Validate Deletions from the Source** + +Now that we know the deletions are real, we need to validate them against the original data source to confirm the source actually removed the data. + +1. **Locate the Textproto File:** To find out exactly where the data came from, check the `textproto` file for this import. For WorldDevelopmentIndicators, the file path is: `google3/datacommons/import/mcf/manifest/international_stats/WorldDevelopmentIndicators.textproto` + * You can search for this file using Google Code Search: [WorldDevelopmentIndicators.textproto Link](screens/wdi_textproto_source.png +2. **Find the Source URL:** Open the `textproto` file and look for the line that says: `provenance_url: "https://datatopics.worldbank.org/world-development-indicators/"` This URL tells us exactly where the data is downloaded from. +3. **Navigate the Source Website:** Open that source URL. On the World Bank website, click on the **Explore Data** option, and then click on **Access Data**. This will allow you to see their entire dataset. +4. **Manually Verify Deleted Records:** check at your `nodes_deleted.mcf` file and pick out 5 to 6 specific deleted records. Go back to the World Bank website and manually filter the data by entering the parameters for those specific records. + * *If you do not get any data for them (or the data is missing), this confirms that the data was accurately deleted from the source.* + + *Example :* The job failed because some specific data points present in the previous production version are missing in the current version. + +* **Source Deletion:** The record has been removed from the source. +1) [Source Screenshot](screens/3oK73RXvox3xeTM.png) | [Differ Screenshot](screens/BWYDfWvTNZoixSE.png) +2) [Source Screenshot](screens/8znJ3XadmGGRmdL.png) | [Differ Screenshot](screens/4JKbhAxnJz6SGbJ.png) +3) [Source Screenshot](screens/5Nw3dA9a3iH7GAM.png) | [Differ Screenshot](screens/7ovdMycUe7tiADw.png) +* These deletions are confirmed as intentional source-side changes from the World Bank’s April 2026 update. + + WDI April 8, 2026 Changelog [Screenshot](screens/497zku465XJzsJt.png) | [Link](https://datatopics.worldbank.org/world-development-indicators/release-note/apr-2026.html) + +5. In my case, the amount of data getting deleted from the source was very huge. When this happens, we cannot just ignore it. We need to: + * **Create a validation error document:** Document the massive deletion properly so there is a clear record of why the job threshold failed and what was removed. [BLS\_CES\_State\_20\_04\_2026](docs_summary/bls_ces_state_deletion_resolution.md + * **Store the data historically:** Keep a record of the deleted data. (need approval from core team) +6. Check for the affected SVs deleted that we find out by taking unique deleted SVs from the differ & check in BigQuery (Table: `[](gcp_variables.md#bq_project).dc_kg_latest.NLStatVars`) if these Svs are present in the NL SVs table. +7. Because the deletions are minor & from the source the next would be store Historical data & because it had recurring failure golden checks \+ threshold increase as per history deletions will also be implemented ( All these steps must be mentioned in the validation error document because we need core team approval to store historical & update the latest\_version.txt) + + ## **How to implement golden checks?** + + Consult the following resource for instructions on incorporating goldens into your import process: [Implementation Guide: Golden Set Validations](docs_summary/golden_set_validations_implementation_guide.md + + ## **How to store historical data ?** + +1. resolve **validation errors** + Preserve the deleted data by storing it in a historical file, which should then be copied to the CNS. Below are the steps: + 1. Verify the production version of the import via [Data Commons/Version](https://datacommons.org/version). + 2. Locate and download the production CSV from the storage bucket, matching the date specified in the Data Commons version. + 3. Run the script in your local environment, then utilize Python code to perform a difference comparison between the latest and production output CSVs. + 4. Run the differ on output & historical data table\_mcf\_nodes.mcf files together & ensure no deletion flags are there. + 5. Once the deleted rows are identified & no deletions with this historical file through the comparison, save them to a file and upload it to CNS as a historical record, path shown below. + +``` +mcf_proto_url: "/cns/jv-d/home/[](gcp_variables.md#base_project)/v3_resolved_mcf/us_bls/ces/state/latest/historical_data/graph.tfrecord@1.gz" + table { + mapping_path: "/cns/jv-d/home/[](gcp_variables.md#base_project)/v3_mcf/wdi/WorldDevelopmentIndicators/historical_data/worldbank.tmcf" + csv_path: "/cns/jv-d/home/[](gcp_variables.md#base_project)/v3_mcf/wdi/WorldDevelopmentIndicators/historical_data/*.csv" + } +``` + + b. To stop these “Deleted” flags from appearing as errors, **the latest\_version.txt file must be updated only after the core team approves.** This ensures the differ recognizes the change as deliberate to the dataset rather than a data loss error. + +``` +experimental/users/ajaits/[](gcp_variables.md#base_project)/scripts/import_info.sh -i WorldDevelopmentIndicator-set_latest 2026_04_20_02_42_50_536488_08_00 -note 'details of deletion analysis in b/500945912 reviewed by: ' +``` + + C. Rerun the pipeline to verify it finishes without issues and check that all tests in *validation\_output.csv* have passed. + + + + + diff --git a/tools/import_validation/troubleshooting_guides/minor_deletions.md b/tools/import_validation/troubleshooting_guides/minor_deletions.md new file mode 100644 index 0000000000..622403cd29 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/minor_deletions.md @@ -0,0 +1,85 @@ +# **Validating Deletions for the EurostatData\_Fertility Import** + +To address occasional, low-impact source data deletions in the import **EurostatData\_Fertility**, you must validate that these minor data drops are expected and benign. The resolution process involves: examining the cloud job logs to confirm the deletion percentage is small, verifying the scope of the missing records, and adjusting the validation thresholds to allow the job to succeed. + +## **Check the Cloud Logs to See What Failed** + +First, we need to find out why the job failed by looking at the cloud logs and the error messages. + +1. **Check the Cloud Batch Job:** + Go to the Cloud Batch Jobs console and find the latest run for the import. + * Here is the job I looked at for this import: Cloud Batch Job: `eurostatdata-fertility-1781582401` (Project: `[](gcp_variables.md#auto_refresh_project)`, Region: `us-central1`) +2. **Look for Errors in the Logs:** + Even if the job status says "Succeeded," it might still have errors. Look through the logs (you can use the errors filter to jump straight to them). In my case, the job failed some validation checks because of deletions. + * **Error Message:** Found 0.06% deleted records, which is over the threshold of 0%. +3. **Check the Production Bucket:** + Next, head over to the Datacom production cloud bucket to find the specific files for this job run. Navigate down the correct path for the import: + Project: `[](gcp_variables.md#prod_project)`, Bucket: `[](gcp_variables.md#prod_bucket)`, Path: `scripts/eurostat/regional_statistics_by_nuts/fertility_rate_mother_age/EurostatData_Fertility/` + * Go to the folder with the latest timestamp as the job you just looked at: 2026\_06\_15T04\_07\_31\_200635\_07\_00 +4. **Review the Validation Files:** + Go into the input0/validation/ folder. Here, you want to look at two specific files: + * **validation\_output.csv:** This tells you the exact reason the checks failed. In my case, 3 checks passed, but the check\_deleted\_records\_percent failed because 36 records were deleted. (See table below) + * Location: Project: `[](gcp_variables.md#prod_project)`, Bucket: `[](gcp_variables.md#prod_bucket)`, Path: `scripts/eurostat/regional_statistics_by_nuts/fertility_rate_mother_age/EurostatData_Fertility/2026_06_15T21_03_10_391174_07_00/input0/validation/validation_output.csv` + * **obs\_diff\_log.csv:** This file shows you the actual list of records that were deleted. + * Location: Project: `[](gcp_variables.md#prod_project)`, Bucket: `[](gcp_variables.md#prod_bucket)`, Path: `scripts/eurostat/regional_statistics_by_nuts/fertility_rate_mother_age/EurostatData_Fertility/2026_06_15T21_03_10_391174_07_00/input0/validation/obs_diff_log.csv` + +## **Run the Import Locally** + +Now, we need to run the whole import process on our own computer. This helps us confirm if the deletions are really happening, or if it was just a weird glitch in the cloud job. + +1. **Clone the GitHub Repository:** + Clone the Datacommons data repository to your machine: [https://github.com/datacommonsorg/data.git](https://github.com/datacommonsorg/data.git). Once cloned, navigate to the exact same import path we looked at in the bucket. +2. **Run the Scripts:** + Check the [manifest.json](https://github.com/datacommonsorg/data/blob/master/scripts/eurostat/regional_statistics_by_nuts/fertility_rate_mother_age/manifest.json) file in that folder to understand how the job runs, and manually run the scripts and processes it outlines. + +**Note**: The manifest.json file defines the number of output and TMCF files generated by the import; the presence of directories like input0 or input1 in the bucket is a direct result of these manifest specifications. + + + +3. **Generate the MCF File:** + Run the lint and genmcf tests to generate the table.mcf file, which we will need for the differ tool. Run these commands: + * java \-jar java-jar.jar lint output.csv output.tmcf + * java \-jar java-jar.jar genmcf output.csv output.tmcf +4. **Get the Previous Data:** + To run the differ tool, we need to compare the table.mcf file we just created (the current data) with the table.mcf file from the previous successful run. + * Find the timestamp of the last successful run by checking the `latest_version.txt` file in the cloud bucket. GCS Location: Project: `[](gcp_variables.md#prod_project)`, Bucket: `[](gcp_variables.md#prod_bucket)`, Path: `scripts/eurostat/regional_statistics_by_nuts/fertility_rate_mother_age/EurostatData_Fertility/latest_version.txt` + * Go to that timestamp's folder in the bucket and download its table.mcf file. +5. **Run the Differ Tool:** + Now, run this command to compare the two files: + python3 import\_differ.py \--current\_data= \--previous\_data= \--output\_location= \--file\_format=mcf \--runner\_mode=local +6. **Confirm the Results:** + Open the dc\_generated/obs\_diff\_logs.csv folder and the validation\_output.csv on your local machine. If they match the results we saw in the cloud bucket during Phase 1, we know the results are accurate\! + +## **Validate Deletions from the Source** + +Now that we know the deletions are real, we need to validate them against the original data source to confirm the source actually removed the data. + +1. **Locate the Textproto File:** To find out exactly where the data came from, check the `textproto` file for this import. For EurostatData_Fertility, the file path is: `datacommons/import/mcf/manifest/international_stats/EurostatData_Fertility.textproto` + * You can search for this file using Google Code Search: [EurostatData\_Fertility.textproto Link](screens/eurostat_fertility_textproto_source.png +2. **Find the Source URL:** Open the `textproto` file and look for the line that says: `provenance_url: "https://ec.europa.eu/eurostat/databrowser/view/demo_r_find3/default/table?lang=en"` This URL tells us exactly where the data is downloaded from. +3. **Navigate the Source Website:** Open that source URL. On the Eurostat website, click on the **Explore Data** option, and then click on **Access Data**. This will allow you to see their entire dataset. +4. **Manually Verify Deleted Records:** check at your `obs_diff_logs.csv` file and pick out 5 to 6 specific deleted records. Go back to the EurostatData website and manually filter the data by entering the parameters for those specific records. + * *If you do not get any data for them (or the data is missing), this confirms that the data was accurately deleted from the source.* + + *Example :* The job failed because some specific data points present in the previous production version are missing in the current version. + +* **Source Deletion:** The record has been removed from the source. + + [Source Screenshot](screens/3VNMybHoDVt6kFR.png) | [Differ Screenshot](screens/B25DugjwV2js2TJ.png) + + +* These deletions are confirmed as intentional source-side changes from the Eurostat Official Website. + +5. In my case, the amount of data getting deleted from the source was very less. When this happens, we just ignore it. We need to: + * **Create a validation error document:** Document the deletion properly so there is a clear record of why the job threshold failed and what was removed. + * Add goldens checks & increase minor threshold +6. Check for the affected SVs deleted that we find out by taking unique deleted SVs from the differ & check in [bigquery](screens/4pprS7vyu4Cba6C.png) if these SVs are present in the NL SVs table. + + + ## **How to implement golden checks?** + + Consult the following resource for instructions on incorporating goldens into your import process: [Implementation Guide: Golden Set Validations](docs_summary/golden_set_validations_implementation_guide.md + + + + \ No newline at end of file diff --git a/tools/import_validation/troubleshooting_guides/production_challenges.md b/tools/import_validation/troubleshooting_guides/production_challenges.md new file mode 100644 index 0000000000..2f66699663 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/production_challenges.md @@ -0,0 +1,97 @@ +## **Production Challenges: Source-Side Complications** + +### 1\. Structural and Schema Modifications + +Pipeline disruptions frequently arise from unexpected modifications to source data structures. These breaking changes—including altered formats, deleted records, or revised schema definitions—interfere with established ingestion and processing workflows. + +Information regarding the most frequent types of import failures can be found [DC - Imports Execution Learnings ](docs_summary/dc_imports_execution_learnings.md + +*Historical instances requiring code adjustments:* + +1. WorldBankDatasets: [Reference Documentation](docs_summary/world_bank_datasets_summary.md +2. EurostatData\_lifeexpectency: [Reference Documentation](docs_summary/eurostat_life_expectancy_summary.md +3. UsMontlyRetailSales: [Reference Documentation](docs_summary/us_monthly_retail_sales_summary.md + +**Standard Remediation Protocol** + +1. Verify the failure by searching for the specific import name within *validation\_output.csv* located in the `[](gcp_variables.md#prod_bucket)` bucket. +2. In cases involving data loss, investigate whether the records were intentionally removed at the source or if a mismatch has occurred, then establish the necessary correction. + +### 1\. Data structure changes (Schema changes) + +Modifications to source data are causing pipeline failures. When a data source alters its format, removes records, or changes schema definitions without notice, it introduces breaking changes that disrupt our ingestion and processing workflows. +For example, code fixes were required after changes occurred in the following datasets: + +1. WorldBankDatasets: [WorldBankDatasets ](docs_summary/world_bank_datasets_summary.md +2. EurostatData\_lifeexpectency: [EurostatData\_LifeExpectancy](docs_summary/eurostat_life_expectancy_summary.md +3. UsMontlyRetailSales : [ USMontlyRetailsales](docs_summary/us_monthly_retail_sales_summary.md + + + +**Standard procedures for addressing the problem** + +1. Analyze *validation\_output.csv* in the `[](gcp_variables.md#prod_bucket)` bucket to confirm the import failure using the import name. +2. If the error is due to deletions, verify if data was removed at the source or if there is a data mismatch. Determine the appropriate fix. +3. If a 'missing reference' error occurs, identify the missing entity or mapping (e.g., place DCID or statistical variable). Update the relevant MCF file in Cider and submit a CL. + + + +### 2\. Data deletion at source + +The data has been officially removed from the source system. In this case, the standard deletion handling rules apply: + +* For minor data deletions, implement golden checks to protect critical information and do a subsequent increase in the import threshold. + * In instances of large data removal, ensure that all deleted records are preserved as historical data. + +1. In case of deletions, we check for generated o/p CSV \-\> in o/p folder of an import/timestamp folder + + For example, Historical has been stored in [EurostatData\_Employment\_Per\_Sector](docs_summary/eurostat_employment_per_sector_summary.md & latest\_version.txt has been updated in [EIA\_NuclearOutages](docs_summary/eia_nuclear_outages_summary.md + +**Standard procedures for addressing the problem** + +1. Review the latest run of the job in the `[](gcp_variables.md#prod_bucket)` bucket under the `[](gcp_variables.md#base_project)` project using the import name. +2. If the job folder timestamp is older than one week, re-trigger the job to generate the latest output. +3. Examine `input0/validation/nodes_deleted.mcf` to review deletion details. +4. Validate whether the deletions originated from the source or were caused by a pipeline/code issue. +5. Apply the appropriate resolution based on the identified deletion scenario. + +### 3\. Downtime or modifications to source URLs + +These failures are typically caused by one of the following reasons: + +* The source URL is completely broken or no longer active. +* The external source changed the URL structure or moved the data to a new location. +* The external source's server is temporarily unresponsive. +* A firewall is blocking our connection to the source URL. + The `UsCensusPep_xxx` data import pipeline has experienced frequent failures (28 occurrences to date) due to issues with the source URLs. eg: [USCensusPEP\_AgeSexRace](docs_summary/uscensuspep_agesexrace_summary.md + +**Standard procedures for addressing the problem** + +1. Restart the pipeline; if the issue persists, monitor the URL performance over the next several hours/days. +2. In cases where the source URLs have been fully replaced or relocated, modify the codebase or the relevant configuration settings accordingly. +3. If the URL is completely deleted and there is no new link: +* Use **historical data** if a large amount of data is missing. (only if core teams approves) +* Update the **`latest_version.txt`** file to keep the pipeline running.(only if core team approves) + +### 4\. API Issues (Limits & Failures) + +The pipeline can fail if we hit the API rate limit (too many requests) or if the API stops working entirely. + +The BLS\_CES import process utilizes an API for data retrieval; however, executing the import twice can occasionally trigger API rate limits: [BLS Imports Issues](docs_summary/bls_imports_issues_summary.md + +**Standard procedures for addressing the problem** + +1. **Wait it Out:** If we hit a temporary rate limit or timeout issue, wait for the lockout period to end and try again later. +2. **Switch the API:** If the API is completely broken, no longer supported, or constantly failing, change the code to use an alternative API with a fallback logic. + +### 5.Missing reference errors due to additional source data + +The pipeline fails with a `missingReferenceObservationAbout` or `missingReferencesVariableMeasured` error when the source website adds new data (like new locations or new variables) that do not exist in our system yet. Because Data Commons doesn't recognize these new entities, the import fails. + +The EIA\_SEDS & US\_SAT\_ACT\_Participation imports has some data additions [EIA\_SEDS-2026-](docs_summary/eia_seds_summary.md [Support P2 - Auto Refresh Failed Imports: US\_SAT\_ACT\_Participation](bugs_summary/507394518.md) + +**Standard procedures for addressing the problem** + +1. Identify the missing place DCID or Statistical Variable from the error log. Add it to the existing `.mcf` file if the additions are valid. +2. Raise a CL to merge the updates and fix the pipeline. + diff --git a/tools/import_validation/troubleshooting_guides/rca_runbook.md b/tools/import_validation/troubleshooting_guides/rca_runbook.md new file mode 100644 index 0000000000..6a0933d883 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/rca_runbook.md @@ -0,0 +1,79 @@ +# Production Failure RCA Runbook + +# 1\. Overview + +This document provides an overview of the process that can be followed for doing the Root Cause Analysis (RCA) of the errors popping up in the auto refresh pipelines of Data Commons. + +# 2\. Background on Auto Refresh Pipelines + +The auto refresh pipelines are configured in the Cloud Batch service of the project “[](gcp_variables.md#auto_refresh_project)”. The purpose of these pipelines is to refresh the data on a regular basis for the eligible datasets that have been ingested into Data Commons already. The pipelines perform the following tasks: + +1. Download the data from the source +2. Pre-process the data as required +3. Generate .csv and .tmcf files (via PV-MAPS or custom .py scripts) +4. Validate data against the last execution +5. Ingest the data to Data Commons if every step is successfully completed + +# 3\. Errors in auto refresh pipelines + +While these pipelines are meant to refresh the data in Data Commons, it is necessary to apply checks and perform data validation before the final ingestion. These checks ensure that the right data is ingested to Data Commons and reduce the probability of wrong data ingestion in Data Commons. Consequently, these pipelines fail in case the necessary checks and validations fail while data processing. + +Apart from data checks and validation there could be other probable reasons for the pipeline failures including the compute resources exhaustion, unavailability of data at source, deletion of source, addition of new places in the source data etc. + +The process to tackle each of these issues can be different. However, the steps to perform the RCA can be generalised. The next section will cover some of the important steps that can be taken into consideration for performing RCA. + +# 4\. Steps for performing RCA + +This section highlights the steps that can be taken to perform RCA for the failed production pipelines. + +Step 1: Go to the looker [dashboard](https://lookerstudio.google.com/c/reporting/e88fda74-50c9-46c6-88aa-c84342ceba48/page/eaXdF) and get the latest status of all the pipelines from the dashboard. The probable different states of each pipelines could be as below: + +1. VALIDATION: Data validation failure. +2. FAILURE: Import job failed. +3. STAGING: Import job completed, ready for ingestion. +4. SUCCESS: Data ingestion completed +5. SKIP: Incoming data is same as production data + +Pipelines having the state as “VALIDATION” or “FAILED” are the candidates for performing RCA. + +Step 2: Next, for each failed pipeline check the priority status. The same can be determined from the [Code Search](screens/code_search_portal.png) portal. Search for the import name in the Code Search portal and determine the respective import group for each of the pipelines from the respective manifest.json files. + +| Classification | Priority Ranking | import\_groups | +| :---- | :---- | :---- | +| **P0** | 1 | SearchBranch | +| **P0** | 2 | LaeLaps | +| **P0** | 3 | SearchAim | +| **P2** | 4 | Auto1W | +| **P2** | 5 | Auto2W | +| **P2** | 6 | USCensus | +| **P2** | 7 | USBLS | +| **P2** | 8 | WorldBank | +| **P2** | 8 | OECD | +| **P2** | 9 | CDC | +| **P2** | 10 | EuroStat | +| **P2** | 11 | UNSDG | +| **P2** | 12 | BRFSS | + +Step 3: Next, go to the Cloud Batch service of the `[](gcp_variables.md#auto_refresh_project)` project and search for the respective import name in the search bar. [Screenshot](screens/Bah7SXkdpNu5r7u.png) + +Step 4: Based on the type of the failure check the logs for the latest execution of the pipeline. +**Note: For validation failures, the Cloud Batch portal might show the pipeline “Succeeded” but it is important to check the logs to verify the same** + +Step 5: Next, traverse the logs of the pipelines and search for the respective reason of the failures. A few common reasons of the failure are as mentioned below: + +1. Exit code 137 : Signifies that the pipeline failed due to exhaustion of resources +2. Validation and lint errors: Signifies that the pipeline failed because of validation and lint errors. [Screenshot](screens/3nosGztEJKr2pZV.png) +3. Pre processing script failed: Signifies that the script responsible for either downloading the data or processing the data has failed. [Screenshot](screens/9KjTNR3EF8644v8.png) + +Step 6: Based on RCA, plan the error resolution. Discuss issues with the CORE TEAM on a need basis. Also, prepare a document capturing the issues and the next actions. [Template](docs_summary/wdi_auto_failure_validation_template.md + +Step 7: Change the code (if required) and raise [CL/PR] as appropriate. + +Step 8: Once the code is merged with production, force execute the production pipelines using the command below: + +``` +/import-automation/executor/run_import.sh -p -d dc-test-executor-$USER -cloud -a -batch +``` + + + diff --git a/tools/import_validation/troubleshooting_guides/recurring_failures.md b/tools/import_validation/troubleshooting_guides/recurring_failures.md new file mode 100644 index 0000000000..3aad6a02b1 --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/recurring_failures.md @@ -0,0 +1,62 @@ +# Analysis of Recurring Production Job Failures + +## 1\. Overview + +This document outlines several recurring production job failures, some of these failures are caused by minor or temporary data deletions. By increasing the threshold, we can prevent these jobs from failing due to non-permanent issues. However, before adjusting these thresholds, we will implement new validation rules to ensure that crucial data is never lost, even during minor deletions. +More details can be found here [DC - Imports Execution Learnings ](docs_summary/dc_imports_execution_learnings.md + +## 2\. Most Recurring Job Failures + +In the previous quarter, the following production jobs experienced the highest frequency of failures: + +| Import name | Occurrences | Bug IDs | +| :---- | :---- | :---- | +| BLS\_CES\_State | 4 | [b/482946661](bugs_summary/482946661.md), [b/500945912](bugs_summary/500945912.md), [b/502090898](bugs_summary/502090898.md) | +| USCensusPEP\_Sex | 5 | [b/472605922](bugs_summary/472605922.md), [b/478186511](bugs_summary/478186511.md), [b/483219293](bugs_summary/483219293.md) | +| WorldDevelopmentIndicators | 3 | [b/470415967](bugs_summary/470415967.md), [b/482902862,](bugs_summary/482902862.md) [b/489948342](bugs_summary/489948342.md) | +| EurostatData\_Education\_Enrollment | 3 | [b/474326901](bugs_summary/474326901.md), [b/481243356,](bugs_summary/481243356.md) [b/496059688](bugs_summary/496059688.md) | +| USCensusPEP\_By\_Sex\_Race | 5 | [b/485260648](bugs_summary/485260648.md), [b/502079403](bugs_summary/502079403.md) | +| USCensusPEP\_PopulationEstimatebyRace | 5 | [b/486801970](bugs_summary/486801970.md), [b/493190090](bugs_summary/493190090.md), [b/497802532](bugs_summary/497802532.md) | +| EurostatData\_Education\_Attainment | 3 | [b/472606851](bugs_summary/472606851.md), [b/481245546](bugs_summary/481245546.md), [b/504879314](bugs_summary/504879314.md) | +| WorldBankDatasets | 3 | [b/472258775](bugs_summary/472258775.md), [b/506961224](bugs_summary/506961224.md) | +| EurostatData\_Fertility | 3 | [b/498154643](bugs_summary/498154643.md) | +| USCensusPEP\_AgeSexRace | 3 | [b/500622108](bugs_summary/500622108.md) | +| USCensusPEP\_Annual\_Population | 3 | [b/479399481](bugs_summary/479399481.md), [b/4935600377](bugs_summary/4935600377.md) | + +### 2.1 Minor Deletions imports + +The following imports frequently experience minor data losses due to broken source URLs or the removal of individual data points at the origin. To address this, we recommend increasing the deletion thresholds based on historical percentages, supplemented by automated validation rules to maintain data integrity. + +1. USCensusPEP\_Sex +2. USCensusPEP\_By\_Sex\_Race +3. USCensusPEP\_PopulationEstimatebyRace +4. USCensusPEP\_AgeSexRace +5. USCensusPEP\_Annual\_Population +6. EurostatData\_Education\_Attainment +7. EurostatData\_Fertility +8. EurostatData\_Education\_Enrollment + +### 2.2 Major Deletions Imports + +The following import jobs have experienced significant data deletions resulting from annual benchmarking conducted by the data sources, who have officially announced these changes: + +1. BLS\_CES\_State +2. WorldDevelopmentIndicators +3. WorldBankDatasets + +While these major deletions are unavoidable due to source updates, we intend to raise the threshold to accommodate and manage any minor deletions that may occur with some validation rules. + +**Note:** The current baseline threshold for deletions is **0.01%** (doesn’t contain crucial data) any deletion greater than the threshold should be stored as historical. + +## 3\. Imports with a Single Deletion Incident + +The following imports have experienced exactly one minor deletion incident from the source. If these deletions recur, we will increase the thresholds and apply automated validation rules. + +1. EurostatData\_Employment\_Per\_Sector +2. EIA\_Electricity +3. EIA\_NaturalGas +4. EIA\_Petroleum +5. EIA\_SEDS +6. EIA\_NuclearOutages +7. EurostatData\_GDP + diff --git a/tools/import_validation/troubleshooting_guides/resolution_runbook.md b/tools/import_validation/troubleshooting_guides/resolution_runbook.md new file mode 100644 index 0000000000..9ed747bb4b --- /dev/null +++ b/tools/import_validation/troubleshooting_guides/resolution_runbook.md @@ -0,0 +1,106 @@ +**PRODUCTION INCIDENT** + +**RESOLUTION RUNBOOK** + +**Organization / Product: DATA COMMONS** + +*\[\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\_\]* + +| Version | *\[e.g1. 1.0\]* | +| :---- | :---- | +| **Last Reviewed** | *\[Date\]* | +| **Runbook Owner** | *\[Name / Team\]* | +| **On-call Rotation** | *\[Link or name\]* | +| **Status Page URL** | *\[URL\]* | +| **Incident Channel** | *\[e.g. \#incidents\]* | + +## **Overview:** + +This runbook outlines resolutions for production failures, serving as a continuous guide for troubleshooting and support. + +## **Types of Production Failure** + +| TYPE | ERROR NAME | MEANING | +| :---- | :---- | :---- | +| **VALIDATION ERROR** | Found 0.01% deleted records | Production environment data deletion detected | +| **LINT ERROR** | Existence\_MissingReference\_observationAbout | Data Commons does not contain the specified mapping/place DCID. | +| **MISSING REFERENCE** | Existence\_MissingReference\_variableMeasured | The Data Commons repository is currently missing the required SV/dcid reference. | +| **Command Failed ExitCode 1** | Subprocess Failed | Failed due to code/development issue | +| **FALSE NEGATIVE** | Could be any failure out of above 4 | Failed once because of an issue with the cloud, server, or source. | + +1. ## **Validation Failure** + +| If a job fails due to data deletions, first identify the source of the deletion. Data deletions can occur in two scenarios: | +| :---- | + + **Case 1: Source-Level Deletion** + The data has been officially removed from the source system. In this case, the standard deletion handling rules apply: + + * If the deleted data is less than **0.01%** and does not contain critical data, the import threshold can be increased. + * If the deleted data is greater than **0.01%**, the deleted records must be stored as historical data. + + **Case 2: Pipeline/Code Issue** + The data deletion is caused unintentionally due to a pipeline, transformation, or code issue. In this scenario, investigate and fix the underlying issue before proceeding further. + +**Resolution Steps** + +1. Review the latest run of the job in the `[](gcp_variables.md#prod_bucket)` bucket under the `[](gcp_variables.md#base_project)` project using the import name. +2. If the job folder timestamp is older than one week, re-trigger the job to generate the latest output. +3. Examine `input0/validation/nodes_deleted.mcf` to review deletion details. +4. Validate whether the deletions originated from the source or were caused by a pipeline/code issue. +5. Apply the appropriate resolution based on the identified deletion scenario. + +**How to store Historical data** + +1. Verify the production version of the import via [Data Commons/Version](https://datacommons.org/version). +2. Locate and download the production CSV from the storage bucket, matching the date specified in the Data Commons version. +3. Run the script in your local environment, then utilize Python code to perform a difference comparison between the latest and production output CSVs. +4. Run the differ on output & historical data table\_mcf\_nodes.mcf files together & ensure no deletion flags are there. +5. Once the deleted rows are identified & no deletions with this historical file through the comparison, +6. Upload this historical data (along with its tmcf file) to the 'unresolved' path for that import, and use the importer to write it to the Knowledge Graph (KG). +7. Once the data is saved to the Knowledge Graph, use the COPY Service to move the file to CNS. +8. Add the historical path where you have added the file in the CNS as a historical record + +**Note:** If a folder for historical data already exists for this import, place the new file inside it. If not, create a new folder. + +## **LINT ERROR** + +| When dcid/mapping is missing in Data Commons | +| :---- | + +1. Review the latest run of the job in the `[](gcp_variables.md#prod_bucket)` bucket under the `[](gcp_variables.md#base_project)` project using the import name. +2. If the job folder timestamp is older than one week, re-trigger the job to generate the latest output. +3. Check for **Existence\_MissingReference\_observationAbout** in the report.json inside the input0/genmcf/report.json +4. Look for the observationAbouts’s that are throwing the error +5. Run the import locally an update or create the MCF file with the necessary mappings, keep adding mappings until the error is gone from report.json. +6. Add these mappings to the current CNS file or create a new one by creating a CL. + +| When the Data Commons API fails to locate the dcid or mapping | +| :---- | + + 1\. Existence\_FailedDcCall\_observationAbout +Existence\_FailedDcCall\_observationAbout + +## **MISSING REFERENCE** + +## + +| When SV or schema is missing in Data Commons | +| :---- | + +1. Review the latest run of the job in the `[](gcp_variables.md#prod_bucket)` bucket under the `[](gcp_variables.md#base_project)` project using the import name. +2. If the job folder timestamp is older than one week, re-trigger the job to generate the latest output. +3. Check for **Existence\_MissingReference\_variableMeasured** in the report.json inside the input0/genmcf/report.json +4. Look for the Variable measured (value-ref) that are throwing the warnings +5. Run the import locally and update or create the MCF file with the necessary mappings,keep adding mappings until the error is gone from report.json +6. Add these mappings to the current CNS file or create a new one by creating a CL. + +## **Command Failed ExitCode 1** + +| When the code fails due to any issue | +| :---- | + +1. Restart the job in the cloud environment; if it succeeds, the previous failure likely resulted from a transient environmental problem. +2. If the issue persists after re-triggering the job, execute the process in a local environment to identify and resolve code defects. + + diff --git a/tools/import_validation/troubleshooting_guides/screens/3VNMybHoDVt6kFR.png b/tools/import_validation/troubleshooting_guides/screens/3VNMybHoDVt6kFR.png new file mode 100644 index 0000000000..951bf909b5 Binary files /dev/null and b/tools/import_validation/troubleshooting_guides/screens/3VNMybHoDVt6kFR.png differ diff --git a/tools/import_validation/troubleshooting_guides/screens/3nosGztEJKr2pZV.png b/tools/import_validation/troubleshooting_guides/screens/3nosGztEJKr2pZV.png new file mode 100644 index 0000000000..58ef6aae83 Binary files /dev/null and b/tools/import_validation/troubleshooting_guides/screens/3nosGztEJKr2pZV.png differ diff --git a/tools/import_validation/troubleshooting_guides/screens/3oK73RXvox3xeTM.png b/tools/import_validation/troubleshooting_guides/screens/3oK73RXvox3xeTM.png new file mode 100644 index 0000000000..dcacfd8e00 Binary files /dev/null and b/tools/import_validation/troubleshooting_guides/screens/3oK73RXvox3xeTM.png differ diff --git a/tools/import_validation/troubleshooting_guides/screens/497zku465XJzsJt.png b/tools/import_validation/troubleshooting_guides/screens/497zku465XJzsJt.png new file mode 100644 index 0000000000..dd4b305e2d Binary files /dev/null and b/tools/import_validation/troubleshooting_guides/screens/497zku465XJzsJt.png differ diff --git a/tools/import_validation/troubleshooting_guides/screens/4JKbhAxnJz6SGbJ.png b/tools/import_validation/troubleshooting_guides/screens/4JKbhAxnJz6SGbJ.png new file mode 100644 index 0000000000..a9ff201c45 Binary files /dev/null and b/tools/import_validation/troubleshooting_guides/screens/4JKbhAxnJz6SGbJ.png differ diff --git a/tools/import_validation/troubleshooting_guides/screens/4pprS7vyu4Cba6C.png b/tools/import_validation/troubleshooting_guides/screens/4pprS7vyu4Cba6C.png new file mode 100644 index 0000000000..e2329b443e Binary files /dev/null and b/tools/import_validation/troubleshooting_guides/screens/4pprS7vyu4Cba6C.png differ diff --git a/tools/import_validation/troubleshooting_guides/screens/5Nw3dA9a3iH7GAM.png b/tools/import_validation/troubleshooting_guides/screens/5Nw3dA9a3iH7GAM.png new file mode 100644 index 0000000000..5be646e142 Binary files /dev/null and b/tools/import_validation/troubleshooting_guides/screens/5Nw3dA9a3iH7GAM.png differ diff --git a/tools/import_validation/troubleshooting_guides/screens/7ovdMycUe7tiADw.png b/tools/import_validation/troubleshooting_guides/screens/7ovdMycUe7tiADw.png new file mode 100644 index 0000000000..a8da14af0e Binary files /dev/null and b/tools/import_validation/troubleshooting_guides/screens/7ovdMycUe7tiADw.png differ diff --git a/tools/import_validation/troubleshooting_guides/screens/8znJ3XadmGGRmdL.png b/tools/import_validation/troubleshooting_guides/screens/8znJ3XadmGGRmdL.png new file mode 100644 index 0000000000..26df4e7491 Binary files /dev/null and b/tools/import_validation/troubleshooting_guides/screens/8znJ3XadmGGRmdL.png differ diff --git a/tools/import_validation/troubleshooting_guides/screens/9KjTNR3EF8644v8.png b/tools/import_validation/troubleshooting_guides/screens/9KjTNR3EF8644v8.png new file mode 100644 index 0000000000..4e8ccacde2 Binary files /dev/null and b/tools/import_validation/troubleshooting_guides/screens/9KjTNR3EF8644v8.png differ diff --git a/tools/import_validation/troubleshooting_guides/screens/B25DugjwV2js2TJ.png b/tools/import_validation/troubleshooting_guides/screens/B25DugjwV2js2TJ.png new file mode 100644 index 0000000000..63c1c7c706 Binary files /dev/null and b/tools/import_validation/troubleshooting_guides/screens/B25DugjwV2js2TJ.png differ diff --git a/tools/import_validation/troubleshooting_guides/screens/BWYDfWvTNZoixSE.png b/tools/import_validation/troubleshooting_guides/screens/BWYDfWvTNZoixSE.png new file mode 100644 index 0000000000..4b0b5d963a Binary files /dev/null and b/tools/import_validation/troubleshooting_guides/screens/BWYDfWvTNZoixSE.png differ diff --git a/tools/import_validation/troubleshooting_guides/screens/Bah7SXkdpNu5r7u.png b/tools/import_validation/troubleshooting_guides/screens/Bah7SXkdpNu5r7u.png new file mode 100644 index 0000000000..0ae30c99bd Binary files /dev/null and b/tools/import_validation/troubleshooting_guides/screens/Bah7SXkdpNu5r7u.png differ diff --git a/tools/import_validation/troubleshooting_guides/screens/code_search_portal.png b/tools/import_validation/troubleshooting_guides/screens/code_search_portal.png new file mode 100644 index 0000000000..671d87aaed Binary files /dev/null and b/tools/import_validation/troubleshooting_guides/screens/code_search_portal.png differ diff --git a/tools/import_validation/troubleshooting_guides/screens/eurostat_fertility_textproto_source.png b/tools/import_validation/troubleshooting_guides/screens/eurostat_fertility_textproto_source.png new file mode 100644 index 0000000000..0c12f7c74c Binary files /dev/null and b/tools/import_validation/troubleshooting_guides/screens/eurostat_fertility_textproto_source.png differ diff --git a/tools/import_validation/troubleshooting_guides/screens/wdi_textproto_source.png b/tools/import_validation/troubleshooting_guides/screens/wdi_textproto_source.png new file mode 100644 index 0000000000..ac1e4a6d42 Binary files /dev/null and b/tools/import_validation/troubleshooting_guides/screens/wdi_textproto_source.png differ