You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: CONTRIBUTING.md
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -58,7 +58,7 @@ poetry install --with lint
58
58
59
59
## Installation for Development
60
60
61
-
We are utilising Poetry for build dependency management and packaging. If you're on a system that has `Make` available, you can simply run `make install` to setup a local virtual environment with all the dependencies installed (this won't install Poetry for you).
61
+
We are utilising Poetry for build dependency management and packaging. To install simply run `poetry install`. If you need to update any of the dependencies then you will need to run `poetry lock` before the install command. **Please always review the contents of the lock before installing. We have a `min-release-age` of `10` which means that packages will not be updated until they have been released for at least 10 days.
[](https://github.com/NHSDigital/data-validation-engine/actions/workflows/ci_testing.yml)
The Data Validation Engine (DVE) is a configurationdriven data validation library built and utilised by NHS England. Currently the package has been reverted from v1.0.0 release to a 0.x as we feel the package is not yet mature enough to be considered a 1.0.0 release. So please bear this in mind if reading through the commits and references to a v1+ release when on v0.x.
12
+
The Data Validation Engine (DVE) is a configuration-driven data validation library created and used by NHS England. It lets users define validation rules once and apply them across multiple dataset collections - supporting consistent, accurate data checks.
13
13
14
-
As mentioned above, the DVE is "configuration driven" which means the majority of development for you as a user will be building a JSON document to describe how the data will be validated. The JSON document is known as a `dischema` file and example files can be accessed [here](https://github.com/NHSDigital/data-validation-engine/tree/main/tests/testdata). If you'd like to learn more about JSON document and how to build one from scratch, then please read the documentation [here](https://nhsdigital.github.io/data-validation-engine/).
14
+
__The DVE offers__:
15
15
16
-
Once a dischema file has been defined, you are ready to use the DVE. The DVE is typically orchestrated based on four key "services". These are...
16
+
- SQL configuration-based validations
17
+
- Format normalization to Parquet for a unified data representation
18
+
- Data modelling and typecasting
19
+
- Business-rule validations executed on supported backends such as Spark and DuckDB, with the option to add custom backends
20
+
- Deriving new fields and entities
21
+
- Clear validation reporting, including summary insights and record-level error messages
17
22
18
-
|| Service | Purpose |
19
-
| -- | ------- | ------- |
20
-
| 1. | File Transformation | This service will take submitted files and turn them into stringified parquet file(s) to ensure that a consistent data structure can be passed through the other services. |
21
-
| 2. | Data Contract | This service will validate and perform type casting against a stringified parquet file using [pydantic models](https://docs.pydantic.dev/1.10/). |
22
-
| 3. | Business Rules | The business rules service will perform more complex validations such as comparisons between fields and tables, aggregations, filters etc to generate new entities. |
23
-
| 4. | Error Reports | The error reports service will take all the errors raised in previous services and surface them into a readable format for a downstream users/service. Currently, this implemented to be an excel spreadsheet but could be reconfigured to meet other requirements/use cases. |
24
-
25
-
If you'd like more detailed documentation around these services the please read the extended documentation [here](https://nhsdigital.github.io/data-validation-engine/).
26
-
27
-
The DVE has been designed in a way that's modular and can support users who just want to utilise specific "services" from the DVE (i.e. just the file transformation + data contract). Additionally, the DVE is designed to support different backend implementations. As part of the base installation of DVE, you will find backend support for `Spark` and `DuckDB`. So, if you need a `MySQL` backend implementation, you can implement this yourself. Given our organisations requirements, it will be unlikely that we add anymore specific backend implementations into the base package beyond Spark and DuckDB. So, if you are unable to implement this yourself, I would recommend reading the guidance on [requesting new features and raising bug reports here](#requesting-new-features-and-raising-bug-reports).
28
-
29
-
Additionally, if you'd like to contribute a new backend implementation into the base DVE package, then please look at the [Contributing](#Contributing) section.
23
+
As mentioned above, the DVE is "configuration driven" which means the majority of development for you as a user will be building a JSON document to describe how the data will be validated. The JSON document is known as a `dischema` (data ingest schema) file and example files can be accessed [here ↗️](https://github.com/NHSDigital/data-validation-engine/tree/main/tests/testdata). If you'd like to learn more about JSON document and how to build one from scratch, then please read the documentation [here ↗️](https://nhsdigital.github.io/data-validation-engine/).
30
24
31
25
## Installation and usage
32
26
33
-
The DVE is a Python package and can be installed using package managers such as [pip](https://pypi.org/project/pip/). As of the latest release we support Python 3.10 & 3.11, with Spark v3.4 and DuckDB v1.1. In the future we will be looking to upgrade the DVE to working on a higher versions of Python, DuckDB and Spark.
34
-
35
-
If you're planning to use the Spark backend implementation, you will also need OpenJDK 11 installed.
36
-
37
-
Python dependencies are listed in `pyproject.toml`.
38
-
39
-
To install the DVE package you can simply install using a package manager such as [pip](https://pypi.org/project/pip/).
40
-
41
-
```
42
-
pip install data-validation-engine
43
-
```
44
-
45
-
*Note - Only versions >=0.6.2 are available on PyPi. For older versions please install directly from the git repo or build from source.*
46
-
47
-
Once you have installed the DVE you are ready to use it. For guidance on how to create your dischema JSON document (configuration), please read the [documentation](https://nhsdigital.github.io/data-validation-engine/).
48
-
49
-
Version 0.0.1 does support a working Python 3.7 installation. However, we will not be supporting any issues with that version of the DVE if you choose to use it. __Use at your own risk__.
27
+
Please see the documentation [here ↗️](https://nhsdigital.github.io/data-validation-engine/user_guidance/install/).
50
28
51
29
## Requesting new features and raising bug reports
52
30
**Before creating new issues, please check to see if the same bug/feature has been created already. Where a duplicate is created, the ticket will be closed and referenced to an existing issue.**
53
31
54
-
If you have spotted a bug with the DVE then please raise an issue [here](https://github.com/nhsengland/Data-Validation-Engine/issues) using the "bug template".
32
+
If you have spotted a bug with the DVE then please raise an issue [here ↗️](https://github.com/nhsengland/Data-Validation-Engine/issues) using the "bug template".
55
33
56
34
If you have feature request then please follow the same process whilst using the "Feature request template".
57
35
@@ -63,7 +41,7 @@ Below is a list of features that we would like to implement or have been request
63
41
| Uplift to Python 3.11 | 0.2.0 | Yes |
64
42
| Uplift Pyspark to 3.5 | 0.8.0 | Yes |
65
43
| Allow DVE to run on Python 3.12+ | 0.8.0 | Yes |
66
-
| Upgrade to Pydantic 2.0 | 0.9.0 |No|
44
+
| Upgrade to Pydantic 2.0 | 0.9.0 |Yes|
67
45
| Uplift Pyspark to 4.0+ | TBA | No |
68
46
| Polars upgrade to v1+ | TBA | No |
69
47
| DuckDB upgrade to v1.5+ | TBA | No |
@@ -73,7 +51,7 @@ Below is a list of features that we would like to implement or have been request
73
51
If you are interested in getting any of the unreleased features listed above available, then please read the [Contributing](#Contributing) section and then submit us a pull request.
74
52
75
53
## Contributing
76
-
Please see guidance [here](https://github.com/NHSDigital/data-validation-engine/blob/main/CONTRIBUTE.md).
54
+
Please see guidance [here ↗️](https://github.com/NHSDigital/data-validation-engine/blob/main/CONTRIBUTE.md).
77
55
78
56
## Legal
79
57
This codebase is released under the MIT License. This covers both the codebase and any sample code in the documentation.
We take security and the protection of private data extremely seriously. If you believe you have found a vulnerability or other issue which has compromised or could compromise the security of any of our systems or private data managed by our systems, please do not hesitate to contact us using the method outlined below.
4
+
5
+
## Table of contents
6
+
7
+
-[Security](#security)
8
+
-[Table of contents](#table-of-contents)
9
+
-[Reporting a vulnerability](#reporting-a-vulnerability)
"description": "A complete entity specification within data ingest",
@@ -10,7 +10,8 @@
10
10
"description": "A mapping of field names to their Python types. These will either be strings representing Python types (if there are no argumements to the type), and field specification objects otherwise",
11
11
"additionalProperties": {
12
12
"$ref": "field.schema.json"
13
-
}
13
+
},
14
+
"minProperties": 1
14
15
},
15
16
"aliases": {
16
17
"description": "A mapping of field name to allowed field alias",
"description": "A mapping of field names to the custom error code and message required if these fields were to fail validation during the data contract phase. For nested fields, these should be specified using struct '.' notation (eg. fieldA.fieldB.fieldC)",
0 commit comments