> ## Documentation Index
> Fetch the complete documentation index at: https://docs.datazone.co/llms.txt
> Use this file to discover all available pages before exploring further.

# Data Quality

> Define, run, and monitor data quality rules against your datasets with git-backed Rulesets

<Frame>
  <img src="https://mintcdn.com/datazone/goPCua848baak0G_/images/covers/data-quality.png?fit=max&auto=format&n=goPCua848baak0G_&q=85&s=1ac3eb533f30b70365147ecf8bf0b6ff" alt="Data Quality Cover" width="1920" height="741" data-path="images/covers/data-quality.png" />
</Frame>

Datazone's **Data Quality** feature lets you attach a **Ruleset** — a versioned, git-backed collection of rules — to a dataset. Each rule checks one thing (nulls, duplicates, freshness, a value range, a custom SQL condition, …) and rolls up into a per-ruleset and per-project **quality score** you can track over time.

### Key Features

* **Git-backed** - rulesets live in your project's repository as YAML, the same way [Flows](/reference/flows/overview), [Actions](/reference/development/actions), and [Knowledge Objects](/reference/knowledge-objects/overview) do
* **Typed checks** - missing values, duplicates, format/regex, value ranges, freshness, row count minimums, schema compliance, or fully custom SQL
* **Severity-weighted scoring** - `critical`/`high`/`medium`/`low` rules contribute proportionally to the quality score
* **Multiple triggers** - run automatically after every dataset load, on a cron schedule, or manually
* **AI-assisted authoring** - Orion can propose rules for a dataset from a natural-language prompt
* **A shared Rule Library** - reusable built-in and organization-defined rule templates you can drop into any ruleset

<Note>
  Rulesets currently target **datasets** only. `view` and `knowledge_object` are accepted by the YAML schema but are not yet resolvable — a ruleset on anything other than a dataset is skipped.
</Note>

## Key Concepts

| Concept       | Description                                                                                                                   |
| ------------- | ----------------------------------------------------------------------------------------------------------------------------- |
| **Ruleset**   | The stable object bound to one dataset — has a status (`active`/`paused`/`draft`), a trigger, and a denormalized `rule_count` |
| **Rule**      | One check inside a ruleset — identified by its `alias`, which also doubles as its id                                          |
| **Check**     | The type of validation a rule performs (see [Check Types](#check-types))                                                      |
| **Dimension** | The quality category a rule measures — `completeness`, `uniqueness`, `validity`, `freshness`, `schema`, or `volume`           |
| **Threshold** | The pass/warning/fail boundary for a rule's metric                                                                            |
| **Run**       | One evaluation of a ruleset (or of every active ruleset in a project) — produces per-rule results and a score                 |
| **Issue**     | A rule whose most recent result is `warning` or `failed`                                                                      |

<Warning>
  The rest of this page documents the full git-backed model, including per-branch definitions and branch-aware preview. In the current UI, rulesets are only ever edited and evaluated on the **`main`** branch — there is no branch switcher yet, even though the API itself is branch-aware.
</Warning>

## Configuration

A ruleset is declared as a `rulesets:` entry in your project's `config.yml`, parallel to `flows:`, `actions:`, and `objects:`:

```yaml config.yml theme={null}
project_name: my-project
project_id: proj_abc123

rulesets:
  - path: rulesets/customers_ruleset_ab12cd.yaml
```

The referenced file holds the ruleset's rules:

```yaml rulesets/customers_ruleset_ab12cd.yaml theme={null}
entity_alias: customers
version: 1

rules:
  - alias: not_null_email
    name: Email must be populated
    severity: critical
    check:
      type: missing_percent
      column: email
    thresholds:
      warning: 0.5
      failure: 1
      unit: percent

  - alias: unique_customer_id
    name: Customer id must be unique
    severity: high
    check:
      type: duplicate_count
      column: customer_id
    thresholds:
      failure: 0
```

| Field                        | Type   | Default   | Description                                                        |
| ---------------------------- | ------ | --------- | ------------------------------------------------------------------ |
| `entity_alias` / `entity_id` | string | —         | The dataset this ruleset checks — exactly one of the two, required |
| `entity_type`                | string | `dataset` | Only `dataset` is resolvable today                                 |
| `version`                    | int    | —         | Schema version of the file                                         |
| `rules`                      | array  | —         | The list of rules, see below                                       |

### Rule fields

| Field                  | Type   | Description                                                                                                                             |
| ---------------------- | ------ | --------------------------------------------------------------------------------------------------------------------------------------- |
| `alias`                | string | Required. URL-safe (`^[a-zA-Z0-9_-]+$`), unique within the file, and doubles as the rule's id                                           |
| `name`                 | string | Required. Human-readable label                                                                                                          |
| `severity`             | string | Required. One of `critical`, `high`, `medium`, `low` — weights the rule's contribution to the quality score                             |
| `dimension`            | string | (Optional) One of `completeness`, `uniqueness`, `validity`, `freshness`, `schema`, `volume`. Inferred from `check.type` when omitted    |
| `check`                | object | Either `check` or `template`+`binding`, never both — see [Check Types](#check-types)                                                    |
| `thresholds`           | object | `warning` / `failure` values and a `unit` (`percent`, `count`, or `hours`) — see [Thresholds](#thresholds)                              |
| `template` / `binding` | object | Instantiates a [Rule Library](#rule-library) template instead of a raw `check` — `template.id` (+ optional `version`), `binding.column` |

<Warning>
  Renaming a rule's `alias` is treated as deleting the old rule and creating a new one — its evaluation history and sparkline reset.
</Warning>

## Check Types

| Check (YAML alias) | Canonical type      | Requires                        | Default dimension                 | What it measures                                                |
| ------------------ | ------------------- | ------------------------------- | --------------------------------- | --------------------------------------------------------------- |
| `missing_percent`  | `missing_value`     | `column`                        | `completeness`                    | Percent of null values in the column                            |
| `duplicate_count`  | `duplicate_value`   | `column`                        | `uniqueness`                      | Count of duplicate values (`count() - uniqExact(column)`)       |
| `format_check`     | `format_regex`      | `column`, `pattern`             | `validity`                        | Rows whose value doesn't match the regex `pattern`              |
| `range_check`      | `range_check`       | `column`, a `between` threshold | `validity`                        | Rows outside a `[min, max]` value range                         |
| `freshness`        | `freshness`         | `column` (date/datetime)        | `freshness`                       | Hours since `max(column)`                                       |
| —                  | `row_count_minimum` | —                               | `volume`                          | Total row count                                                 |
| —                  | `schema_compliance` | —                               | `schema`                          | Diffs live columns against the dataset's last known schema      |
| `custom_query`     | `custom_query`      | `sql`                           | none (set `dimension` explicitly) | Any SQL selecting `rows_scanned`, `failed_rows`, `metric_value` |

```yaml theme={null}
# Custom SQL check
- alias: r1
  name: Orders total is never negative
  severity: low
  check:
    type: custom_query
    sql: |
      SELECT
        count() AS rows_scanned,
        countIf(amount < 0) AS failed_rows,
        countIf(amount < 0) AS metric_value
      FROM orders
```

<Note>
  `duplicate_count` and `format_check`/`range_check`/`schema_compliance` are equality-style checks — they fail on any nonzero count, independent of the configured `unit`. `freshness` and `row_count_minimum` have no concept of a "failed row", only a metric value compared against the threshold.
</Note>

## Thresholds

```yaml theme={null}
thresholds:
  warning: 0.5   # soft boundary — result is "warning" past this point
  failure: 1     # hard boundary — result is "failed" past this point
  unit: percent  # percent | count | hours
```

`range_check` is the one check type that needs an explicit **`between`** comparator with both bounds — anything else on a `range_check` rule fails validation:

```yaml theme={null}
thresholds:
  comparator: between
  value: 0
  secondary_value: 100000
  unit: count
```

A rule with no `thresholds` at all is valid but is reported as `not_evaluated` — useful while you're still deciding on a boundary.

## Rule Library

Datazone ships a set of built-in rule templates, and you can add your own project- or organization-scoped ones. A ruleset can reference a template instead of writing out `check`/`thresholds`:

```yaml theme={null}
- alias: from_library_email_format
  name: Email format
  severity: high
  template:
    id: 6aa93df1442d41e95fbfcd47
    version: 1
  binding:
    column: email
```

Built-in templates include **Not Null Check**, **Primary Key Uniqueness**, **Email Format**, **UK Post Code Format**, **Value Range**, **Daily Freshness SLA**, **Minimum Row Count**, and **Schema Compliance**. The **Add from Library** action in the ruleset editor searches templates by column type and dimension and splices a ready-to-use YAML fragment into your rules.

## AI-Assisted Rule Proposals

The ruleset editor's **Generate with Orion** action asks an AI agent to propose rules for the dataset you're editing, from an optional free-text prompt and the rules you already have in the draft buffer. The agent samples the dataset, searches the Rule Library first, and only falls back to a custom SQL check when nothing in the library fits — every proposal is re-validated server-side before it's returned, so you always get YAML that's ready to save. Nothing is written until you review and accept the proposed entries. See [Orion AI](/reference/intelligent-apps/orion-ai) for more on how Datazone's AI agents are grounded in your data.

## Validate, Preview & Apply

Editing a ruleset follows the same **git-backed** flow as Flows and Knowledge Objects — nothing is evaluated or scored until it's committed:

1. **Validate** — parses the YAML and checks it against the dataset's live schema; returns per-field issues without touching git.
2. **Preview** — additionally diffs the resolved rules against what's already committed on the target branch, returning `to_create` / `to_update` / `to_remove` (with field-level changes) so you can see exactly what applying will do.
3. **Apply** — commits the YAML to the project repository (creating `rulesets/<slug>_<hash>.yaml` and registering it in `config.yml` on first save) and immediately materializes the ruleset and its rules — the same code path a repository sync would run.

<Note>
  Because apply is a real commit, ruleset writes are gated by the same `write_repository` branch permission used for every other project resource — see [Policy → Branch Protection](/reference/development/policy#branch-protection) if `main` is protected in your project.
</Note>

## Running Rulesets

A ruleset's `trigger` controls when it runs:

| Trigger                               | Behavior                                                                    |
| ------------------------------------- | --------------------------------------------------------------------------- |
| `after_dataset_transaction` (default) | Runs automatically every time the dataset receives a new transaction (load) |
| `on_schedule`                         | Runs on the cron `expression` you set on the ruleset                        |
| `manual_only`                         | Only runs when explicitly triggered                                         |

Every ruleset can also be run on demand from the UI ("Run Now") or via the API. Each run evaluates **every rule in the ruleset** — there's no single-rule run.

Each rule's result is one of:

| Result          | Meaning                                                      |
| --------------- | ------------------------------------------------------------ |
| `passed`        | Metric is within the threshold                               |
| `warning`       | Metric crossed the soft `warning` boundary but not `failure` |
| `failed`        | Metric crossed the `failure` boundary                        |
| `error`         | The check itself failed to execute                           |
| `no_data`       | The metric came back null (e.g. empty table)                 |
| `not_evaluated` | The rule has no threshold configured                         |

A run also produces a **project-level** aggregate alongside its ruleset-level result, rolling up every active ruleset in the project — this is what powers the project-wide Data Quality overview.

## Issues & History

A **quality issue** is simply a rule whose latest cached result is `warning` or `failed` — issues aren't a separate record, they're derived live from every active ruleset's most recent run. The issues list supports the same `filters=[field][$op]:value` query syntax as other list endpoints, letting you scope to an entity, a project, or a specific ruleset.

For a single rule, its result history across past runs is available as a time-ordered list — this is what drives the sparkline shown in the rule detail panel.

## Quality Score

The quality score is a **severity-weighted pass rate**:

* Each rule's severity contributes a weight: `critical=5`, `high=3`, `medium=2`, `low=1`
* Each result contributes a factor: `passed=1`, `warning=0.5`, `failed=0`
* `error`, `no_data`, and `not_evaluated` results are excluded entirely from the calculation
* `score = round(sum(weight × factor) / sum(weight) × 100)`, or `0` if nothing was scored

Scores are computed at both **ruleset** scope and **project** scope on every run, and a daily snapshot is kept for history — score-history endpoints return one point per calendar day over a range (default: the last 30 days), with `null` for days that had no run, so a chart can show a real gap instead of a false zero. A live summary (current score, active/evaluated rule counts, coverage percent) is also available without waiting for history to catch up.

## Authorization

Data Quality introduces two new project-scoped resource types alongside the ones described in [Policy](/reference/development/policy):

| Resource        | Actions                                        | Description                                                    |
| --------------- | ---------------------------------------------- | -------------------------------------------------------------- |
| `ruleset`       | `read`, `write`, `create`, `delete`, `execute` | A dataset's ruleset. `execute` gates manually triggering a run |
| `rule_template` | `read`, `write`, `create`, `delete`            | Entries in the Rule Library                                    |

Both follow the standard project hierarchy (`project:<id>:ruleset:*`, `project:<id>:ruleset:<id>`), so you can grant Data Quality access the same way you'd grant access to datasets or flows.

## Where to Find It

Data Quality has two entry points in the app, both showing **Overview / Rulesets / Issues** tabs:

* **Project → Catalog → Data Quality** — rulesets and issues across every dataset in the project
* **Dataset → Data Quality** — scoped to a single dataset, with "Run all" and "Add ruleset"

Rules are edited YAML-first — there's no per-check-type form. The editor is a YAML pane with column-name autocomplete next to a read-only, parsed view of the same rules, plus **Add from Library**, **Generate with Orion**, **Validate**, and **Preview** actions before you save. Opening a rule shows its check configuration, current metric, threshold, generated SQL (read-only, for verification), and its recent history.

## Related Resources

<CardGroup cols={2}>
  <Card title="Policy" icon="lock" href="/reference/development/policy">
    Grant access to rulesets and rule templates
  </Card>

  <Card title="Projects" icon="folder" href="/reference/development/project">
    How git-backed project resources are structured and synced
  </Card>

  <Card title="Flows" icon="diagram-project" href="/reference/flows/overview">
    Another git-backed, YAML-defined project resource
  </Card>

  <Card title="Orion AI" icon="sparkles" href="/reference/intelligent-apps/orion-ai">
    How Datazone's AI agents are grounded in your data
  </Card>
</CardGroup>
