Skip to main content
Data Quality Cover
Datazone’s Data Quality feature lets you attach a Ruleset — a versioned, git-backed collection of rules — to a dataset. Each rule checks one thing (nulls, duplicates, freshness, a value range, a custom SQL condition, …) and rolls up into a per-ruleset and per-project quality score you can track over time.

Key Features

  • Git-backed - rulesets live in your project’s repository as YAML, the same way Flows, Actions, and Knowledge Objects do
  • Typed checks - missing values, duplicates, format/regex, value ranges, freshness, row count minimums, schema compliance, or fully custom SQL
  • Severity-weighted scoring - critical/high/medium/low rules contribute proportionally to the quality score
  • Multiple triggers - run automatically after every dataset load, on a cron schedule, or manually
  • AI-assisted authoring - Orion can propose rules for a dataset from a natural-language prompt
  • A shared Rule Library - reusable built-in and organization-defined rule templates you can drop into any ruleset
Rulesets currently target datasets only. view and knowledge_object are accepted by the YAML schema but are not yet resolvable — a ruleset on anything other than a dataset is skipped.

Key Concepts

The rest of this page documents the full git-backed model, including per-branch definitions and branch-aware preview. In the current UI, rulesets are only ever edited and evaluated on the main branch — there is no branch switcher yet, even though the API itself is branch-aware.

Configuration

A ruleset is declared as a rulesets: entry in your project’s config.yml, parallel to flows:, actions:, and objects::
config.yml
The referenced file holds the ruleset’s rules:
rulesets/customers_ruleset_ab12cd.yaml

Rule fields

Renaming a rule’s alias is treated as deleting the old rule and creating a new one — its evaluation history and sparkline reset.

Check Types

duplicate_count and format_check/range_check/schema_compliance are equality-style checks — they fail on any nonzero count, independent of the configured unit. freshness and row_count_minimum have no concept of a “failed row”, only a metric value compared against the threshold.

Thresholds

range_check is the one check type that needs an explicit between comparator with both bounds — anything else on a range_check rule fails validation:
A rule with no thresholds at all is valid but is reported as not_evaluated — useful while you’re still deciding on a boundary.

Rule Library

Datazone ships a set of built-in rule templates, and you can add your own project- or organization-scoped ones. A ruleset can reference a template instead of writing out check/thresholds:
Built-in templates include Not Null Check, Primary Key Uniqueness, Email Format, UK Post Code Format, Value Range, Daily Freshness SLA, Minimum Row Count, and Schema Compliance. The Add from Library action in the ruleset editor searches templates by column type and dimension and splices a ready-to-use YAML fragment into your rules.

AI-Assisted Rule Proposals

The ruleset editor’s Generate with Orion action asks an AI agent to propose rules for the dataset you’re editing, from an optional free-text prompt and the rules you already have in the draft buffer. The agent samples the dataset, searches the Rule Library first, and only falls back to a custom SQL check when nothing in the library fits — every proposal is re-validated server-side before it’s returned, so you always get YAML that’s ready to save. Nothing is written until you review and accept the proposed entries. See Orion AI for more on how Datazone’s AI agents are grounded in your data.

Validate, Preview & Apply

Editing a ruleset follows the same git-backed flow as Flows and Knowledge Objects — nothing is evaluated or scored until it’s committed:
  1. Validate — parses the YAML and checks it against the dataset’s live schema; returns per-field issues without touching git.
  2. Preview — additionally diffs the resolved rules against what’s already committed on the target branch, returning to_create / to_update / to_remove (with field-level changes) so you can see exactly what applying will do.
  3. Apply — commits the YAML to the project repository (creating rulesets/<slug>_<hash>.yaml and registering it in config.yml on first save) and immediately materializes the ruleset and its rules — the same code path a repository sync would run.
Because apply is a real commit, ruleset writes are gated by the same write_repository branch permission used for every other project resource — see Policy → Branch Protection if main is protected in your project.

Running Rulesets

A ruleset’s trigger controls when it runs: Every ruleset can also be run on demand from the UI (“Run Now”) or via the API. Each run evaluates every rule in the ruleset — there’s no single-rule run. Each rule’s result is one of: A run also produces a project-level aggregate alongside its ruleset-level result, rolling up every active ruleset in the project — this is what powers the project-wide Data Quality overview.

Issues & History

A quality issue is simply a rule whose latest cached result is warning or failed — issues aren’t a separate record, they’re derived live from every active ruleset’s most recent run. The issues list supports the same filters=[field][$op]:value query syntax as other list endpoints, letting you scope to an entity, a project, or a specific ruleset. For a single rule, its result history across past runs is available as a time-ordered list — this is what drives the sparkline shown in the rule detail panel.

Quality Score

The quality score is a severity-weighted pass rate:
  • Each rule’s severity contributes a weight: critical=5, high=3, medium=2, low=1
  • Each result contributes a factor: passed=1, warning=0.5, failed=0
  • error, no_data, and not_evaluated results are excluded entirely from the calculation
  • score = round(sum(weight × factor) / sum(weight) × 100), or 0 if nothing was scored
Scores are computed at both ruleset scope and project scope on every run, and a daily snapshot is kept for history — score-history endpoints return one point per calendar day over a range (default: the last 30 days), with null for days that had no run, so a chart can show a real gap instead of a false zero. A live summary (current score, active/evaluated rule counts, coverage percent) is also available without waiting for history to catch up.

Authorization

Data Quality introduces two new project-scoped resource types alongside the ones described in Policy: Both follow the standard project hierarchy (project:<id>:ruleset:*, project:<id>:ruleset:<id>), so you can grant Data Quality access the same way you’d grant access to datasets or flows.

Where to Find It

Data Quality has two entry points in the app, both showing Overview / Rulesets / Issues tabs:
  • Project → Catalog → Data Quality — rulesets and issues across every dataset in the project
  • Dataset → Data Quality — scoped to a single dataset, with “Run all” and “Add ruleset”
Rules are edited YAML-first — there’s no per-check-type form. The editor is a YAML pane with column-name autocomplete next to a read-only, parsed view of the same rules, plus Add from Library, Generate with Orion, Validate, and Preview actions before you save. Opening a rule shows its check configuration, current metric, threshold, generated SQL (read-only, for verification), and its recent history.

Policy

Grant access to rulesets and rule templates

Projects

How git-backed project resources are structured and synced

Flows

Another git-backed, YAML-defined project resource

Orion AI

How Datazone’s AI agents are grounded in your data