
Key Features
- Git-backed - rulesets live in your project’s repository as YAML, the same way Flows, Actions, and Knowledge Objects do
- Typed checks - missing values, duplicates, format/regex, value ranges, freshness, row count minimums, schema compliance, or fully custom SQL
- Severity-weighted scoring -
critical/high/medium/lowrules contribute proportionally to the quality score - Multiple triggers - run automatically after every dataset load, on a cron schedule, or manually
- AI-assisted authoring - Orion can propose rules for a dataset from a natural-language prompt
- A shared Rule Library - reusable built-in and organization-defined rule templates you can drop into any ruleset
Rulesets currently target datasets only.
view and knowledge_object are accepted by the YAML schema but are not yet resolvable — a ruleset on anything other than a dataset is skipped.Key Concepts
Configuration
A ruleset is declared as arulesets: entry in your project’s config.yml, parallel to flows:, actions:, and objects::
config.yml
rulesets/customers_ruleset_ab12cd.yaml
Rule fields
Check Types
duplicate_count and format_check/range_check/schema_compliance are equality-style checks — they fail on any nonzero count, independent of the configured unit. freshness and row_count_minimum have no concept of a “failed row”, only a metric value compared against the threshold.Thresholds
range_check is the one check type that needs an explicit between comparator with both bounds — anything else on a range_check rule fails validation:
thresholds at all is valid but is reported as not_evaluated — useful while you’re still deciding on a boundary.
Rule Library
Datazone ships a set of built-in rule templates, and you can add your own project- or organization-scoped ones. A ruleset can reference a template instead of writing outcheck/thresholds:
AI-Assisted Rule Proposals
The ruleset editor’s Generate with Orion action asks an AI agent to propose rules for the dataset you’re editing, from an optional free-text prompt and the rules you already have in the draft buffer. The agent samples the dataset, searches the Rule Library first, and only falls back to a custom SQL check when nothing in the library fits — every proposal is re-validated server-side before it’s returned, so you always get YAML that’s ready to save. Nothing is written until you review and accept the proposed entries. See Orion AI for more on how Datazone’s AI agents are grounded in your data.Validate, Preview & Apply
Editing a ruleset follows the same git-backed flow as Flows and Knowledge Objects — nothing is evaluated or scored until it’s committed:- Validate — parses the YAML and checks it against the dataset’s live schema; returns per-field issues without touching git.
- Preview — additionally diffs the resolved rules against what’s already committed on the target branch, returning
to_create/to_update/to_remove(with field-level changes) so you can see exactly what applying will do. - Apply — commits the YAML to the project repository (creating
rulesets/<slug>_<hash>.yamland registering it inconfig.ymlon first save) and immediately materializes the ruleset and its rules — the same code path a repository sync would run.
Because apply is a real commit, ruleset writes are gated by the same
write_repository branch permission used for every other project resource — see Policy → Branch Protection if main is protected in your project.Running Rulesets
A ruleset’strigger controls when it runs:
Every ruleset can also be run on demand from the UI (“Run Now”) or via the API. Each run evaluates every rule in the ruleset — there’s no single-rule run.
Each rule’s result is one of:
A run also produces a project-level aggregate alongside its ruleset-level result, rolling up every active ruleset in the project — this is what powers the project-wide Data Quality overview.
Issues & History
A quality issue is simply a rule whose latest cached result iswarning or failed — issues aren’t a separate record, they’re derived live from every active ruleset’s most recent run. The issues list supports the same filters=[field][$op]:value query syntax as other list endpoints, letting you scope to an entity, a project, or a specific ruleset.
For a single rule, its result history across past runs is available as a time-ordered list — this is what drives the sparkline shown in the rule detail panel.
Quality Score
The quality score is a severity-weighted pass rate:- Each rule’s severity contributes a weight:
critical=5,high=3,medium=2,low=1 - Each result contributes a factor:
passed=1,warning=0.5,failed=0 error,no_data, andnot_evaluatedresults are excluded entirely from the calculationscore = round(sum(weight × factor) / sum(weight) × 100), or0if nothing was scored
null for days that had no run, so a chart can show a real gap instead of a false zero. A live summary (current score, active/evaluated rule counts, coverage percent) is also available without waiting for history to catch up.
Authorization
Data Quality introduces two new project-scoped resource types alongside the ones described in Policy:
Both follow the standard project hierarchy (
project:<id>:ruleset:*, project:<id>:ruleset:<id>), so you can grant Data Quality access the same way you’d grant access to datasets or flows.
Where to Find It
Data Quality has two entry points in the app, both showing Overview / Rulesets / Issues tabs:- Project → Catalog → Data Quality — rulesets and issues across every dataset in the project
- Dataset → Data Quality — scoped to a single dataset, with “Run all” and “Add ruleset”
Related Resources
Policy
Grant access to rulesets and rule templates
Projects
How git-backed project resources are structured and synced
Flows
Another git-backed, YAML-defined project resource
Orion AI
How Datazone’s AI agents are grounded in your data