Discover every sensitive column. Mask it irreversibly. Never touch the source.
Agentless discovery classifies PII and PHI across Postgres, MySQL, CSV, and JSON — then deterministic, plan-able masking runs against a clone, never your system of record. Audit-first and API-first, end to end.
One masking engine. 20+ pluggable strategies.
Four steps from raw data to a safe, realistic clone
No agents installed. No writes to your system of record. Every step is one HTTP call away — controlled re-identification included, and always audit-logged.
Connect & discover
Register Postgres, MySQL, CSV, or JSON sources — through an SSH tunnel if private. A classifier ensemble scans names and real data to flag every sensitive column.
Rule
Discovery becomes a versioned YAML rule set you review, edit, and commit — every revision kept, so you always know exactly what a run applied.
Run on a clone
Each run targets a structurally-isolated session clone of the source. The system of record has no write path from a run, by construction.
Audit & re-identify
Every run, audit, and re-identification attempt is recorded immutably. Controlled recovery of a single flagged row is recomputation-based and admin-gated.
Engineering-grade masking, not a checkbox feature
Deterministic, irreversible hashing
Every mask is a one-way function: SHA-512 of the value keyed by a per-run secret that is never persisted. The same value masks the same way across columns, tables, and runs — without any reversible mapping ever stored.
Plugin strategy architecture
Every strategy implements one contract — new ones plug in by registration, never by editing the engine. Seed files, composites, key renumbering: the format stays valid, relations stay intact.
Confidence-scored discovery
Column-name heuristics, data-pattern sampling, and sequential-ID detection fuse into a confidence score per column — with a suggested strategy ready to turn into rules.
Referential integrity preserved
Key renumbering records each mapping, foreign keys propagate in dependency order, and real constraints are handled safely — so masked data stays relationally coherent for the applications that consume it.
Durable background runs
Close the tab, redeploy the worker — runs keep going, checkpoint progress after every batch, and email you a download link the moment they finish.
API-first, always
Every button in the UI is one curl call away. Build masking into CI, not into a click-through ritual.
Built on three non-negotiables
01 · Never touch the source
Masking runs operate exclusively on a cloned session. The original connection has no write path from a run, ever.
02 · Irreversible by default
Standard masking uses an ephemeral, per-run secret that is never persisted in a reconstructable form.
03 · Every action is logged
Runs, audits, and re-identification attempts — including zero-match attempts — are recorded immutably, no exceptions.
Stop shipping real data to staging.
Connect your first database and see a generated rule set in under five minutes.