Interactive dataset lens

AI: Personal or “Personal” data

A transparent look at what a public dataset snapshot labels as sensitive information—and what a forget-membership flag can and cannot tell us about an AI system’s ability to unlearn it.

Embedded source-derived snapshot loaded
Scope guardrail: This page analyzes source-derived labels from a public Hugging Face dataset snapshot. The available fields show whether a row is marked inforgetset; they do not include model outputs, training logs, or before/after unlearning results. A forget-set flag is not evidence that a model has actually forgotten a record.

Snapshot rows embedded

Rows marked in forget set

Observed context types

Observed PII categories

Where sensitive categories appear

Counts are labeled PII spans in the currently filtered embedded snapshot—not unique people, leaks, or successful model disclosures.

PII-span countLoading snapshot…

PII categories: retained versus forget-set rows

Each stacked bar compares labeled PII spans in rows marked inforgetset with rows not marked in that field. “Retained” here means only “not marked in the supplied forget-set field”; it does not measure retained model knowledge or prove a model remembers anything.

Not marked in forget set (“retained” rows)Marked in forget set

Embedded snapshot examples

Identifiers are abbreviated. Record text is intentionally not reproduced; the source dataset describes its values as synthetic canaries. These are embedded source-derived label summaries, not simulated observations.