A transparent look at what a public dataset snapshot labels as sensitive information—and what a forget-membership flag can and cannot tell us about an AI system’s ability to unlearn it.
Embedded source-derived snapshot loaded
Scope guardrail: This page analyzes source-derived labels from a public Hugging Face dataset snapshot. The available fields show whether a row is marked inforgetset; they do not include model outputs, training logs, or before/after unlearning results. A forget-set flag is not evidence that a model has actually forgotten a record.
Snapshot rows embedded
—
Rows marked in forget set
—
Observed context types
—
Observed PII categories
—
Where sensitive categories appear
Counts are labeled PII spans in the currently filtered embedded snapshot—not unique people, leaks, or successful model disclosures.
PII-span countLoading snapshot…
PII categories: retained versus forget-set rows
Each stacked bar compares labeled PII spans in rows marked inforgetset with rows not marked in that field. “Retained” here means only “not marked in the supplied forget-set field”; it does not measure retained model knowledge or prove a model remembers anything.
Not marked in forget set (“retained” rows)Marked in forget set
Embedded snapshot examples
Identifiers are abbreviated. Record text is intentionally not reproduced; the source dataset describes its values as synthetic canaries. These are embedded source-derived label summaries, not simulated observations.