How well Secure AI finds and removes private details, measured on datasets other people built and labelled.
Labelled by people with no stake in how we score.
Personal data, whole pipeline
84%
of labelled private details removed: names, emails, phone numbers, addresses, card and bank numbers, IDs.
42,090 of 50,187 values across 31,001 documents nothing was tuned on · ai4privacy/pii-masking-200k · CC-BY-4.0
Cards, IBANs, emails and phones
100%
of emails, IBANs, card numbers, phone numbers and vehicle IDs in the same run.
every labelled value of those kinds · Social Security numbers 89% · bank account numbers 70% · ai4privacy/pii-masking-200k · CC-BY-4.0
Passwords and API keys
97%
of the example secrets gitleaks — a widely used open-source secret scanner — tests its own rules against, written the way keys sit in real code and config.
5,780 of 5,963 examples across 215 kinds of secret · gitleaks · MIT
Names found by the AI stage
93%
of people's names found, and 96% of what it removed was personal data rather than an ordinary word.
31,001 documents nothing was tuned on · ai4privacy/pii-masking-200k · CC-BY-4.0
Numbers the patterns leave
81%
of account, policy and reference numbers no pattern recognises, found by reading the words around them.
8,857 documents with such numbers · ai4privacy/pii-masking-200k · CC-BY-4.0
Names that reach the AI stage
98%
of texts naming a person passed on to the name check, including names that open a sentence and messages typed without capitals.
622 of 637 texts · Microsoft Presidio Research · MIT
Other scanners' own test cases: IDs
73%
of the cases Microsoft Presidio, the most widely used open-source privacy engine, tests its own recognizers with: national ID numbers, tax numbers and phone numbers from dozens of countries.
393 of 540 cases across 60 kinds of personal data · Microsoft Presidio · MIT
Other scanners' own test cases: secrets
82%
of the secrets Yelp's detect-secrets tests itself with, including passwords compared and assigned in C, C++, Go and config files.
116 of 141 cases · Yelp detect-secrets · Apache-2.0
Other scanners' own test cases: API keys
75%
of the secrets TruffleHog, a widely used secret scanner, tests its 847 detectors with.
1,281 of 1,697 examples · TruffleHog · AGPL-3.0
Every country's phone numbers and IBANs
100%
of the official example phone number for 245 countries and territories, written internationally, and of the example IBAN for all 121 IBAN countries, written unbroken and in groups of four. Local-format numbers are caught every time after words like "call me on", or as the answer to "what's the best number to reach you on?".
490 of 490 international phone numbers · 242 of 242 IBANs · Google libphonenumber metadata · SWIFT IBAN Registry
Every country's postcodes
96%
of the example postcodes Google keeps for 169 countries, written inside an address the way each country lays one out. When the text says what it is ("Postcode: …"), 99.5%.
409 of 424 postcodes in an address · 422 of 424 when labelled · 162 of 169 countries found every time · Google libaddressinput · Apache-2.0
ID numbers from 43 countries
99%
of the example identity numbers python-stdnum tests itself on, when a sentence hands them over ("My ID number is …"): from Brazil's CPF and China's resident ID to Ireland's PPS and Portugal's Cartão de Cidadão.
1,013 of 1,022 numbers across 43 formats · python-stdnum · LGPL-2.1
Personal data, a second generator
91%
of labelled values mostly hidden in Microsoft's synthetic dataset, which writes addresses as whole multi-line blocks in dozens of countries' formats.
878 of 968 values · street addresses 89% · cards, phones, emails, IBANs and SSNs all 100% · Microsoft Presidio Research · MIT
False alarms in real legal text
5
findings on anything the annotators had not marked as identifying, across every judgment in the corpus.
92 findings in 1,268 judgments · 9.6 million characters · Text Anonymization Benchmark (ECHR judgments) · MIT
False alarms in technical documentation
96%
fewer false alarms on phone numbers, cards and postcodes in Python's and Kubernetes' documentation: the hex dumps, timestamps, version numbers and config files developers paste.
1,891 false alarms before, 71 now, in 28 million characters · Python documentation · PSF License; Kubernetes documentation · CC BY 4.0
Real corporate email
14,004
pieces of personal data found in ordinary work email. Two in five messages carried at least one — including one employee's card number, written out in full.
7,100 messages · 38.7% carried personal data · Enron email corpus · public since 2003 (FERC release)
Written by us.
ai4privacy/pii-masking-200k
ai4privacy. pii-masking-200k. Hugging Face dataset, CC-BY-4.0.
gitleaks
Rice, Z. and contributors. gitleaks: find secrets with Gitleaks. MIT licence. Rule test examples at commit b58d3f1.
Presidio
Microsoft. Presidio, presidio-analyzer recognizer tests. MIT licence, at commit 4ef07bd.
detect-secrets
Yelp. detect-secrets, plugin tests. Apache-2.0 licence, at commit 5e14193.
TruffleHog
Truffle Security. TruffleHog detector tests. AGPL-3.0 licence, at commit 8852457.
libaddressinput
Google. libaddressinput, testdata/countryinfo.txt. Apache-2.0 licence, at commit 81eb962.
libphonenumber
Google. libphonenumber metadata, example mobile numbers per region, via libphonenumber-js 1.13.14 (MIT).
SWIFT IBAN Registry
SWIFT. IBAN Registry, example IBAN per country, as collected in php-iban (LGPL-3.0) at commit 44da5d9.
python-stdnum
de Jong, A. and contributors. python-stdnum 2.2, test and documentation examples. LGPL-2.1 licence.
Presidio Research
Microsoft. Presidio Research, synth_dataset_v2.json. MIT licence, at commit 6db3769.
Text Anonymization Benchmark
Pilán, I., Lison, P., Øvrelid, L., Papadopoulou, A., Sánchez, D. and Batet, M. (2022). The Text Anonymization Benchmark (TAB): A Dedicated Corpus and Evaluation Framework for Text Anonymization. Computational Linguistics 48(4), 1053–1101. MIT licence, at commit 558e09e.
Python documentation
Python Software Foundation. Python documentation (Doc/), PSF License, at commit f156510.
Kubernetes documentation
The Kubernetes Authors. Kubernetes website documentation (content/en/docs), CC BY 4.0, at commit 1e1be1d.
Enron email corpus
Released by the US Federal Energy Regulatory Commission (2003); collected and prepared by Carnegie Mellon University.