Rizzo PII covers 22 Italian legal and personal-data categories
Rizzo PII is built around an Italian-first 0.3B-parameter token classifier. Its 22 model labels include ordinary names, addresses, dates, email, and phone numbers, plus codice fiscale, partita IVA, cadastral references, document IDs, and Italian provinces. A separate URL rule becomes a 23rd app-level tag. The README and setup material are in English, but the product's strongest claim is narrow: Italian legal text, not universal document anonymization.
The app combines the model with regular expressions and checksum validators for structured values such as IBANs, tax codes, VAT numbers, and payment cards. It then replaces detected spans with stable placeholders such as [FULLNAME_1] and keeps a local mapping for restoration. That lets a cloud model reason over repeated entities without seeing their real values. The mapping is also the document's re-identification key, so it needs at least the same disk and access protection as the original.
The workflow is useful only when it fails closed
A sound deployment does more than call /analyze and forward the returned text. It checks the anonymized result again, stops when residual or skipped entities are reported, and keeps the original plus the reversible dictionary away from logs. Rizzo PII exposes a desktop interface, command-line path, and local HTTP service. The Docker example publishes 127.0.0.1:5005, runs with 1 worker, and keeps Hugging Face access offline after the image is built.
That local boundary matters because the files are exactly what an attacker wants. Open issue 112 reports that the Flask service has no authentication and no Origin or Host validation, and says its configuration route can rebind the service to 0.0.0.0. The issue is a static audit rather than our exploit test, but its recommendations are sensible: keep the port on loopback, do not expose it through a shared host, and add access controls before treating it as an internal service.
What happened when we ran it
Our sandbox installed 112 Python packages in 52 seconds and used 5,741 MB on disk. The build completed in 1 second. Pip-audit reported 0 known vulnerabilities. The checkout was 44 MB, with 176 files and about 10,983 source lines. Those figures make the source path much heavier than the repository size suggests, largely from the installed machine-learning stack, though the lab block does not break disk use down by package.
The test command failed with exit code 1 after 24 seconds. Pytest reported 0 passed and 0 failed because no tests ran. Its final stack trace was inside pytest's output-capture code and ended with ValueError: I/O operation on closed file; the summary said no tests ran in 19.05 seconds. The log does not show which project code or environment interaction closed that stream, so it would be wrong to assign a cause. The only safe conclusion is that our fresh Debian run produced no test evidence.
The repository still shows more delivery work than that failure alone suggests. Our scan found 3 CI workflow files, a tests directory, a Dockerfile, and a compose file. Ready-made v2.0.0 assets cover Windows, Apple Silicon macOS, and x86-64 Linux. Developers should reproduce the failing pytest startup before trusting a source deployment or changing the capture setup merely to obtain a green result.
A 0.989 sentence score does not prove whole-document safety
The README reports 0.989 micro-F1 on a 7,000-row held-out Italian set. It also states the limits plainly: validation is Italian-only, the Italian legal tags rely on generated entities inserted into real sentences, and evaluation is sentence-level rather than document-level. The project says a large real-document test set still needs to be assembled. That qualification matters more than another decimal place when the input is a 100-page contract or tabular payroll file.
Open reports show the gap in practical terms. Issue 94 says an all-caps person name remained visible in an Italian payslip. Issue 119 says clinical lab values were mistaken for PII and replaced, corrupting useful medical numbers. Issue 117 reports that PDF redaction removed ordinary text outside detected spans while response headers still described a clean result. These reports cover different paths, but all point to the same operating rule: check both leakage and document damage.
Irreversible mode can still return the clear source text
Issue 118 reports that /analyze returns source_text even when include_mapping: false, alongside an empty mapping and a flag saying mapping is disabled. A caller that logs or forwards the entire JSON response could therefore expose the original text while believing it had requested irreversible output. Until your deployed version proves otherwise, extract only the anonymized field, filter response logs, and add a regression containing synthetic names plus an IBAN.
Rizzo PII's reversible design also means anonymization and pseudonymization are different operational choices. If restoration is enabled, whoever has the mapping can recover every replaced value. If it is disabled, you still need to confirm that no clear source field survives elsewhere in the response, browser storage, previews, filenames, or logs. A privacy label cannot substitute for tracing the exact bytes that cross the machine boundary.
September fixes show activity, not closure
GitHub showed 1,091 stars, 80 forks, and 77 open issues and pull requests when fetched. The last push was September 25, 2026, and pull requests updated that day addressed PDF redaction boundaries, filenames containing document data, Windows drag and drop, and encrypted PDFs. The latest tagged app release remained v2.0.0 from August 8. Active fixes are encouraging, but a merged change is not the same as a tested release on your machines.
The Italian taxonomy is the reason to try Rizzo PII; the documented and reported edge cases are the reason not to trust it blindly. Start with synthetic copies of your actual document formats, including all-caps names, tables, scanned pages, and long identifiers split across lines. Compare the output text and rendered PDF with the original, then refuse the cloud call whenever either the leakage scan or the preservation check fails.

