mrkeyoor.com_
Mon 28 Sept 15:17 UTC
AI Toolsevaluationupdated 28 Sept 2026

rizzo-pii review

Rizzo PII is an Italian-first desktop app and local HTTP service for finding personal data, replacing it with typed placeholders, and restoring it after a cloud model responds. Its main README and setup guidance are available in English, while the detector concentrates on Italian legal identifiers such as codice fiscale, partita IVA, and cadastral references.

Verdict

Our Rizzo PII install consumed 5,741 MB, and pytest crashed before collecting a single test, so the clean audit and 1-second build do not establish a safe document pipeline. The Italian legal taxonomy is unusually useful, but use it only behind a fail-closed review step and keep the service on a trusted local boundary. Do not send anonymized payroll, clinical, or legal files to a cloud model until a second pass confirms that no identifier remains and no ordinary text was damaged.

We ran it

Lab card: what happened when we ran rizzo-piiScreenshot of rizzo-pii (rizzo-ai-academy.github.io/rizzo-pii)
Install✓ · 52s112 packages · 5741 MB
Build✓ · 1s
Tests✗ · 24s0 passed · 0 failed of 0 (pytest)
Known vulns0(pip-audit)
Repo176 files~10,983 lines of source · 44 MB · 3 CI workflows · Dockerfile · tests dir

Answers from our run

Does rizzo-pii build from source?

Dependencies installed in 52 seconds (112 packages), and the build succeeded in 1 seconds. We cloned commit 30d6f9c into a clean Debian container with 3 CPUs and no project-specific setup.

Do rizzo-pii's tests pass?

Yes: 0 of 0 passed when we ran the project's own test command (pytest). Some failures need services or credentials a bare container does not have.

Does rizzo-pii have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use rizzo-pii?

Organizations that need certified coverage outside Italian: the README says the other seven trained languages are not validated.

What are the alternatives to rizzo-pii?

Presidio, OpenMed, OpenPipe PII Redaction. Our Rizzo PII install consumed 5,741 MB, and pytest crashed before collecting a single test, so the clean audit and 1-second build do not establish a safe document pipeline.

Setup2/552-second install used 5,741 MB; pytest collected no tests
Docs4/5English setup, API, Docker, taxonomy, and limits are detailed
Community4/51,091 stars with active fixes and issue reports in September
Maturity2/5v2.0.0 ships apps, but privacy and PDF defects remain open

Who it’s for

Italian law firms, accountants, and document teams willing to verify every anonymized file before it leaves the machine.
Developers building a fail-closed privacy step in front of a cloud language model.
Desktop users who want packaged Windows, Apple Silicon macOS, or x86-64 Linux releases.
Researchers studying Italian PII detection who understand the stated sentence-level and synthetic-data limits.

Who it’s NOT for

Organizations that need certified coverage outside Italian: the README says the other seven trained languages are not validated.
Anyone treating a high benchmark score as permission to send output unchecked: open reports describe missed all-caps names, damaged clinical numbers, and PDF text removed outside detected entities.
Shared or LAN deployments without added access controls: issue 112 reports no authentication or Origin and Host validation on the local service.
Pipelines assuming irreversible mode removes the original from the API response: issue 118 reports that source_text still contains the clear text.
Small containers or casual source installs: our environment pulled 112 packages and occupied 5,741 MB before any document was processed.

Setup reality

Our sandbox installed 112 packages in 52 seconds and occupied 5,741 MB. The build succeeded in 1 second. Tests failed with exit code 1 after 24 seconds: pytest ran 0 tests, with 0 passed and 0 failed. The log ended in pytest's capture code with ValueError: I/O operation on closed file. Pip-audit found 0 known vulnerabilities.

End users can avoid the Python environment by downloading a packaged desktop release or building the provided CPU Docker image. Source use needs Python 3.11 or newer, PyTorch, Transformers, Flask, PyMuPDF, and a separately downloaded model. Normal inference is local and needs no hosted API key; dataset generation uses Gemini and Hugging Face credentials.

The 176-file repository includes 3 CI workflows, a Dockerfile, a compose file, and a tests directory. The container binds the service to localhost, keeps model access offline at runtime, and persists preferences in a volume. Long or sensitive document use still needs output checks, mapping protection, and attention to the open localhost-service report.

Rizzo PII covers 22 Italian legal and personal-data categories

Rizzo PII is built around an Italian-first 0.3B-parameter token classifier. Its 22 model labels include ordinary names, addresses, dates, email, and phone numbers, plus codice fiscale, partita IVA, cadastral references, document IDs, and Italian provinces. A separate URL rule becomes a 23rd app-level tag. The README and setup material are in English, but the product's strongest claim is narrow: Italian legal text, not universal document anonymization.

The app combines the model with regular expressions and checksum validators for structured values such as IBANs, tax codes, VAT numbers, and payment cards. It then replaces detected spans with stable placeholders such as [FULLNAME_1] and keeps a local mapping for restoration. That lets a cloud model reason over repeated entities without seeing their real values. The mapping is also the document's re-identification key, so it needs at least the same disk and access protection as the original.

The workflow is useful only when it fails closed

A sound deployment does more than call /analyze and forward the returned text. It checks the anonymized result again, stops when residual or skipped entities are reported, and keeps the original plus the reversible dictionary away from logs. Rizzo PII exposes a desktop interface, command-line path, and local HTTP service. The Docker example publishes 127.0.0.1:5005, runs with 1 worker, and keeps Hugging Face access offline after the image is built.

That local boundary matters because the files are exactly what an attacker wants. Open issue 112 reports that the Flask service has no authentication and no Origin or Host validation, and says its configuration route can rebind the service to 0.0.0.0. The issue is a static audit rather than our exploit test, but its recommendations are sensible: keep the port on loopback, do not expose it through a shared host, and add access controls before treating it as an internal service.

What happened when we ran it

Our sandbox installed 112 Python packages in 52 seconds and used 5,741 MB on disk. The build completed in 1 second. Pip-audit reported 0 known vulnerabilities. The checkout was 44 MB, with 176 files and about 10,983 source lines. Those figures make the source path much heavier than the repository size suggests, largely from the installed machine-learning stack, though the lab block does not break disk use down by package.

The test command failed with exit code 1 after 24 seconds. Pytest reported 0 passed and 0 failed because no tests ran. Its final stack trace was inside pytest's output-capture code and ended with ValueError: I/O operation on closed file; the summary said no tests ran in 19.05 seconds. The log does not show which project code or environment interaction closed that stream, so it would be wrong to assign a cause. The only safe conclusion is that our fresh Debian run produced no test evidence.

The repository still shows more delivery work than that failure alone suggests. Our scan found 3 CI workflow files, a tests directory, a Dockerfile, and a compose file. Ready-made v2.0.0 assets cover Windows, Apple Silicon macOS, and x86-64 Linux. Developers should reproduce the failing pytest startup before trusting a source deployment or changing the capture setup merely to obtain a green result.

A 0.989 sentence score does not prove whole-document safety

The README reports 0.989 micro-F1 on a 7,000-row held-out Italian set. It also states the limits plainly: validation is Italian-only, the Italian legal tags rely on generated entities inserted into real sentences, and evaluation is sentence-level rather than document-level. The project says a large real-document test set still needs to be assembled. That qualification matters more than another decimal place when the input is a 100-page contract or tabular payroll file.

Open reports show the gap in practical terms. Issue 94 says an all-caps person name remained visible in an Italian payslip. Issue 119 says clinical lab values were mistaken for PII and replaced, corrupting useful medical numbers. Issue 117 reports that PDF redaction removed ordinary text outside detected spans while response headers still described a clean result. These reports cover different paths, but all point to the same operating rule: check both leakage and document damage.

Irreversible mode can still return the clear source text

Issue 118 reports that /analyze returns source_text even when include_mapping: false, alongside an empty mapping and a flag saying mapping is disabled. A caller that logs or forwards the entire JSON response could therefore expose the original text while believing it had requested irreversible output. Until your deployed version proves otherwise, extract only the anonymized field, filter response logs, and add a regression containing synthetic names plus an IBAN.

Rizzo PII's reversible design also means anonymization and pseudonymization are different operational choices. If restoration is enabled, whoever has the mapping can recover every replaced value. If it is disabled, you still need to confirm that no clear source field survives elsewhere in the response, browser storage, previews, filenames, or logs. A privacy label cannot substitute for tracing the exact bytes that cross the machine boundary.

September fixes show activity, not closure

GitHub showed 1,091 stars, 80 forks, and 77 open issues and pull requests when fetched. The last push was September 25, 2026, and pull requests updated that day addressed PDF redaction boundaries, filenames containing document data, Windows drag and drop, and encrypted PDFs. The latest tagged app release remained v2.0.0 from August 8. Active fixes are encouraging, but a merged change is not the same as a tested release on your machines.

The Italian taxonomy is the reason to try Rizzo PII; the documented and reported edge cases are the reason not to trust it blindly. Start with synthetic copies of your actual document formats, including all-caps names, tables, scanned pages, and long identifiers split across lines. Compare the output text and rendered PDF with the original, then refuse the cloud call whenever either the leakage scan or the preservation check fails.

Alternatives

ProjectWhat it isPick it when
PresidioA configurable PII detection and anonymization framework for text, images, and structured data.pick this instead when language customization, recognizer composition, and service integration matter more than ready-made Italian legal coverage.
OpenMed gh↗A local clinical NER and de-identification project aimed at healthcare data across many languages.pick this instead when medical documents and multilingual clinical models are the primary workload.
OpenPipe PII RedactionA focused local detector and redactor for common PII classes.pick this instead when you need a smaller general-purpose redaction component rather than a desktop PDF workflow.

What people are saying

  1. [github-trending] Rizzo-AI-Academy/rizzo-pii

Sources

  1. Rizzo PII README
  2. Rizzo PII v2.0.0 release
  3. Issue 112: localhost service security report
  4. Issue 118: source text in irreversible response
  5. Issue 119: false positives in clinical reports
  6. Issue 117: PDF redaction removes ordinary text
  7. Issue 94: all-caps names missed in payslips

More ai tools reviews

redamon · SkillOpt · awesome-ai-agent-platforms · guizang-yingzao-skill · vdn-minimax-h3 · image-to-3d-pipeline · the whole board →