English | 繁體中文
FastAPI wrapper for OpenAI Privacy Filter, with Docker, Docker Compose, and GitHub Container Registry publishing support.
This project turns OpenAI Privacy Filter into a small self-hosted service that exposes both a REST API and a browser-based web interface from a single app for PII detection and text redaction.
The web interface mirrors the official openai/privacy-filter Hugging Face Space: paste text, detect and highlight personal identifiers, and get a redacted output with label placeholders. A single page lets you choose where each request runs — on your device with WebGPU, on the server, or Auto — all backed by the same service and the same API.
OpenAI Privacy Filter is a strong local model for detecting and masking sensitive text such as names, emails, phone numbers, dates, addresses, account numbers, private URLs, and secrets. The upstream repo ships a Python package and CLI. This repo adds the missing deployment layer many teams want:
- REST API and web interface unified in one FastAPI service
- Optional on-device (WebGPU) inference with automatic server fallback, from the same page
- Local-first deployment
- Docker and Docker Compose support
- GitHub Actions workflow to publish container images
- Small integration surface for internal tools, RAG pipelines, ETL jobs, and document preprocessing
If you want to run OpenAI Privacy Filter as a backend service instead of calling the CLI directly, this repo is for you.
- Redact PII before sending text to an LLM
- Sanitize support tickets, chat logs, and transcripts
- Clean internal documents before indexing into RAG systems
- Build a privacy gateway for AI apps
- Add an on-prem redaction layer to compliance-sensitive workflows
GET /unified web interface (Auto / On-device WebGPU / Server modes)GET /webgpubackward-compatible redirect to/GET /healthhealth checkGET /configon-device inference configuration for the web interfacePOST /redactfull redaction response for a single text (spans + summary)POST /redact/texttext-only redaction responsePOST /redact/batchbatch redaction response with detected spans and latency- Configurable model device and checkpoint through environment variables
- Docker image build and Compose-based local startup
- pm2-managed deployment via
setup.sh/run.sh/stop.shandecosystem.config.js
After starting the service, open http://127.0.0.1:8080/ in your browser. The API and the web interface are the same service — no separate process or page.
The single page provides:
- A processing mode selector (Auto / On-device WebGPU / Server)
- A text input for content that may contain PII
- A "Detect & Redact" action that highlights detected entities
- A redacted text output with a copy button
- A per-label summary of detected entities
- Multilingual quick examples
The old
/webgpupage has been folded into/and now redirects there.
The three modes:
- Auto (default): runs in the browser with WebGPU if the browser supports
it and the device looks capable; otherwise it uses the server model. A
software/fallback WebGPU adapter scores
0, so weak machines automatically use the server even when WebGPU is technically "supported". - On-device (WebGPU): forces in-browser inference — text never leaves the machine. The first run downloads the model (cached afterward). If on-device processing isn't available (no WebGPU / disabled) or fails, your text is not sent anywhere; you're prompted and can explicitly choose to use the server.
- Server: always uses the higher-fidelity backend model — the same
openai/privacy-filtermodel exposed by the REST API.
Privacy note: automatic, silent server fallback happens only in Auto mode. In On-device mode the app never uploads your text without an explicit click.
On-device detection combines an in-browser NER model (names & locations, run with WebGPU via transformers.js) with local regex detectors for structured PII (email, phone, URL, date, account numbers, secrets). It is an approximation of the server model and may differ from its results.
How the engine is chosen and falls back:
flowchart TD
A[Submit text] --> B{Mode}
B -->|Server| S[Server /redact]
B -->|On-device| C{WebGPU supported?}
B -->|Auto| D{WebGPU supported<br/>and score >= min?}
C -->|yes| E[Run in browser]
C -->|no| K[Ask for consent first]
D -->|yes| E
D -->|no| S
E -->|error & forced On-device| K
E -->|error & Auto| S
E -->|ok| R[Render result]
K -->|user consents| S
S --> R
Relevant settings live in .env (exposed to the page via GET /config):
OPF_CLIENT_ENABLE: set tofalseto force every request to the serverOPF_CLIENT_MODEL: override the in-browser token-classification model idOPF_TRANSFORMERS_URL: override the transformers.js module URL (e.g. self-hosted)OPF_CLIENT_MIN_SCORE: device capability score (0-100) required for Auto to run on-device
Returns service status and whether the model is loaded.
Returns the full redaction result for a single text, including the original text, redacted text, detected spans, and a summary. This is the endpoint used by the web interface.
Request:
{
"text": "Email me at alice@example.com"
}Response:
{
"schema_version": 0,
"text": "Email me at alice@example.com",
"redacted_text": "Email me at [EMAIL]",
"detected_spans": [
{
"label": "private_email",
"start": 12,
"end": 29,
"text": "alice@example.com",
"placeholder": "[EMAIL]"
}
],
"summary": { "output_mode": "typed", "span_count": 1, "by_label": { "private_email": 1 }, "decoded_mismatch": false },
"warning": null,
"latency_ms": 123.45
}Request:
{
"text": "Alice lives at 1 Main Street and her email is [email protected]"
}Response:
{
"redacted_text": "[PRIVATE_PERSON] lives at [PRIVATE_ADDRESS] and her email is [PRIVATE_EMAIL]",
"latency_ms": 123.45
}Request:
{
"texts": [
"Alice was born on 1990-01-02.",
"Call Bob at +1 415 555 0114."
]
}Returns per-item redaction results, detected spans, summary metadata, and total latency.
The repo ships three scripts that wrap setup and process management. Deployment is managed by pm2, which keeps the service alive (auto-restart), centralizes logs, and can resurrect it on reboot.
./setup.sh # create .venv, install torch + privacy-filter + API deps + pm2, create .env
./run.sh # start (or reload) the web interface + API under pm2
./stop.sh # stop and remove the pm2 processsetup.shinstalls pm2 globally via npm (setSKIP_PM2=1to skip; requires Node.js). If a global install isn't possible,run.sh/stop.shfall back tonpx pm2.run.shusesecosystem.config.jsandpm2 startOrReload, so re-running it performs a zero-downtime reload. Pass--foreground(or-f) to bypass pm2 and run uvicorn directly in the foreground (useful for debugging).stop.shdeletes the pm2 process; pass--keep(or-k) to only stop it sopm2 restart openai-privacy-filterworks later.- Host and port are read from
.env(HOST,PORT), defaulting to0.0.0.0:8080.
Useful pm2 commands:
pm2 status # list processes
pm2 logs openai-privacy-filter # tail logs
pm2 restart openai-privacy-filter
pm2 startup && pm2 save # enable start-on-bootOnce running, open http://127.0.0.1:8080/ for the web UI or http://127.0.0.1:8080/docs for the API docs.
python -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
pip install ./privacy-filter
python main.pyThe API and web interface start on http://127.0.0.1:8080.
CPU image:
Build the image:
docker build -t openai-privacy-filter .Run the container:
docker run --rm -p 8080:8080 --env-file .env openai-privacy-filterStart the service:
docker compose up --buildRun in the background:
docker compose up --build -dStop it:
docker compose downCommon options:
PORT: API port, default8080OPF_DEVICE:cpu,cuda,mps, orauto; defaultcpuOPF_OUTPUT_MODE: OpenAI Privacy Filter output mode, defaulttypedOPF_CHECKPOINT: optional custom checkpoint path
Example .env:
PORT=8080
OPF_DEVICE=cpu
OPF_OUTPUT_MODE=typedThis is not the official OpenAI repo. It is a deployment-focused wrapper around the official OpenAI Privacy Filter project.
Upstream project:
If you need the core model, training flow, or evaluation tooling, start with the upstream repo. If you want to expose it as an API service quickly, use this repo.
OpenAI Privacy Filter API, OpenAI Privacy Filter FastAPI, PII redaction API, PII masking service, self-hosted privacy filter, Docker privacy filter, local PII detection, OpenAI privacy filter server, privacy filter for RAG, privacy filter for LLM preprocessing.
This wrapper repo does not change the upstream model license. Review the upstream project for model and code licensing details: