Exactly how HideReveal works
This page describes how the HideReveal desktop app works at a technical level — for people in IT/security roles who want to know exactly what happens to data before rolling the app out. If you're after step-by-step instructions instead, see the user guide. GDPR compliance is covered separately in the Privacy Policy.
Contents
- Architecture: local processing only
- Where the data dictionary is stored
- Token format
- How Excel anonymization works
- How Word / PowerPoint / TXT anonymization works — and its limitations
- How de-anonymization (reveal) works
- What HideReveal does not do
1. Architecture: local processing only
HideReveal is a regular desktop application (Python + Tkinter) that runs as a process on your own computer. The app makes no network connections as part of its operation — it does not upload files, the values extracted from them, or the token↔value dictionary to any HideReveal server, because no such server exists for this purpose. The entire anonymization and de-anonymization process happens in the app's memory and in the local database file described below.
This is different from hidereveal.eu, the website you're reading now — with your consent, this website uses Google Analytics 4 for visit statistics. The desktop app has nothing to do with that and contains no telemetry or analytics of its own. Details about the website are covered in the Privacy Policy.
2. Where the data dictionary is stored
The mapping of "original value → token" is kept in a local SQLite database (the file anonymizer.db), in the standard, OS-appropriate application data directory:
- Windows:
%APPDATA%\HideReveal\anonymizer.db(the "Roaming" folder of your user profile) - macOS:
~/Library/Application Support/HideReveal/anonymizer.db - Linux:
$XDG_DATA_HOME/HideReveal/anonymizer.db, or~/.local/share/HideReveal/anonymizer.dbif that variable isn't set
This is a single, shared location for every project created on that computer — each project is its own set of entries in this database, distinguished by its own ID. The database file is never synced or sent anywhere — it's a plain file on disk that you can protect the same way you would any other sensitive file on that machine (e.g. with disk encryption).
For "quick anonymize" (without setting up a persistent project), the mappings still go into this same database — the difference is that the ID of that ad-hoc project only lives in the app's memory for the duration of that run and isn't recorded anywhere outside the token database itself, so you can't "get back" to it from the interface after restarting the app — unless you still have an output file with that ID embedded in it (see below).
3. Token format
Every token has the shape SHORTCODE_TAG_RANDOM, e.g. A1B2C3_PESEL_XZ9KQP2M:
- SHORTCODE — a short hash of the project the value came from,
- TAG — a semantic tag code (e.g.
PESEL,EMAIL,KWOTA), independent of the interface language, - RANDOM — eight random characters (letters A-Z and digits), generated with Python's standard-library
secrets.choice— a source suitable for cryptographic use, not an ordinary, predictable random number generator.
A token is self-contained: it doesn't reveal the original value by itself, but carries enough context (what kind of data this is, which project it belongs to) to be human-readable and to be reliably decoded later, even if rows from different anonymized files happen to end up in the same spreadsheet or document.
4. How Excel anonymization works
This is the process that runs first in any given project (or quick-anonymize session), and it's what builds the dictionary described in section 2 — every value anonymized from the Excel file is added to it along with the token assigned to it. Only once that dictionary exists can related Word, PowerPoint, or text files be anonymized afterward (see section 5).
Once you pick which columns to anonymize, HideReveal reads every sheet in the workbook (not just the active one) and writes each one back under its original name — including sheets where none of the selected columns appear (those pass through unchanged). All columns are read as text rather than numbers, to avoid a common Excel pitfall: stripping a leading zero from a number-looking value (e.g. a national ID starting with "0").
The same value in the same column always gets the same token within one project — the app first pulls the unique values out of a column, translates each one once, and only then applies the substitution to the whole column.
A hidden _METADATA sheet with the project ID is added to the output file — this is purely informational (e.g. so the interface can show which project a file belongs to) and isn't required for de-anonymization itself (see section 6); it's dropped again when a de-anonymized file is written back out.
5. How Word / PowerPoint / TXT anonymization works — and its limitations
Word documents, PowerPoint slides and text files (.txt, .md, .csv) are anonymized using the same dictionary already built from anonymizing an Excel file in that project — meaning that literal occurrences of values already known from that dictionary get replaced, not any newly-detected sensitive data in the document itself. So there has to be something in the dictionary to look for before a document can be anonymized: anonymize the related Excel file in the same project first.
Matching happens on word boundaries and has a few deliberate simplifications:
- It is case-sensitive — "John Smith" and "john smith" are two different values as far as matching is concerned.
- It does not handle grammatical inflection — if the dictionary knows "Jan Kowalski", forms like "Kowalskiego" or "Kowalskim" (Polish declension) won't be recognized or replaced. This mostly affects Polish and German text; it's largely a non-issue in English.
- Longer values are matched before shorter ones (e.g. "Jan Kowalski" before plain "Jan"), to avoid replacing only part of a longer value.
On top of that, in Word and PowerPoint a paragraph's text can internally be split across several formatting fragments ("runs") — due to manual formatting, autocorrect, or how the text was pasted in. To avoid missing a match split across two such fragments, HideReveal joins the whole paragraph's text before searching for values to replace, then writes the result back into the first fragment of that paragraph. The cost of this simplification is losing any formatting that varied within a changed paragraph (e.g. if only one word in it was bold) — formatting of the paragraph as a whole, and of its first fragment, is preserved.
6. How de-anonymization (reveal) works
De-anonymization works differently than you might expect: HideReveal doesn't need to know which project a pasted text or loaded file came from. The app simply scans the content for anything that looks like a token (a pattern of uppercase letters, digits and underscores in at least three segments) and checks each candidate against the local database. Something that happens to look like a token but isn't one simply won't be found in the database and is left unchanged in the text.
That's why the "Paste & reveal" tab works on any text pasted from an AI chat window, without needing to know which project the data came from — and an Excel/Word/PowerPoint/TXT file can be de-anonymized even without access to the original project, as long as the tokens inside it still exist in that computer's local database.
7. What HideReveal does not do
- It does not send files, or values extracted from them, to any server — the whole process is local.
- It contains no telemetry or analytics of its own (unlike the website, see above).
- It does not judge whether a given column "definitely" contains personal data — the suggested tag is only a guess based on the column name; the decision of what to anonymize always stays with the person using the app.
- It does not replace a risk assessment or legal advice — it's a supporting tool, and responsibility for complying with applicable data protection law stays with the person using it (see the Privacy Policy).