Anonymization

When a program sends a request to a model, everything you put into it leaves together with the question: keys, passwords, document numbers, personal data. Secret hiding is a filter that replaces such values with conventional placeholders before sending, and puts them back in the reply. The model works with the placeholders, and you get a meaningful answer with the real data.

Enabled by default? No

Honestly, from the start: the feature is off by default. Enabled, it rewrites requests, and turning it on should be a conscious decision made after looking at the rules. The switch lives in the settings and applies within about 15 seconds, without a restart.

While the feature is off, nothing changes at all: requests go out as they are, and the product behaves exactly as if this capability did not exist.

How a rule is built

A rule consists of three parts:

  • Name — what you will call the rule, for example “Client personal data”.
  • The search pattern — a description of what to look for. It is a “pattern” in the technical sense: a special notation that can describe, say, “any eleven-digit number” or “the word after ‘number:’”. Composed once, works on its own afterwards.
  • The output pattern — how to mark what was found. The pattern holds exactly one $ sign, which the number is substituted into, for example CLIENT_$.

A rule also carries the list of the secrets themselves — the concrete values to hide.

What gets replaced, and what does not

The search pattern only finds candidate places, but the product replaces just the values listed in the rule’s secrets. That is deliberate, so nothing extra is touched: a similar-looking number or a different key stays as it is.

The practical conclusion: the secrets list has to be maintained by hand. If a secret is not on the list, it goes to the model as it is — even when the search pattern finds it.

Numbering and putting back

The numbering is continuous and permanent: the same value gets the same placeholder in all requests. If a secret first became CLIENT_1, it will also go out as CLIENT_1 the tenth time. That way the model does not get confused, and you understand the replies.

Putting back works in all replies, including the ones that arrive in parts. It keeps working even after the rule was changed or deleted: the product separately remembers every placeholder it has ever issued. Words that merely look similar, but are not secrets, are not spoiled — they pass through as they are.

The safe refusal

When a rule’s check cannot be completed, the request does not go to the vendor — you get an error. The logic is simple: better to get no answer than to send the secrets out. An overly “greedy” search pattern that takes longer than the allotted time to scan the request body behaves the same way.

Leak reports

Every replacement is recorded: the rule, the model, the key, the time, the placeholder and how many times the value occurred. This is the “Leak reports” section — it answers “what exactly are we hiding, and how often”. The secrets themselves are masked in the reports, so the history can be reviewed months later.

The default retention for the reports is 90 days, after which old records remove themselves. The value can be changed in the settings. When the same request arrives again, it is not reported a second time — the reports show no artificial “leak spike”.

The rule tester

Before putting a rule to work, you can check it on a live example: paste a fragment of text and see exactly what will be replaced. The numbers in the preview are their own, separate ones — they do not consume the real numbering. That is the cheapest way to make sure the search pattern does not grab too much.

Where rules and mappings live

The rules live in the anonymization.yml file in the configuration folder — together with the other settings files. The mappings between secrets and placeholders live separately, in the secret-hiding vault inside the working-state volume: the PROXYAGENT_ANONYMIZER_DB folder, by default /apps/proxy/anonymizer. If the working-state volume is not mounted, recreating the container loses the mappings: the old placeholders stop coming back, and the new numbering starts from scratch. Mounting the volume is described in the Hosting & Data section.

What the feature does not do

  • It does not touch the AI agent chat — the conversation in the chat goes past this filter.
  • It does not touch proxy service traffic — ordinary, non-AI traffic is filtered separately, if at all.

Next: how to route ordinary traffic through the same server — in the Tunnel Service section.

← Back to the documentation index