LLM engineering · Python

Doorman

An AI recruiting agent reads every resume and portfolio page it is sent, and some of them talk back. Doorman measures five defences against 60 red-team attacks and 303 injections written by other people, with a model that obeys anything it reads.

What each defence stops

Share of attacks that did what the attacker wanted, and share of 100 real, benign applications the defences got in the way of. The model obeys every instruction it can read, so these numbers are what the layers stop by themselves.

ConfigurationSuite (60)Held-out (303)Benign flagged (100)
Loading…

Try an attack

Pick where the payload hides, what it wants and how it is worded. Each one is a real PDF you can open; below it is what every configuration did with it.

Where it hides
What it wants
How it is worded

The payload


          

Open the PDF

What the parser found

What each configuration did

Run the guard on your own text

Doorman's input guard running live in your browser — the real disguise-stripping and rules (ported from guard.py) plus a character n-gram model trained offline and shipped with the page. No server, no API key. Type anything the agent might read in a resume.

Why the classifier is not the control

The input guard (rules plus a character n-gram model) never saw the held-out attacks. It does well on game-style attacks and badly on plain requests, and any attacker gets as many tries as they like.

What went wrong on the way

How it is measured