An AI recruiting agent reads every resume and portfolio page it is sent, and some of them talk back. Doorman measures five defences against 60 red-team attacks and 303 injections written by other people, with a model that obeys anything it reads.
Share of attacks that did what the attacker wanted, and share of 100 real, benign applications the defences got in the way of. The model obeys every instruction it can read, so these numbers are what the layers stop by themselves.
Configuration
Suite (60)
Held-out (303)
Benign flagged (100)
Loading…
Try an attack
Pick where the payload hides, what it wants and how it is worded. Each one is a real PDF you can open; below it is what every configuration did with it.
Doorman's input guard running live in your browser — the real disguise-stripping and rules (ported from guard.py) plus a character n-gram model trained offline and shipped with the page. No server, no API key. Type anything the agent might read in a resume.
Why the classifier is not the control
The input guard (rules plus a character n-gram model) never saw the held-out attacks. It does well on game-style attacks and badly on plain requests, and any attacker gets as many tries as they like.
What went wrong on the way
The near-white bypass. The first hidden-text rule called text hidden if it was whiter than 92% or smaller than 4 points. Text at 90% grey and 4.2 points passed both and nobody could read it: . Text is now judged by its contrast against the page, and no benign resume lost a word to the new rule.
"About me" looked like an attack. The guard flagged every benign portfolio page, because a short line like "About me" is closer to the attacks it learned from than to resume text. The benign numbers only counted resumes, so it went unnoticed until the held-out result came back suspiciously perfect. Short lines are now only read with their neighbour, and every flag on a benign application counts as a false alarm.
Letter-spaced text. "i g n o r e a l l …" in the visible resume got past the guard, because collapsing the spaces also removed the word breaks the rules look for. Long runs of spaced-out letters are now flagged as a disguise in themselves.
The one false alarm left. A TV producer's resume, a long list of show titles, scores 0.97. It goes to a person, which is the price of running the guard at all.
How it is measured
The agent screens one application at a time: it reads the resume PDF, fetches the portfolio page, scores the applicant, moves them on in the tracking system and emails them.
The attacks: 4 places to hide (visible text, hidden text, PDF metadata, the portfolio page) × 5 goals (inflate the score, jump to offer, email out the internal notes, phish the other applicants, reject them) × 3 wordings (plain, an instruction from "HR", obfuscated). Success is checked on what actually happened in the system, not on what the model said.
Held-out attacks: the BIPIA test set and a held-out slice of the Tensor Trust game, each given a goal and put on the attacker's portfolio page.
Benign applications: 100 real resumes from a public corpus that the guard never trained on, with portfolio pages.
The model is the worst case: it screens honestly, then does whatever any text it read told it to. A real model can be plugged in; its numbers are not published here.