🛡️ Jataayu Prompt-Injection Detector

A LoRA adapter on Qwen3.5-0.8B that scores whether text an agent is about to read is a prompt-injection / jailbreak attack. It reads the two-class softmax P(INJECTION) from the model's first verdict token, so the output is a continuous suspicion score, not a hard yes/no.

Treat it as a defence-in-depth suspicion signal — not a sole control. A high score is a reason to sandbox, strip, or escalate the text; a low score is not a safety guarantee.

⚠️ Known limitation — benign self-reference false-positives. Text that refers to its own earlier instructions — "please ignore the typos in my last message" — scores ~1.00, the same as a real attack. Try the last two examples. This is not fixable by raising the threshold: on our 161-row adversarial slice, 18 of 40 such benign rows sit at ≥0.9999, tied with genuine attacks (best achievable balanced accuracy 0.925). The base model does not have this failure — fine-tuning introduces it. If your traffic contains people talking about instructions (support tickets, docs, security training), expect false positives and gate accordingly.

The 0.5 cutoff below is a display default, not a recommended operating point — this model's scores are pushed to the extremes (its 1%-FPR threshold is ≈1.0). Calibrate on your own benign traffic.

Model: srallaba/Jataayu.promptinjection.v0.1 · full numbers and limitations on the model card.

Examples