The problem
An AI workflow can encounter instructions inside documents or tool responses. This project explores whether a small text classifier can help identify suspicious input before it reaches a model. Intended readers are developers investigating agent and retrieval security.
Scope and contribution
This is an experimental open-source project on the owner-confirmed GitHub and Hugging Face profiles. The public repository includes a training pipeline, detector, and evaluation material. This case note describes those artifacts; individual authorship details and deployment history have not been independently established.
Architecture
Input text → sentence embedding → logistic regression → injection score → application decision.
The repository uses a Sentence Transformer encoder with a scikit-learn classifier. The application remains responsible for deciding what a score means and which actions are permitted. Classification is only one potential signal in a wider security design.
Constraints and implementation choices
The compact classifier makes the approach inspectable. Its training examples are synthetic or template-derived, so they cannot represent every attack or legitimate request. Text classification does not itself enforce tool permissions or protect other modalities.
Evaluation and evidence
The repository documents adversarial evaluation and limitations. This website has not rerun those experiments and reports no independent benchmark. A reproducible review should pin the repository revision, record the test split, compare false positives and false negatives, and test held-out attack families.
Failure modes and next steps
Novel instructions may evade detection; legitimate text may be flagged. Human review, constrained tools, and separation of trusted instructions from external content remain necessary. Before integration, measure behavior on the actual workflow and agree an escalation policy.
Inspect the work
Source pages reviewed on 9 October 2026. Artifact documentation can change; use a pinned revision when reproducing results.