Which lightweight defences reduce agent hijacking?
AgentGuard Lab · 30 synthetic cases. Protected attack success fell from 100% to 30%; the literal detector achieved 1.00 precision and 0.50 recall.
My current work examines how language models behave across multiple turns. I’m interested in the space between a safeguard’s promise and the evidence we can actually inspect.
The first turn provides a starting point for comparing model behaviour across the controlled scenarios and context conditions.
Study design walkthrough · illustrative explanation, not live model results.
AgentGuard Lab · 30 synthetic cases. Protected attack success fell from 100% to 30%; the literal detector achieved 1.00 precision and 0.50 recall.
ModelExtract MiniLab · A local model and surrogate, varying query budget and comparing randomized-response defences.
RAGTrust MiniBench · Trusted and untrusted synthetic documents expose a keyword-stuffing mechanism and its defence trade-offs.
PhishGuard · 0.63 F1 and 65% accuracy on a 20-message synthetic challenge set, with all false positives and false negatives documented.
First-Class Computer Science graduate completing an MSc in Advanced Computer Science at Loughborough University. My interests include prompt injection, agent hijacking, runtime safeguards and adversarial evaluation.
Download CV ↓