Mohamed AjabSoftware / security / curiosityLet’s talk
RESEARCH / ACTIVE QUESTIONSLoughborough University · MSc

Safety is a
conversation.
What changes
along the way?

My current work examines how language models behave across multiple turns. I’m interested in the space between a safeguard’s promise and the evidence we can actually inspect.

MULTI-TURN LLM SAFETY EVALUATIONIN PROGRESS / 2026

Establish the first response.

The first turn provides a starting point for comparing model behaviour across the controlled scenarios and context conditions.

Study design walkthrough · illustrative explanation, not live model results.

How I approach it

Keep the experiment
open to inspection.

Controlled comparison
Fixed model endpoints, explicit scenarios, context conditions and repeated runs make the comparison easier to follow.
Human annotation
Responses are coded against an explicit scheme, with TF-IDF and SVD supporting comparative text analysis.
Reproducible records
Structured calls, append-only evidence and offline validation preserve a trail from input to observation.
Clearly stated limits
Small synthetic studies are labelled as exploratory. A benchmark result is not a production security guarantee.

Related experiments

01

Which lightweight defences reduce agent hijacking?

AgentGuard Lab · 30 synthetic cases. Protected attack success fell from 100% to 30%; the literal detector achieved 1.00 precision and 0.50 recall.

Open project ↗
02

How much can an attacker learn through queries?

ModelExtract MiniLab · A local model and surrogate, varying query budget and comparing randomized-response defences.

Open project ↗
03

Can provenance help protect retrieval?

RAGTrust MiniBench · Trusted and untrusted synthetic documents expose a keyword-stuffing mechanism and its defence trade-offs.

Open project ↗
04

Can a compact classifier explain its mistakes?

PhishGuard · 0.63 F1 and 65% accuracy on a 20-message synthetic challenge set, with all false positives and false negatives documented.

Open project ↗

First-Class Computer Science graduate completing an MSc in Advanced Computer Science at Loughborough University. My interests include prompt injection, agent hijacking, runtime safeguards and adversarial evaluation.

Download CV ↓