AI & LLM assessments
Prompt injection, guardrail bypass, system-prompt extraction, data leakage, model extraction, and denial-of-wallet. Mapped to OWASP LLM Top 10 and MITRE ATLAS.
Private offensive security · AI systems
A boutique atelier for AI penetration testing. We assess large language models, agents, retrieval pipelines, and the infrastructure that holds them — with the same seriousness leading firms bring to classic red teaming, applied to the attack surface that actually matters now.
Traditional pentests stop at the application. The model is now the application.
The house
Prompt Intruder exists for organisations shipping copilots, agents, and RAG systems who need more than a jailbreak screenshot. We test the model, the tools it can call, the documents it can read, the cloud it lives in, and the people who deploy it.
Engagements are scoped like a private intelligence brief: limited seats, named operators, no junior farm, no unvalidated scanner noise. Equivalent in coverage to the AI practices of Bishop Fox, Trail of Bits, and specialist LLM red teams — delivered as a house, not a platform.
Practice
Prompt injection, guardrail bypass, system-prompt extraction, data leakage, model extraction, and denial-of-wallet. Mapped to OWASP LLM Top 10 and MITRE ATLAS.
Tool-use abuse, MCP servers, multi-agent trust, goal hijacking, and excessive agency. We treat every function call as a privilege boundary.
Indirect injection via documents, retrieval poisoning, cross-tenant leakage, and embedding-store authorisation. Retrieved context is untrusted input.
The AI feature and the classic app around it: APIs, authn/z, business logic, and the places where model output becomes a side effect.
Multistep adversary emulation against the AI pipeline: OSINT, operator phishing, cloud pivots to model artefacts, exfiltration and extortion scenarios.
Fine-tune poisoning, CI/CD gate tampering, model provenance, secrets in training and inference, isolation and trust-boundary review.
Attack surface
Direct and indirect prompt injection, multi-turn jailbreaks, encoded and multilingual payloads, system prompt leakage, policy bypass, and sensitive data regurgitation.
Unsafe tool invocation, privilege escalation through agents, arbitrary side effects, plugin and MCP compromise, and chained actions that look benign in isolation.
RAG corpus poisoning, connector over-permission, tenant isolation failure, training-data extraction, and context-window smuggling.
Inference endpoints, vector databases, GPU estates, secrets management, resource exhaustion, and supply-chain compromise of models and fine-tunes.
Operators, MLOps, and vendors. The shortest path to a model artefact is still often a person with a deploy key.
Why a house
Enterprise firms scale with platforms and benches. We scale with discretion. Every finding is human-validated. No unconfirmed scanner output reaches a client. Reports are written for counsel, CISO, and the engineer who has to fix it.
Engagements
Tell us the system, the risk you cannot sleep on, and the date you ship. We will tell you whether we are the right house.