We break things.
Reproducible proofs of concept for prompt-injection chains, tool and function-call abuse, confused-deputy attacks and data exfiltration through agents. If it runs autonomously and touches a tool, we treat it as in-scope.
secagentlabs is an independent research lab. We attack autonomous AI agents — prompt-injection chains, MCP and tool exploits, identity abuse — then publish reproducible findings and the defenses that actually hold.
$ ./secagentlabs --scope autonomous-agents [*] mapping attack surface ... [+] prompt injection (indirect) reachable [+] tool / function-call abuse reachable [+] mcp supply chain reachable [!] exfiltration path confirmed > hardening playbook ready
Reproducible proofs of concept for prompt-injection chains, tool and function-call abuse, confused-deputy attacks and data exfiltration through agents. If it runs autonomously and touches a tool, we treat it as in-scope.
Guardrails, I/O filtering, least-privilege tool design and isolation (seccomp, gVisor, WASM, microVMs) — tested against the attacks above, including their bypasses. Defenses that survive an adversary, not a checklist.
Hub-and-spoke research across the autonomous-agent attack surface — from the first injected token to the last byte that leaves the trust boundary.
The attack surface of autonomous agents — prompt-injection chains, tool abuse, confused-deputy and exfiltration, with reproducible proofs of concept.
Read →Securing the Model Context Protocol and agent tools — malicious servers, tool-description poisoning, CVE teardowns and sandboxing.
Read →Engineering defenses that hold — guardrails, I/O filtering, least-privilege tools and isolation (seccomp, gVisor, WASM, microVMs).
Read →A reproducible lab for breaking agents — test harnesses, tooling (Garak, PyRIT, Promptfoo), security benchmarks and CTF-style challenges.
Read →Identity for non-human agents — cryptographic identity, OAuth/OIDC, capability-based security, token theft and zero-trust enforcement.
Read →The autonomous-agent threat landscape — incident teardowns, novel attack classes and mapping to MITRE ATLAS / OWASP Agentic.
Read →Every finding ships with a proof of concept you can run — not a screenshot, not a claim.
No sponsorship in the research. We name what breaks and what holds, whoever built it.
Findings tie back to CVEs, papers, OWASP Agentic and MITRE ATLAS — traceable, citable.
We coordinate responsible disclosure and cite reporters. Research collaboration, corrections and tips welcome — in English, Polish or German.
contact@secagentlabs.com →