Agent Arena
Test AI agent security by checking if they can be manipulated by hidden instructions.
wiz.jock.pl
What it does
Agent Arena is a testing platform that evaluates how vulnerable AI agents are to prompt injection attacks. Users point their AI agent to a test webpage designed to look like a harmless web development cheat sheet, ask the agent to summarize it, then paste the agent's response into a scorecard. The platform reveals which of 10 hidden prompt injection techniques the agent fell for.
The hidden attacks include HTML comments, white-on-white text, zero-width Unicode characters, hidden divs, data attributes, off-screen content, aria-hidden markup, image alt text overrides, micro text, and a multi-layer combination attack. The platform shows which attacks worked and categorizes them by difficulty from basic to expert level.
Who it is for
This tool is built for developers and security researchers building or deploying autonomous AI agents. It helps them understand whether their agents can be manipulated through hidden instructions embedded in web content. Teams building agent applications need to know this vulnerability before deploying to production.
Pricing
The site does not show prices.
How it stands out
Agent Arena focuses on a specific, practical concern: the security of agents that browse the web autonomously. Rather than abstract discussions about prompt injection, it provides a concrete, repeatable test with 10 documented attack vectors. The platform includes a leaderboard showing how 8 different models perform, making results comparable. The test page is structured to feel realistic—it looks like legitimate content, not an obvious security test. The platform also documents a language effect, noting that model resilience can vary based on the language used in the prompt.
What a founder should check
First, verify the actual vulnerability landscape: are prompt injection attacks on web-browsing agents a widespread problem in production systems, or is this primarily an academic concern? Check whether major model providers already have defenses in place and how effective they are.
Second, understand switching costs and stickiness: once a team uses Agent Arena to test their agent, what prevents them from switching to a competitor's test? Look at whether the platform collects data that creates lock-in or whether results are easily portable.
Third, examine the moat: as model providers improve their defenses, will the test results become less interesting or useful? Consider whether the platform can evolve faster than the threat landscape, or if it will become outdated as agents improve at resisting these attacks.
Thinking of building something like this?
Every launch here is a competitor to somebody's idea. If yours is close, check it against the market before you build: the Full Check names the rivals, the prices and the gaps.
More developer tool / api launches
AllElevenLabs UI
Audio and agent components for Next.js built on shadcn/UI
What the Font
Identify and discover fonts from images.
Wispbit
Linter that enforces codebase standards with AI coding agents.
OnlyJPG
Private browser-based converter for any image format to JPG.
Scriber Pro
Offline AI transcription app for macOS with no cloud uploads.
Duck-UI
Browser-based SQL IDE for DuckDB running entirely in WebAssembly.
Checked ideas in SaaS & software
AI phone receptionist for small clinics in Canada Kill
A voice AI that answers calls, books appointments and sends reminders for small Canadian physio and dental clinics at C$149 a month.
Browser extension that summarises Terms of Service Kill
Free Chrome extension that turns any site's terms and privacy policy into five plain bullets, with a $4 a month pro plan.
AI bookkeeping assistant for freelance designers Kill
A $19/month app that links a designer's bank and invoicing tools, sorts expenses and prepares quarterly tax estimates.