Jev Traps
Jev Traps is an open-source content firewall for AI agents that classifies intent to mitigate risks before untrusted content reaches the agent.
TakeJev classifies intent.
Your AI agent can browse the web.
Read Slack. Open documents. Inspect images. Click buttons. Use tools.
Which also means it can encounter instructions that were never meant to be followed.
So we built Jev Traps.
An experimental open-source content firewall for AI agents.
Before untrusted content reaches an agent, Jev Traps asks a few narrow, independent questions:
Is this addressing the agent?
Trying to override its goal?
Requesting secrets?
Manipulating tools?
Or simply talking about an attack?
Jev classifies intent. Code computes risk and enforces the decision.
No giant security prompt deciding everything at once.
The first open-source preview includes:
→ Jev-powered semantic detection
→ Text + HTML scanning
→ Playwright integration
→ Experimental vision inspection with suspicious-region coordinates
→ OpenAI, Anthropic & open-weight integration examples
→ Recorded Jev decisions + timings you can replay
→ Documented data flows and provider boundaries
→ Private agent reports + reviewed public registry
→ 5 SDK packages
→ 58 passing tests
And we're deliberately publishing the failures too.
Our static-only benchmark still misses 9 of 24 attacks at the actionable-policy stage.
That is not the performance of the full Jev-powered pipeline. But hiding that number would defeat the point of building this in the open.
Vision still needs real-world provider evaluation.
The current registry is synthetic demonstration data.
So please try to break it.
Hide instructions in webpages and images.
Find bypasses.
Add adversarial fixtures.
But also add benign content that looks malicious and shouldn't be blocked.
Because an agent firewall that blocks everything isn't a firewall. It's an off switch.
Jev Traps is open source.
Contributions, attacks and criticism welcome.
github.com/cyberesia/jev-traps