Inspired by this research (and priors), and, @GoodfireAI catches reward hacks in the activations
Inspired by this research (and priors), and, @GoodfireAI catches reward hacks in the activations, but closed coding agents don’t give you those. So I built the other half: structural denies on graders/hidden tests + a @typesafeai Jev sidecar that scores tool trajectories as typed hazard Nouls. Opaque denials to the agent; scores stay with you. Video is an actual Exo run side by side with what rh-guard stopped b/c it judged it to be reward hacking. Has initial support for most agents (claude, codex, DSH, pi, etc) via agent hooks, but I'm targeting @AlexKrentsel et al's Exo harness b/c RSI is still cool even though Jev is all anyone is talking about anymore, but it doesn't have the exact hooks so still a 2nd class citizen. Fail-closed option. Repo in reply, annoying, but it's X....