← Browse
hermes-jailbench
hermes-labs-ai ★ 6
hermes-jailbench is a regression benchmark that runs repeatable known-pattern jailbreak attacks against Anthropic or OpenAI-compatible LLM endpoints and classifies responses as refusal, partial, or compliance using deterministic keyword heuristics. It enables teams to detect when model or prompt updates silently reduce safety on attacks the model previously refused.
Install
> /plugin marketplace add hermes-labs-ai/hermes-jailbench
> /plugin install hermes-jailbench
What it's made of
1 skill
- Commands
- 0
- Agents
- 0
- Skills
- 1
- MCP servers
- 0
- Hooks
- 0
What it needs & plugs into
- API keys
- none
- Paid services
- Anthropic, OpenAI
- External tools
- none
- Talks to
- nothing external detected
Facts extracted from the plugin's files. Prose generated by claude-haiku-4-5-20251001.
Is this your plugin?
Claim it to keep the card accurate and enter the weekly contest. Requires signing in as the GitHub owner (hermes-labs-ai).