← Browse
eval-kit
TracineHQ ★ 0
eval-kit is a Claude Code skill and CLI for statistical significance testing that sizes evaluation runs before testing, applies Welch's or paired t-tests appropriately, and returns "can't tell yet" when data cannot support a conclusion. It provides zero-dependency JavaScript and scipy-backed Python implementations, both cross-validated in CI.
Install
> /plugin marketplace add TracineHQ/eval-kit
> /plugin install eval-kit
What it's made of
no extra components
- Commands
- 0
- Agents
- 0
- Skills
- 0
- MCP servers
- 0
- Hooks
- 0
What it needs & plugs into
- API keys
- none
- Paid services
- none detected
- External tools
- none
- Talks to
- nothing external detected
Analyzed . Facts extracted from the plugin's files. Prose generated by claude-haiku-4-5-20251001.
Is this your plugin?
Claim it to keep the card accurate and enter the weekly contest. Requires signing in as the GitHub owner (TracineHQ).