model-evaluation
This plugin provides comprehensive model evaluation tools including classification metrics with bootstrap confidence intervals, probability calibration analysis, sliced evaluation by demographic subgroup, fairness metrics, and LLM output quality assessment. It covers sklearn metrics (MCC, AUC-PR, AUC-ROC), calibration correction, fairness testing (Fairlearn, AI Fairness 360), and text evaluation (BERTScore, RAGAS) with statistical comparison between model versions.
Install
Source: https://github.com/HermeticOrmus/LibreMLOps-Claude-Code/tree/HEAD/plugins/model-evaluation
What it's made of
1 command · 1 agent · 1 skill
- Commands
- 1
- Agents
- 1
- Skills
- 1
- MCP servers
- 0
- Hooks
- 0
What it needs & plugs into
- API keys
- none
- Paid services
- none detected
- External tools
- none
- Talks to
- nothing external detected
Analyzed . Facts extracted from the plugin's files. Prose generated by claude-haiku-4-5-20251001.
Is this your plugin?
Claim it to keep the card accurate and enter the weekly contest. Requires signing in as the GitHub owner (HermeticOrmus).