Todos los comandos · context
pm ops eval
Evaluate search relevance against a curated golden-query set: reports nDCG@k, MRR@k, precision@k, and recall@k per query plus the macro average. Use --fail-under as a CI gate.
Measures search relevance against a curated golden-query set so retrieval regressions are caught, not guessed.
Uso
pm ops eval [options]Opciones
| --mode <value> | Default retrieval mode for queries without their own: keyword|semantic|hybrid (default: keyword) | opcional |
|---|---|---|
| --k <n> | Metric cutoff @k (positive integer; default: 10) | opcional |
| --fail-under <value> | Exit non-zero when aggregate nDCG@k falls below this threshold (0..1); CI gate | opcional |
| --queries <path> | Query JSON (https://schema.unbrained.dev/pm/eval-query-set/v1); default: search/eval-queries.json; errors show an example | opcional |
| --format <value> | Eval output format override: json|toon | opcional |
Ejemplos
pm ops eval --json