All commands · context

pm ops eval

Evaluate search relevance against a curated golden-query set: reports nDCG@k, MRR@k, precision@k, and recall@k per query plus the macro average. Use --fail-under as a CI gate.

Measures search relevance against a curated golden-query set so retrieval regressions are caught, not guessed.

Usage

pm ops eval [options]

Options

--mode <value>Default retrieval mode for queries without their own: keyword|semantic|hybrid (default: keyword)optional
--k <n>Metric cutoff @k (positive integer; default: 10)optional
--fail-under <value>Exit non-zero when aggregate nDCG@k falls below this threshold (0..1); CI gateoptional
--queries <path>Query JSON (https://schema.unbrained.dev/pm/eval-query-set/v1); default: search/eval-queries.json; errors show an exampleoptional
--format <value>Eval output format override: json|toonoptional

Examples

pm ops eval --json

Related commands