Skip to content
Apothem
Command pipeline

/eval

Reference for the /eval command: invocation, arguments, inputs and outputs.

Model-agnostic language-model evaluation campaign. Defines an evaluation dataset and scorer using the eval-harness skill, runs a candidate model or prompt over the dataset, scores every output with the prompt-evaluator agent, aggregates campaign metrics plus a per-category breakdown, and emits a report that surfaces regressions against the prior baseline. Operates against any model provider — no single vendor is assumed; the dataset, scorer, and candidate endpoint are all operator-supplied. Output lands at the consuming suite's _inputs/eval-findings.md with per-category scores, aggregate metrics, and a regression ledger ready for downstream release-readiness review.

Invocation

/eval [--dataset PATH] [--scorer NAME]

The definition sets disable-model-invocation: true, so a harness that honors that key starts this command only when you type it.

Pipeline position

Diagnostic surface for the model-quality review track. It consumes the evaluation dataset, the scorer definition, and the candidate endpoint, and emits the campaign findings artifact that release-readiness consumers read. The command mutates neither the candidate nor the dataset; the findings are read-only diagnostics.

Inputs

ArgumentTypeRequiredDescription
--dataset PATHPathNoPath to the evaluation dataset (JSONL / CSV / Parquet). MUST carry an input column, an expected column where the scorer is reference-based, and a category column for the per-category breakdown. When omitted, Phase 1 surfaces the dataset choice as an inquiry.
--scorer NAMEEnumNoScorer to apply — exact-match · semantic-similarity · rubric-graded · pairwise-preference · operator-defined. When omitted, Phase 1 surfaces the scorer choice as an inquiry with the candidate set annotated per rules/option-annotation.md.

Next step

Invoke /perf-audit to measure the candidate endpoint's runtime against the per-class budgets once the model-quality campaign is clean; the regression ledger at _inputs/eval-findings.md is the prerequisite evidence the release-readiness sign-off consumes.

Source

Generated from src/apothem/commands/eval.md, the command definition every harness installs.

On this page