Skip to main content
Workspace admin/knowledge-base/:kbId/evaluation

Running an evaluation

An evaluation scores a knowledge base against questions whose answer you already know, so a change in retrieval or settings shows up as a number rather than a hunch. You need a golden set before you can run one.

What you'll need
  • The Knowledge Base module visible in your left-hand navigation.
  • A knowledge base with indexed documents, and one golden set. Every section is visible to every user, so a refused read is where an under-entitled account finds out.

If Knowledge Base isn't in your navigation, your role doesn't have visibility for it — ask a workspace administrator, or see Permissions and module visibility.

Steps

  1. Open the knowledge base, select Evaluation, then Add golden set. Give it a Name and fill the Items list — each item needs a Question, and optionally an Expected answer, Expected documents and Tags. Use Add item for more. Items with no question are dropped on save.

  2. Select Run evaluation. It reads Starting…, then Evaluation run started. There is no picker — the run always uses the first golden set in the list, so put the set you care about at the top by creating it last.

  3. The run appears in the table under RUN, TRIGGER, HIT-RATE, GROUNDEDNESS, RESULT and WHEN. Select its Run button to open the Evaluation Run panel, which refreshes itself until the run is finished.

  4. The Metrics block gives Faithfulness, Answer relevancy, Context precision, Context recall, Groundedness and Retrieval hit-rate, each with its vs baseline movement, above a Regression detected or No regression pill. Per-item breakdown shows the same per question.

  5. Back on the list, the cards summarise the newest finished run: Retrieval hit-rate with its vs prev run delta, Answer correctness, Avg. groundedness and the size of your Golden set. The banner above them says whether the gate passed.

What success looks like

  • A run row carries a HIT-RATE and a GROUNDEDNESS figure, and a RESULT of Passed or Regression.
  • The Metrics block lists all six metrics instead of the not-completed notice.
  • The banner states the verdict, and the cards move with it.

Reading the verdict honestly

Gating is informational, never blocking — a flagged regression stops nothing. Answer correctness on the cards is the same number the panel calls Faithfulness, under a second name. A first run has nothing to compare against, so it shows no banner and a under RESULT.

If something goes wrong

SymptomLikely causeWhat to do
Run evaluation is greyed out.There is no golden set yet — the empty state says Add a golden set first, then run an evaluation.Select Add golden set and save at least one item with a Question.
No metrics yet — the run hasn’t completed.The run is still Queued or Running.Leave the panel open — it refreshes on its own. A failed run shows its error instead.
Save stays disabled on a golden set.The Name is empty, or no item has a Question.Fill both; other item fields are optional.
An expected document reads Not in this knowledge base.Only the first 200 documents are offered; that id is outside them or has been deleted.Remove it and pick a document from the list.
A search returns nothing you can see.The box filters the current page only, by trigger or run id.Clear it and page through the table instead.

Next

Steps verified against elie-tenant-ui 0.1.2 on . Something wrong with this page?