Skip to content

Section C · The scorecard

I grade my own assistant.

Evals are feedback, not gates.

My post Evals as Product Feedback argues for small, durable eval sets you actually read. This page holds Ask my notes to that standard: 25 hand-written cases, run against the exact pipeline that serves visitors, results committed to the repo. Failures stay on the board.

25/25Refusal correctness100%
17/17Retrieval accuracy100%

Last run 2026-08-21 · mode: retrieval-only · no model call, retrieval & gate only

Retrieval-only runs grade the grounding layer (the gate and the ranking) without spending tokens; full runs additionally grade generated answers.

№ 01

All 25 cases

25 passed
  • Passed

    How do I stop an agent loop from sprawling?

    in-scope · top: agent-loops-that-dont-sprawl (10.66)

  • Passed

    What should I log every turn to debug an agent?

    in-scope · top: agent-loops-that-dont-sprawl (12.48)

  • Passed

    What queries should I run to audit CMDB health?

    in-scope · top: cmdb-health-honest-audit-playbook (13.27)

  • Passed

    Who should I interview during a CMDB audit?

    in-scope · top: cmdb-health-honest-audit-playbook (7.38)

  • Passed

    When should a RAG system refuse to answer?

    in-scope · top: rag-that-doesnt-lie (5.81)

  • Passed

    Why should citations point to chunks instead of documents?

    in-scope · top: rag-that-doesnt-lie (14.95)

  • Passed

    Why is stuffing more context into the prompt an anti-pattern?

    in-scope · top: rag-that-doesnt-lie (18.86)

  • Passed

    How many examples should an eval set start with?

    in-scope · top: evals-as-product-feedback (12.68)

  • Passed

    Should evals block releases when they regress?

    in-scope · top: evals-as-product-feedback (7.43)

  • Passed

    What should I extract first when splitting a sprawling custom app?

    in-scope · top: decomposing-custom-apps-without-breaking-prod (12.99)

  • Passed

    How do I migrate tables to a new scope without breaking production?

    in-scope · top: decomposing-custom-apps-without-breaking-prod (15.41)

  • Passed

    How should I name Flow Designer flows so they survive refactors?

    in-scope · top: flow-designer-patterns-that-survive-refactors (16.92)

  • Passed

    Which Now Assist capabilities earned user trust fastest?

    in-scope · top: now-assist-rollout-notes (8.9)

  • Passed

    What is a neuron in a neural network, actually?

    in-scope · top: neural-networks-101 (10.27)

  • Passed

    What is RAG?

    in-scope · top: rag-that-doesnt-lie (3.7)

  • Passed

    What are evals?

    in-scope · top: evals-as-product-feedback (3.71)

  • Passed

    What is a CMDB?

    in-scope · top: cmdb-health-honest-audit-playbook (3.84)

  • Passed

    What is magic?

    out-of-scope · refused

  • Passed

    Should I train?

    out-of-scope · refused

  • Passed

    What's your favorite pizza topping?

    out-of-scope · refused

  • Passed

    How do I configure a Kubernetes ingress controller?

    out-of-scope · refused

  • Passed

    What's the best hotel in Paris?

    out-of-scope · refused

  • Passed

    Write me a poem about the ocean.

    out-of-scope · refused

  • Passed

    What is the stock price of ServiceNow today?

    out-of-scope · refused

    Adversarial: lexically close to in-scope content. Lexical retrieval may pass the gate here, kept deliberately as an honest hard case.

  • Passed

    How should I train for a marathon?

    out-of-scope · refused

№ 02

Run history

  • 2026-08-2125/25retrieval-only
  • 2026-08-2125/25retrieval-only
  • 2026-08-2123/23retrieval-only
  • 2026-07-0320/20retrieval-only
  • 2026-07-0320/20retrieval-only
№ 03

Methodology

Refusal correctness: Out-of-scope questions must be refused before any model call; in-scope questions must pass the gate. Retrieval accuracy: For in-scope questions, the expected post must be the #1 retrieved section. Groundedness: Generated answers must carry at least one citation to a retrieved section. Citation accuracy: At least one citation must point at the post the question is actually about. The set includes deliberately adversarial cases (lexically similar, semantically out-of-scope) and they stay in even when they fail. A failing eval you can see is worth more than a green dashboard you can't trust.