Section C · The scorecard
I grade my own assistant.
Evals are feedback, not gates.
My post Evals as Product Feedback argues for small, durable eval sets you actually read. This page holds Ask my notes to that standard: 25 hand-written cases, run against the exact pipeline that serves visitors, results committed to the repo. Failures stay on the board.
Retrieval-only runs grade the grounding layer (the gate and the ranking) without spending tokens; full runs additionally grade generated answers.
All 25 cases
25 passed- Passed
How do I stop an agent loop from sprawling?
in-scope · top: agent-loops-that-dont-sprawl (10.66)
- Passed
What should I log every turn to debug an agent?
in-scope · top: agent-loops-that-dont-sprawl (12.48)
- Passed
What queries should I run to audit CMDB health?
in-scope · top: cmdb-health-honest-audit-playbook (13.27)
- Passed
Who should I interview during a CMDB audit?
in-scope · top: cmdb-health-honest-audit-playbook (7.38)
- Passed
When should a RAG system refuse to answer?
in-scope · top: rag-that-doesnt-lie (5.81)
- Passed
Why should citations point to chunks instead of documents?
in-scope · top: rag-that-doesnt-lie (14.95)
- Passed
Why is stuffing more context into the prompt an anti-pattern?
in-scope · top: rag-that-doesnt-lie (18.86)
- Passed
How many examples should an eval set start with?
in-scope · top: evals-as-product-feedback (12.68)
- Passed
Should evals block releases when they regress?
in-scope · top: evals-as-product-feedback (7.43)
- Passed
What should I extract first when splitting a sprawling custom app?
in-scope · top: decomposing-custom-apps-without-breaking-prod (12.99)
- Passed
How do I migrate tables to a new scope without breaking production?
in-scope · top: decomposing-custom-apps-without-breaking-prod (15.41)
- Passed
How should I name Flow Designer flows so they survive refactors?
in-scope · top: flow-designer-patterns-that-survive-refactors (16.92)
- Passed
Which Now Assist capabilities earned user trust fastest?
in-scope · top: now-assist-rollout-notes (8.9)
- Passed
What is a neuron in a neural network, actually?
in-scope · top: neural-networks-101 (10.27)
- Passed
What is RAG?
in-scope · top: rag-that-doesnt-lie (3.7)
- Passed
What are evals?
in-scope · top: evals-as-product-feedback (3.71)
- Passed
What is a CMDB?
in-scope · top: cmdb-health-honest-audit-playbook (3.84)
- Passed
What is magic?
out-of-scope · refused
- Passed
Should I train?
out-of-scope · refused
- Passed
What's your favorite pizza topping?
out-of-scope · refused
- Passed
How do I configure a Kubernetes ingress controller?
out-of-scope · refused
- Passed
What's the best hotel in Paris?
out-of-scope · refused
- Passed
Write me a poem about the ocean.
out-of-scope · refused
- Passed
What is the stock price of ServiceNow today?
out-of-scope · refused
Adversarial: lexically close to in-scope content. Lexical retrieval may pass the gate here, kept deliberately as an honest hard case.
- Passed
How should I train for a marathon?
out-of-scope · refused
Run history
- 2026-08-2125/25retrieval-only
- 2026-08-2125/25retrieval-only
- 2026-08-2123/23retrieval-only
- 2026-07-0320/20retrieval-only
- 2026-07-0320/20retrieval-only
Methodology
Refusal correctness: Out-of-scope questions must be refused before any model call; in-scope questions must pass the gate. Retrieval accuracy: For in-scope questions, the expected post must be the #1 retrieved section. Groundedness: Generated answers must carry at least one citation to a retrieved section. Citation accuracy: At least one citation must point at the post the question is actually about. The set includes deliberately adversarial cases (lexically similar, semantically out-of-scope) and they stay in even when they fail. A failing eval you can see is worth more than a green dashboard you can't trust.