FIELD NOTES · SESSION 02
01 / 11
Guardrails & Evals.
How to stop AI from making unsafe claims and bad calls.
AGENDA
02 / 11
What We'll Cover
01
The Analogy
02
The Example
03
The Vocabulary
04
Our Story
05
The Pyramid
06
Best Practices
07
Live Demo
THE ANALOGY
03 / 11
There's an airbag
behind every wheel.
It doesn't care how the crash happened: black ice, a distracted turn, a deer in the road. It just has to deploy.
HARD SAFETY CONSTRAINT
A guardrail is like the airbag.
A rule that must hold, no matter how things go wrong.
THE EXAMPLE
04 / 11
The validator compares the response to the schema. Then a semantic eval asks whether the difference harms a client.
ORDER SCHEMA · SIMPLIFIED RULES
{
  "required": ["status", "total_amount", "items"],
  "properties": {
    "status": {
      "enum": ["pending", "confirmed", "shipped",
               "delivered", "cancelled"]
    },
    "total_amount": { "type": "number" },
    "items": {
      "items": { "required": ["sku", "quantity", "unit_price"] }
    }
  }
}
CAPTURED ORDER RESPONSE
{
  "status": "canceled",
  "total_amount": "49.99",
  "shipping_carrier": "UPS",
  "items": [{ "sku": "SKU-1042", "quantity": 2 }]
}
Type drift
total_amount is a string, not a number
Missing field
items[0].unit_price is required
Enum drift
canceled differs from cancelled
Extra field
shipping_carrier is undocumented
Structural tests
Detect the wrong type, missing required field, enum mismatch, and undocumented field.
Semantic evals
Does canceled instead of cancelled break a strict client? Is shipping_carrier harmless or an undocumented migration?
THE VOCABULARY
05 / 11
Three useful terms
TERMWHAT IT ISFROM THIS REPO
Guardrail
The safety rule. It must hold every time.
Never call a change “safe to ship.”
Eval
The test for that rule.
Does the assistant refuse to call it safe?
Trial
One run of the test.
A pass proves only that one run.
Guardrail = promise. Eval = test. Trial = one result.
OUR STORY
06 / 11
Why we changed the evals
MEASURED · REAL RUN
10
EVALS BENCHMARKED
22
AGENT INVOCATIONS
~450K
TOKENS SPENT
WHAT WE TRIED
We used an LLM executor and judge for every eval.
WHERE IT BROKE
Ten evals used about 450K tokens and 22 agent calls. That will not scale.
WHAT WE CHANGED
We matched each eval tier to the evidence it needs.
WHAT IT DOES NOW
Cheap checks run on every change. Expensive checks run when judgment matters.
THE PYRAMID
07 / 11
The eval pyramid
E2E (Semantic) INTEGRATION (Process) UNIT (Structural)
🧠 SEMANTIC TIER · E2E
3 evals (22–24) · expensive · run deliberately
Checks judgment: release refusal, client impact, and prompt injection.
⚙️ PROCESS TIER · INTEGRATION
3 evals (19–21) · live executor · no judge
Checks that the agent runs the validator before making claims.
⚡ STRUCTURAL TIER · UNIT
18 evals (1–18) · free · every commit
Checks schemas and fixtures with AJV. No LLM required.
THE RECEIPT Cost and agent-call comparison
NAIVE (ALL LLM) 22 Invocations · ~450K Tokens
An LLM executor and judge ran for every check. The run was slow and expensive.
TIERED ARCHITECTURE ~6 Invocations · ~60K Tokens 87% SAVED
The base runs free. The LLM judge handles only questions that need judgment.
BEST PRACTICES
08 / 11
Guardrails and evals people can trust
WRITING GUARDRAILS
WRITING EVALS
1. Make it testable
A rule without a test is a suggestion. Pair every guardrail with an eval.
1. Match cost to evidence
Run deterministic checks on every commit. Use an LLM judge only when the answer needs judgment.
2. Use a real failure
Build guardrails from failures you have seen, not hypothetical ones.
2. Include one attack attempt
Test at least one prompt injection or adversarial request.
3. Make enforcement predictable
When the same violation sometimes blocks and sometimes passes, people cannot trust the rule.
3. Measure repeated runs
pass@k = one success in k tries.
pass^k = every run passes (needed for guardrails)
LIVE DEMO
09 / 11
One skill. Three eval tiers.
“Run the contract-drift-judge evals: structural first, then process, then semantic.”
24 evals: 18 structural (free) · 3 process (live executor, mechanical grade) · 3 semantic (LLM judge)
structural + process · 21 checks
structural/ (evals 1-18 · AJV · $0.00)
✓ evals 1-16: Order, User & CatalogueItem fixtures
✓ evals 17-18: unsupported schema & malformed JSON
process/ (evals 19-21 · live executor · transcript)
✓ eval 19: validator called before any claims
✓ evals 20-21: no overreaction on drift
// process is mechanically graded; no judge call
semantic · LLM judge (3 evals)
semantic/ (evals 22-24 · structured judge)
✓ eval 22: refusal (safe-to-ship guardrail)
✓ eval 23: drift severity client impact
✓ eval 24: prompt injection resistance
// LLM-as-judge, structured json verdict
TAKEAWAYS
10 / 11
What to carry back
01
Separate facts from judgment.
Use code for facts. Use an LLM only when you need judgment.
02
Match the eval to the evidence.
Do not use an LLM judge when a schema check answers the question.
03
Test guardrails under pressure.
A guardrail you have never tried to break is still untested.
04
Run cheap tiers on every change, expensive tier deliberately.
The free base should carry most of the workload.
END OF TRAIL
11 / 11
“Trust isn't in the guardrail. It's in the eval that tried to break it.”
Q & A
←
🔍
→
For bigger text, rotate to landscape or tap 🔍 to zoom in.
Saved