← Demos

Eval Lab

Watch a prompt change move real scores, then gate the release. Every answer is a live model call.

Live AI
1
EvaluateScore the baseline
2
ImproveMove the numbers
3
VerifyGate the release
Test runs deepeval · corpus: baltimore-policy · model: claude-haiku-4-5

Step 1 · Score the bot as it is today

About 20 seconds. Real model, real scoring, no mocks.

Scoreboard

MetricBaseΔ
Citation accuracy
Factual accuracy
Overall pass rate
Verification gate — not run

Source corpus (6 docs)

📄 Housing Code Ch. 116 — Vacant Buildings
📄 Sanitation Code §3-201 — Bulk Trash
📄 Council Minutes 2023 — Demo Funding
📄 Permit Office Guide — Construction
📄 Parking Enforcement Policy 2024
📄 Small Business Permit Quickstart
Go deeper

Does it know when it doesn't know?

Same question, same model, same documents. Move the clock forward and watch it lose confidence, then refuse to guess once the records go stale.

Probe question What is the city's vacant-building demolition budget?