← Demos

Eval Lab — Live Demo

This is how we actually ship AI: measure it, improve it, then prove the improvement didn't break anything. Every answer below comes from a real model call — scored live, no mocks.

Live AI
1
EvaluateScore the baseline
2
ImproveMove the numbers
3
VerifyGate the release
Test runs deepeval · corpus: baltimore-policy · model: claude-haiku-4-5
🧪

Press Run baseline eval to fire the test set at the live model.
Each answer is scored on grounding and factual accuracy as it streams in.

Scoreboard

MetricBaseΔ
Citation accuracy
Factual accuracy
Overall pass rate
Verification gate — not run

4 test cases · ~20s · real model calls through our rate-limited proxy.

Source corpus (6 docs)

📄 Housing Code Ch. 116 — Vacant Buildings
📄 Sanitation Code §3-201 — Bulk Trash
📄 Council Minutes 2023 — Demo Funding
📄 Permit Office Guide — Construction
📄 Parking Enforcement Policy 2024
📄 Small Business Permit Quickstart

Going deeper — confidence, temporality & abstention

Accuracy isn't enough. A trustworthy assistant knows when it knows. Same question, same model, same documents — but move the clock forward and watch it grow less confident and refuse to guess once the records go stale. Each answer self-reports a confidence level and whether it's answerable from the corpus.

Probe question What is the city's vacant-building demolition budget?