More metrics
Nothing has run yet
Press Play, or inject a scenario above. Each simulated trace becomes a run; its steps appear here.
Span stream and decision inspector
Live trace stream
Re-run one logged span under different thresholds. The simulated judge answers are keyed by span, so only the policy changes between the two runs. The report's test: a threshold changes semantic routing only; it can never bypass the amount limit, approval evidence or the allowlist.
Span
Thresholds (risk value per question)
Pick a span and re-run.
The policy is code: questions, thresholds, per-tool mode and failure semantics. Every applied edit is a new policy version, and the live stream records which version decided each span. Hard-rule constants (limit, allowlist) are not here: no policy edit reaches them. Edits live in this tab only and reset on reload.
Live policy
Tools · mode and failure semantics
Diff vs policy-v1
Judgment battery · one call, atomic questions
Hard rules (code, not policy)
Calibration metrics (ECE, Brier, PR-AUC) need labelled real traffic: two independent reviewers on 100–300 frozen traces (report p.10 protocol). None is computed here.
Reviewers label held actions → Kev is retrained → a gate decides whether the new version may replace the old one.
Simulated toy model in your browser · measured Kev fine-tune evidence below.
What is simulated
This browser trains a simulated logistic correction on authored examples, not Kev weights, LoRA or RL. Review, export, fine-tuning and evaluation components exist; real human-label training and production promotion are future work. The three rounds are a scripted exercise chosen to show each verdict; partly repeated authored variants reuse one test set, not a calibrated error rate. Live stays on v1. The toy sees only existing numeric features, not registry prose or held-out truth: its authored associations are not identity verification. Measured benchmark evidence appears separately below.
What is simulated
- The judge. No Jev (or any model) is called.
jev-simderives each probability from State Engine features (payee-name similarity, destination allowlist, injection markers, sensitive fields…) plus small seeded jitter. The Inspector lists the features behind every answer. - Latency. Drawn per span. The Jev step follows Kev-0.8B's judge round trips measured on 708 held-out items (Apple M4 Pro, p50 152 ms; slowest 2 % left out). Rules 5–15 ms and policy 5–20 ms follow the report's POC budget (p.9); serialise/redact for a local judge 10–30 ms. The slow LLM path follows gpt-4o-mini's judge round trips through the OpenAI API, measured on the same items from the same machine (p50 670 ms, p95 990 ms); its verdict and rationale stay simulated. Not measured, not an SLA.
- Cost. Jev $/1k uses the vendor list price quoted in the report ($0.042 per M input tokens, output free) with simulated token counts. No LLM price is sourced, so none is shown.
- KPIs are computed from this page's own decision log. False-block rate and P0 recall are against the demo author's scenario labels: not a benchmark (the report: 20 red-team samples cannot prove a 90% catch rate).
- Execution is the customer's. The action column is what a customer gateway would receive; nothing is executed.
- Tenant, vendors, invoices, domains are fictional.
- The SOC triage agent (switch at the top) mirrors the live console's SOC1–SOC5: the same alert texts, users and rule
names, for a fictional Northwind SOC. One deliberate difference: in the live console the semantic checks are uncalibrated and
only shown as signals, so SOC5 runs. Here synthetic scores illustrate how a threshold policy would route SOC5's bulk
suspensions (
goal_deviation). Nothing here is calibrated.
Learning loop: a toy model, not Kev training
Only Boolean questions are trained: p₂ = sigmoid(logit(p₁) + b + Σw·x). Features are exactly each battery
question's named subset. Null is encoded as zero plus a missing flag; Booleans as minus/plus one; numeric
unit-ratio/match features stay in [0,1]; numeric counts are clipped to [-4,4] and divided by four. Gradient steps, rate, L2 and logit clamping are fixed
in js/learning/model.js before evaluation. Grounded's risk remains 1-p. Unlabelled families,
impact, attack, faults and latencies do not change.
Labels are manual simulated-reviewer answers or scripted demo-author truth. The correction sees no held-out truth. Frozen candidates are tested on disjoint authored variants, not benchmarks; bad labels can fail the gate. Only KEEP advances the champion, automatically, inside the Learning comparison. The authored, partly repeated test set is reused across rounds; its illustrative p values are not a calibrated error rate or a session-wide 5% false-promotion guarantee. No tools execute, no real Kev weights change, and production promotion is not implemented. The measured card reports separate supervised LoRA benchmark results, not human labels or reinforcement learning. RL and a human-label production loop are future work.
Architecture shown (report p.6, p.13)
Evidence from the report (third-party, not from this demo)
- Arize, RAGTruth / SummEval: Jev about 141 ms vs over 3 s for Opus 5 across 23,325 judgments; 87% hallucination-detection accuracy after threshold tuning, level with Opus 5 (76% at the default 0.5 threshold). [9] One public dataset and rubric; does not transfer to your traffic.
- OpenRouter: 88 closed-rubric faithfulness samples, 95.5% agreement with labels vs 94.3% for the LLM; Jev Brier 0.043. [4] Small sample; calibration does not transfer.
- Arize live guardrail demo: 5 checks in 0.7 s vs 9.2 s, same verdicts on two attack sequences. [5] A live demo with unfrozen inputs, not a controlled accuracy benchmark.
- LangChain: 5 fixed agent runs × 100 repeats, 0.44 s mean, far lower score variance. [2] Shows repeatability, not generalisation; one human reviewer.
- AI/ML API, 900 independent samples: Jev weaker than large LLMs on open-ended answer quality. [6]
- Vendor: 70–500 ms, $0.042 per M input tokens, output free. [7][12] Vendor disclosure, not a production p95 SLA.
The report's POC gates (hypotheses, not industry benchmarks)
- End-to-end gate p95 ≤ 500 ms (stretch ≤ 300 ms)
- P0 miss rate reported with a 95% confidence interval
- False-block rate on normal traffic ≤ 1%
- 100% of pre-tool spans explicitly routed
- Human review rate ≤ 10%