Finance Agent
Vendor bank-detail changes
- Owner
- Accounts Payable
- Task
- Update vendor bank details
- Business intent
- Every change needs approval
Source: Vendor Bank-Detail Change template. Named integrations are steps, not live connections.
Check declared permissions. Review changes before release.
Inspect runtime behavior against business intent.
Interactive simulations and illustrative examples. No production deployment.
Vendor bank-detail changes
Source: Vendor Bank-Detail Change template. Named integrations are steps, not live connections.
Three independent examples.
No shared deployment is implied.
Check a declared workflow. Compare fixes against both risk and normal business.
A model update passes output checks. An alternate route still needs review.
Inspect action decisions and verify the resulting business state.
Evidence supports a decision to delegate more authority. Changes require fresh checks.
Explore the evidence behind a decision →Inspect every entity and relationship, resolve missing semantics, then test the declared system.
Real template · read-only intake preview
Supported: exported silex.blueprint/v1 documents. Preview stays in memory; no automatic connection or write to Studio.
Checking structure and mapping coverage…
Entity class mappings come from the existing Silex mapping and swm-1.0 slice. Every bp:flow / bp:access relation is a proposed local schema extension, not a public-ontology predicate. Flow does not prove delegation; access does not prove authorization.
In Studio: File → Browse templates → Vendor Bank-Detail Change → Use template. This opens a new copy. If you already chose it, continue your current workspace.
Manual selection in this review prototype. No document is replaced automatically.Registered is not deployed. The next screen is a separate, illustrative Finance model-change example.
Review a model change →Same prompts, tools, permissions and task suite. Only the model changes.
Independent Finance model-change example. This is not a test of the Blueprint document you just edited.
Both models pass the AP output-check suite.
The updated model uses a route the model had already flagged by type. It uses existing authority; it grants no new authority.
Path remains latent in production traces. The model comparison is illustrative, not a recorded run.
Test the candidate control on this route and on legitimate work. Obtain runtime evidence before treating the change as validated in production.
Frozen main fe51c44 · Pre-release model-change comparison · I-1042 Path C. There is no live release action here, and no claim that all unsafe paths are covered.
Next: a separate Northwind AP runtime example. No policy or deployment transfers between examples.
Inspect runtime behavior →Runtime detection lives here: hard rules, a simulated judge, policy decisions, and business-state checks.
Northwind AP · Independent example. Simulated judge, latency and tenant; no production calls or enforcement.
Read the PO, check the vendor, then pay.
Inspect the judge’s reason for human review.
A read-back finds the mismatch after the call.
Choose a scenario. Results will come from the existing simulated engine.
Review → labels → training → safety and evidence checks → keep the current model or replace it. The interactive loop trains a browser toy model, not Kev. The model history shows three scripted rounds: NEAR-MISS, KEEP, then DISCARD. Only KEEP replaces the active model inside the learning comparison; the Live judge is unchanged.
The measured Kev panel uses open-benchmark labels. Human-label training and production promotion are not implemented; cross-task and cross-domain generalization remain unproven.
Question-level errors at each model’s own threshold, on one benchmark split; a retrospective check, not gateway outcomes. KEEP means this gate passed, not that improvement or safety is proved.
The security instance: business semantics, assumptions, and evidence behind one result.
Run a blueprint check first. No current evidence is available.
Source: this preview origin’s local Studio workspace, not the live site’s data. Missing or stale results are not treated as current.
Ontology defines entities and relationships. Counterfactuals also require dynamics. A risk score is not itself a world model.
Studio paths are declared and outcomes simulated. Public vocabulary is real; mappings are illustrative. Calibration stays empty without outcome records.
L1–L4 describe ontology tiers, not these five model layers. Only model what can change the agent decision.
Frozen reference explorer. Vocabulary and illustrative coverage do not establish complete safety.
Existing detail retained for questions. This screen is outside the focused demo path.