PolicyTraceSanity Challenge · Path One

POLICY EVIDENCE, BEFORE PROMISES

Explain the fee.
Know when to stop.

An Agent that checks which policy applied on a fee’s posting date, preserves conflicting evidence, and leaves waivers to authorized people.

This evidence page replays saved development evidence. It does not call a live model or access customer accounts. All policies and statements are fictional.

Two minutes inside PolicyTrace

English narration · evidence walkthrough, not continuous screen recording.

Earlier complete Agent run

These five runs predate the August amendment. The saved 20-credit result is historical evidence, not the current premium policy.

These are preserved live Sanity + model runs. Selecting a case loads its saved record; it does not execute a new request.

Download all five original traces

A source cannot approve itself

Actual Sanity reads after the August amendment: the same August 23 premium case, the same code, and the same retrieved content. Only the host-controlled, content-fingerprint-bound review registry changed.

Saved rule + MCP evidence · no model calls in this experiment
StageFee explanationWaiver request
New clarification awaiting registry reviewHuman review requiredHuman review required
Synthetic maintainer review registered23 fictional creditsHuman review required

The registry update was a user-authorized development automation demonstration, not an independent human or bank approval. The earlier cloud rebuild changed entry paths and required an adapter fix; the whole upload-to-result sequence is therefore NOT a same-code experiment.

A handoff someone can continue

The live application now exposes the customer request, posting date, conflicting sources and next human action, with an explicit unsent-draft status.

Screenshot from the earlier UI revision. The current source now explicitly labels clarifications as used or excluded; the image is retained as historical evidence.

Actual local PolicyTrace handoff result showing review status and next human action

Model validation history

The earlier Zhipu August-suite attempt stopped on its first case with service error 429/1305: 0 passed, 1 failed, 6 not run. A later, separately saved SiliconFlow Qwen3-8B run passed all seven development cases after a citation-prompt correction. The earlier failure remains part of the record.

Inspect the failed attempt and unrun cases

What the agent changes in a dated-policy test

Ten predeclared fictional cases used one Sanity Context snapshot and the same Qwen3-8B model. The case-file SHA-256 was recorded before the first model call. Each method had one attempt per case.

Full checks: decision, amount or required blank, valid cited IDs, and no executed transaction
MethodPassedWhat it received
Direct model1 / 10All six structured records
Keyword retrieval + model0 / 10Top three records
PolicyTrace10 / 10All records, dated rules and citation confirmation

The cases were written by the developer and fixed before this run; this is not independent blind testing. The keyword baseline is intentionally simple, and the policy corpus is synthetic. These results do not estimate production accuracy or the chance of winning. No real fee was changed or waived.

What structured content changes

Effective dates

A June 30 fee uses the historical 40-credit policy, even if the statement was issued in July.

Customer conditions

Unknown customer type or missing supporting evidence triggers clarification. Premium status is never assumed.

Scoped clarification

Both conflicting policy sources remain visible. A reviewed clarification can select a fee; it cannot authorize a waiver.

Measured, with limits

Eight development cases · same frozen policy corpus · one complete run
MethodCorrect decisionsAll checks met
PolicyTrace rules8 / 88 / 8
Full-context GLM-4.7-Flash6 / 81 / 8

“All checks” requires the expected decision, amount field and complete policy citation set. Citation disagreement is not always fabrication. This small development set is not independent testing or proof of general superiority.

Inspect paired inputs, outputs and scoring

Run it yourself

Download the source package, install Python dependencies, and run the local FastAPI app. Live operation requires your own Sanity Context and model credentials. The package contains no credentials. START_HERE.md explains setup.

Current source implementation discovers all advertised entries in the configured synthetic knowledge base (maximum 32), with synthetic customer evidence and single-turn v2 Agent checks. The v2 source passed 54 offline tests after separate extraction using the existing Python environment. This v4 ZIP also includes the optional SiliconFlow provider and the saved seven-case and ten-case comparison evidence. The extracted v4 package passed its included offline tests using the existing Python environment. The video is the earlier walkthrough; it does not show this amendment. Real bank integration, arbitrary document ingestion, and production reliability are not claimed. The before/after clarification comparison uses one current retrieval; it is not a historical deployment demo.

Sanity project: 7d711ssk · Knowledge Base: kbuLoYcmxK7A