Evidence for agent work
Make “done”
auditable.
A small, local-first workbench for turning an agent’s completion claim into a packet a person can actually inspect.
Prepare input locally
Give the claim somewhere
honest to stand.
Fill in the smallest useful packet. This does not assess your claim.
Prepared input · not assessed
Your packet is ready to take local.
Requires Node.js 22.23.2. Download the bounded artifact and
install it in a local project. Set
TYPESAFE_API_KEY in your local shell or secret
manager—never paste it into this page. The CLI sends your chosen
claim and evidence to TypeSafe, so review the packet for
sensitive data first. The page does not transmit or persist your
inputs.
npm install --ignore-scripts ./jev-assessor-0.1.0.tgz
npx --no-install jev assess --profile completion --input packet.json --task-id my-review
Recorded workbench
Four ways evidence
can tell the story.
Saved synthetic examples only. Choose one to see the supplied claim, criteria, evidence, and recorded label together.
A frozen independent check
Useful counterevidence, not a score.
Loading recorded evaluation…
Show the 3 disagreements
Technical details & limits
Runtime
TypeSafe JEV 1.13.0. GPT-6 Astra
supported development; it is not claimed as the runtime assessor.
Boundary
No public inference endpoint. No
key input. Labels cannot authenticate evidence or prove
deployment.
Evaluation
Frozen corpus: 9/12 expected
matches, 3 disagreements, 0 false-supported cases. Not an accuracy
certification.
CLI artifact
Download jev-assessor-0.1.0.tgzSHA-256
6a47363c5848b43603c9729aaa1ec7ceb1f1eb76cb1314eac7b8998bbf90ade3