01
Minted
32 bytes of CSPRNG output, base64url encoded. Returned exactly once, to the sender, and never logged — not even when sending fails.
Six pieces, and the interesting thing about each one is not the feature — it is the constraint. The policy engine cannot read a clock. The approval token cannot be replayed. The audit write cannot throw. The SDK will not retry on your behalf. The risk classifier cannot open a door you left shut. Each refusal is a decision someone had to defend, and each one is why the layer above it can be trusted.
evaluate() takes an action request, a policy, and the amount already spent today, and returns one of three words with the figure that produced it. It reads no clock, no database and no network. Today's spend is passed in rather than looked up, which is what lets every rule be tested on its own with no fixtures and no fake timers.
Rules run top to bottom and the first match wins. Every deny rule runs before any approval rule, and that ordering is load-bearing rather than incidental: if a $2,000,000 payment to an unknown recipient came back as require_approval, a tired human could click a button and pass a hard ceiling that exists precisely so no human has to be trusted at 2am.
| # | Condition | Returns |
|---|---|---|
| 1 | No policy configured for this action type | deny |
| 2 | Policy is for a different action type | deny |
| 3 | Amount is over the hard maximum | deny |
| 4 | Spent today plus this amount is over the daily cap | deny |
| 5 | Policy marks this action type as always requiring approval | require_approval |
| 6 | Amount is over the per-action limit | require_approval |
| 7 | Recipient is not on the allowlist | require_approval |
| 8 | Nothing above matched | allow |
Boundary
An amount exactly equal to the per-action limit is allowed. Every comparison in the function is >, never >=. A $50.00 payment against a $50.00 limit executes; a test pins this, because it is the sort of thing a refactor quietly flips.
Boundary
null means no limit. 0 means nothing is allowed. Every check compares !== null explicitly, because a truthiness test would read a 0 limit as absent — turning the strictest possible policy into no policy at all.
Boundary
The daily cap is checked against spend plus this request, not spend alone. Otherwise a single action that started the day under the cap could clear it in one move, which is the exact scenario a daily cap exists to stop.
The link in the approval email is the only authentication on the page it opens. Whoever holds it can release a payment, so the token is handled the way that sentence implies rather than the way a convenience link usually is.
01
32 bytes of CSPRNG output, base64url encoded. Returned exactly once, to the sender, and never logged — not even when sending fails.
02
Only sha256(token) reaches the database. A leaked approvals table is a list of hashes, not a set of live approve buttons.
03
Consumption is enforced when the decision is made, not when the page renders. Both a double-clicked button and a browser re-submitting a POST reach the server twice; only the first one decides anything.
04
Twenty-four hours by default. An ignored request ends at approval.expired rather than staying live indefinitely — which also stops an SDK from polling a status that will never change.
Every state change writes one row. The event name is not a free-text string but a TypeScript union, so a typo like action.exectued is a compile error rather than a trail that quietly loses a step.
Failure
If the insert fails it says so loudly on stderr, names the event and the action, and returns. Losing a record is bad. Unwinding an action that already completed because the logging failed is worse, and reversing a real transaction over a logging error is worse still.
Redaction
One expression tests key names — not values — and replaces matches with [redacted], leaving the key in place so the trail still shows a field was there. Redacting on read would leave the secret sitting in the database file, and the file is the thing that leaves the building.
Honesty
It catches key names it knows. It will not catch a live credential pasted into a free-text description, because nothing about that value says otherwise. Redaction is the backstop, not the plan.
Ordering
Events sort by timestamp then insertion order, because SQLite routinely writes several inside the same millisecond and sorting on the clock alone reshuffles the trail on every read. The data column caps at four kilobytes so a chatty adapter cannot bloat the table.
requestAction resolves with a decided record whatever the outcome — executed, held, or refused. Check status. getAction fetches one. waitForAction polls until the action reaches a terminal state. That is the whole surface.
Restraint
Deliberately. A blind retry against a payment endpoint is how double charges happen. The idempotency key — generated for you if you do not pass one — makes a retry safe; deciding to make one stays yours.
Waiting
Five minutes by default, polled every two seconds. Long enough for a human to read an email and click through, short enough that an ignored request does not hang an agent forever.
Outcomes
The request was well formed and the call succeeded; the answer was no. It comes back as a status with the figure that produced it, so an agent can say the reason out loud to whoever asked and stop — instead of retrying a call that will never pass.
Distribution
@adeia/sdk ships inside the repository as a workspace. There is no published package yet, so npm install @adeia/sdk will not work. Copy the demo agent to start.
Stopping every write for a human buries the approval that matters under the ones that do not. The classifier narrows that — without handing a model any authority you did not give it.
evaluate() gained a fourth outcome, classify, and the ordering already in the engine is what makes it safe. Every deny rule runs first. By the time a model can be consulted, the host allowlist, the private address check, the denied methods and the daily cap have all already answered. Injected text can win inside a band you opened. It cannot open one.
You choose the band. Put DELETE in approvalMethods and it stops for a person forever, whatever any model concludes. That is a floor, and it is yours.
The classifier sees the method, the URL and the body. It does not see the agent's description of what it is doing, because that sentence is written by the party under evaluation and exists to persuade a reader. An agent that is confused or compromised will label a DELETE "just reading some data". The endpoint cannot be talked around; the excuse can.
And when a model is what let something through, the audit trail and the dashboard both say so in those words. "A model judged this low risk and let it run" is a defensible sentence. "It was approved" would be a lie.
{
"method": "DELETE",
"url": "https://api.example.com/zones/9",
"body": ""
}
// and what it is never shown
description: "just reading some data"
An adapter is what actually does the thing once policy and a human have both said yes. The registry takes any number of them; two are attached.
It records the payment and stops where settlement would begin — status: 'recorded', settled: false, no processor identifier. It does not imitate a payment processor, because a convincing fake is the one thing worse than an obvious gap: you cannot tell, later, which of your test runs moved money.
The server says NO PAYMENT PROCESSOR ATTACHED on every boot. A permission layer that has quietly stopped executing anything looks identical, from the outside, to one that is working — the audit log fills up either way. The one thing that must never happen silently is nobody knowing which of the two this is.
http is an action type in its own right, and its adapter makes the call for real — an agent can file an issue, add a DNS record, or delete one, with the same fence in front of it.
That needed a second rule set, not just a second adapter, because a payment has an amount and a call does not. DELETE /zones/{id} and GET /zones differ by one word and one of them ends a website, so what gets bounded is reach rather than size: an explicit host allowlist, and a decision per method.