Skip to content
adeia fence
Sign in
Proof

How you
know it works

The landing page shows an animation. An animation proves nothing. This page is the evidence underneath it, with the gaps named as plainly as the passing tests — because a page of evidence that hides what it has not tested is not evidence.

489 tests passing
27 test files
2 full runs on real infrastructure
0 payments settled
01 — The suite

489 tests, and where they sit

Run npm test in the repository and this is the output. The three heaviest files are the HTTP adapter and the two policy engines, which is correct — those are the parts where being wrong costs money or reaches something it should not.

Every test file and the number of tests it contains
File Tests What it covers
adapters/http 34 Re-resolution at call time, redirects refused, headers never returned
policy/evaluate 33 Rule order, inclusive limits, null versus zero, every decision path
policy/evaluateHttp 32 Host allowlist, per-method decisions, and that deny beats classify
routes/dashboard 26 Session gate, tenant isolation, key issue and rotation
routes/policyEdit 26 Editing a policy from the dashboard, and what may not be edited
approvals/routes 25 The approve and deny endpoints, replay, expiry, wrong tokens
actions/service 24 Request to decision to execution, and every terminal path
db/repo 24 Every query, idempotency, the daily spend calculation
audit/log 23 Event vocabulary, redaction, the size cap, ordering
routes/actions 20 The HTTP surface, auth, validation, error bodies
sdk/client 20 All three methods, timeouts, polling, error shapes
actions/classified 18 The classifier inside the action path, and every failure meaning ask
approvals/token 17 Minting, hashing, single use, expiry
policy/classify 17 Timeouts, malformed answers, the body cap, unrecognised verdicts
routes/auth 16 GitHub OAuth, single-use state, the cookie it is matched against
env 15 The boot refusals, including half-configured SMTP
routes/decide 15 Approving and denying from the dashboard, behind three gates
routes/audit 14 Reading a trail back, pagination, scoping to a project
demo/agent 13 A real LLM agent driving the tool, end to end
docs/generated 13 That the published schema matches the runtime behaviour
auth/session 12 Session tokens: hashed at rest, compared in constant time, absolute expiry
notify/email 12 Message construction, the approval link, escaping
shared/actions 11 The wire schema, strict mode, integer cents
notify/smtp 10 Transport construction and credential verification
policy/seedPolicy 8 The policy the seed script writes
adapters/ledger 6 That it records and does not settle
auth/apiKey 5 Generation, hashing, constant-time comparison
02 — The real run

Twice, against real infrastructure

A suite that only ever talks to fakes proves the fakes agree with themselves. The whole loop has been run twice against the real thing.

Real Gmail SMTP. A real email arriving on a real phone. A human reading the amount, the recipient and the rule that was crossed, then tapping approve. The agent, which had been blocked in waitForAction, picking up and reporting the result. A complete audit trail afterwards showing every step in order, including the ninety-five seconds where nothing happened because a person was deciding.

That gap in the timestamps is the only part of this that cannot be faked, and it is the whole claim the product makes: action.executing comes after approval.granted, never before.

the ninety-five seconds
14:22:07.114  action.requested
14:22:07.119  policy.evaluated   require_approval
14:22:07.402  approval.requested
14:22:07.988  email.sent

              ── a person is deciding ──

14:23:42.610  approval.granted
14:23:42.615  action.executing
14:23:42.701  action.executed
03 — Hard cases

The tests worth naming

Most of the 489 are ordinary. These six are the ones that catch the mistakes a future change would actually make.

Audit

No terminal status without its event

One test walks every action in the database and fails if any of them sits in a terminal status with no terminal event beside it. It exists to catch a transition somebody adds later and forgets to record — the failure mode that leaves an audit log quietly incomplete.

Audit

Redaction, checked in the stored bytes

Not that the function returns the right thing — that the row on disk contains the harmless field and does not contain the secret. Testing the function would pass even if the write path skipped it.

Policy

Deny always beats approve

The policy suite pins the rule ordering, so a refactor cannot reorder the checks and make a payment past the hard ceiling arrive as something a human is allowed to approve.

Classifier

Deny always beats classify

The same ordering, one layer along. A method you put on the deny list is never handed to a model, whatever band you opened elsewhere — so injected text can win inside a band, and can never open one.

Classifier

Every failure path means ask

Timeout, malformed answer, dead API key, a verdict nobody recognises — each one has its own case, and each one asserts the action went to a person. A missing key must never quietly widen what an agent may do alone.

Approvals

A token cannot be spent twice

Replay, expiry and unknown tokens all have their own cases, because a double-clicked approve button and a browser re-submitting a POST both reach the server twice in normal use.

04 — Unproven

What none of this shows

Read this section as carefully as the one with the numbers in it.

  • No money has moved Zero payments have settled, ever. The ledger adapter records and returns settled: false. Every test and both real runs stop exactly where a payment processor would begin, so nothing here is evidence about what happens when one is attached.
  • No load or concurrency testing The database is SQLite with a single writer, and nobody has driven it with concurrent traffic to find out where that breaks. The suite runs against a fresh database.
  • No external security review Tokens are hashed, compared in constant time, and single use, and that reasoning is written down — but nobody outside this project has tried to break it.
  • The classifier has not faced a real attacker Its authority is bounded by construction and that bound is tested. Whether a determined prompt injection can get a genuinely dangerous call judged low risk inside a band someone opened has not been tested by anyone trying hard.
  • No production users Nobody is running Adeia in front of real money. Two demos and a test suite are not the same thing as a system that has survived being used.