Skip to content

Independent assurance for AI agents

When your agent acts, you own what it did.

Courts have now said so twice, and since 2 August 2026 EU law requires you to keep a record good enough to reconstruct it. Krynix is where that record lives — independent of the agent, and of whoever built the guardrails.

Amazon v. Perplexity, No. 26-1444 (9th Cir., 4 Aug 2026) · EU AI Act, Art. 12

Example run · billing-agentrecording
  • 14:02:06Readconfig/prod.yml
  • 14:02:11Bashpsql -c "select count(*)…"
  • 14:02:40Bashcurl -X POST /v1/refund
  • 14:02:52Bashcurl -X POST /v1/refund
  • 14:03:05Bashcurl -X POST /v1/refund
  • 14:03:19Bashcurl -X POST /v1/refund
Watching controls…
passverified against evidencefailviolated, with the evidence attachedunknownnot computable — no evidence either waysealedsigned and in custody

What is at stake

This stopped being hypothetical in 2026.

Two court rulings, a regulator's €825m fine, a live trade-secret filing, and a law that took effect last month. Each links to its source.

  1. in force

    Art. 12

    record-keeping, in force

    EU AI Act, Article 12 — Record-keeping

    The law now requires the record

    The EU AI Act’s obligations for high-risk systems took effect. Article 12 requires automatic, tamper-evident logging over a system’s lifetime, kept at least six months, sufficient to reconstruct an individual AI-assisted decision after the fact.

    Could you reconstruct one decision your agent made last quarter, and show the record had not been edited since?

    Every run is hash-chained as it arrives and signed at close. Retention is a policy you set, deletions leave tombstones, and a legal hold outranks both.

  2. court holding

    No. 26-1444

    9th Cir.

    Amazon.com Services LLC v. Perplexity AI, Inc.

    The operator owns what the agent did

    The Ninth Circuit vacated Amazon’s injunction against Perplexity’s Comet agent, holding the agent is a tool, not a person — so it is the user, not the tool’s maker, who accesses a system. The first federal appellate ruling on whether AI agents may act on a user’s behalf.

    If your agent is legally your action, can you show what it did on your behalf?

    Krynix records the run as the operator’s own evidence, independent of the agent vendor and of whoever built the enforcement.

  3. under appeal

    €824,990,000

    Dutch DPA, under appeal

    Autoriteit Persoonsgegevens

    Automated decisions carry regulator-scale fines

    The Dutch DPA fined Uber €824,990,000 for making fully automated decisions about drivers — temporary and permanent deactivations triggered without human review. Uber has appealed.

    Which of your automated decisions could you evidence, one by one, if a regulator asked?

    Decisions are captured as characterised evidence and independently re-checked, so “what was decided, under which rule” is answerable per run rather than reconstructed from logs.

  4. alleged — disputed

    Apple trade-secret suit — reported filings

    The agent becomes the mechanism of exposure

    In a trade-secret suit, Apple alleges that confidential circuit schematics were run through simulation software, and that an AI agent had been taught to operate it. The evidence surfaced from forensics on a laptop produced in discovery. The allegations are disputed and untested, and the defendants have moved to dismiss.

    If confidential material passed through one of your agents, would you find out from your own records — or from someone else’s forensics?

    Tool calls, arguments and outputs are captured as they happen, so what an agent touched is a record you already hold rather than something reconstructed under subpoena.

  5. court holding

    2024 BCCRT 149

    BC Civil Resolution Tribunal

    Moffatt v. Air Canada, 2024 BCCRT 149

    Your agent’s words bind you

    A British Columbia tribunal held Air Canada liable for its chatbot’s incorrect statement about fares. The airline argued the chatbot was “a separate legal entity responsible for its own actions”. The tribunal rejected that outright.

    Your agent speaks and acts as you. Where is the record of what it said?

    One place that holds what each agent did and said, per run, retained on your clock rather than a vendor’s.

  6. And the records are on somebody else’s clock

    Analysis of the Perplexity ruling names the discovery problem directly: records of agent-assisted conduct — prompts, session histories, action traces — sit split between user devices and an outside operator’s systems, on that operator’s retention schedule, “held by a company on nobody’s custodian chart”.

    When a hold lands, can you preserve your agents’ evidence — or does it expire on a vendor’s schedule?

    Evidence lands in your own custody as it is produced. Retention is yours to set, and a legal hold stops the sweep whatever the policy says.

Who this is for

The person who has to answer for it.

Three seats end up in the room. They ask different questions, and only one of them is asking about telemetry.

Engineering & platform

“Which agents are running, what are they costing, and where do they break?”

Today

Traces in an observability tool built for services, where a run is a pile of spans and no one owns the question “did this agent behave”.

With Krynix

One view per fleet and per run: spend, tool failures, stalls, retry loops — and an explicit “unknown” wherever the data cannot answer.

Risk, compliance & audit

“Can we evidence what our automated systems decided, if we are asked?”

Today

A screenshot, an export, and an engineer’s word that the log is complete and unaltered.

With Krynix

A signed, hash-chained record per run that a third party can verify without us, with retention and legal hold you control.

Security

“What did our agents touch, and what governs them?”

Today

The agent’s own logs, produced by the thing being investigated, with tool arguments and outputs frequently missing.

With Krynix

Tool calls with their real arguments, and a run that tells you which calls arrived without them instead of leaving the gap for you to find — the plugins and hooks that could intercept them, and, wherever an enforcer produced claims, an independent re-check of those claims against the record.

What it answers

Six questions, six screens.

Each one is a real screen in the product, not a feature we named for this page.

  • Fleet health

    Is my fleet behaving?

    Every run, one view. Spend, failures, and what broke.

  • Agent inventory

    What is running that I do not govern?

    What you registered, next to what actually ran.

  • Controls

    Is it compliant?

    Your rules, checked mid-run — not after the fact.

  • Evidence & custody

    Can I prove it?

    Export a run. Anyone can verify it without us.

  • Findings

    Where did independent verification disagree with the record?

    We re-check the run ourselves. Disagreements surface here.

  • Policies

    What is an agent allowed to do?

    Every version, and who changed what.

Why independent

Anyone can say their agents behaved.

Evidence produced by the thing being investigated is the weakest kind there is. These are the three properties that make ours worth more than our word.

Nobody marks their own work
Your enforcement tool reports what it allowed. We re-check the same run and show you where the two disagree.
Agents cannot switch off the control failing them
The key your agent holds cannot change a control, a policy or a retention window — those take a person. It can still read your evidence; we are narrowing it to write-only.
The proof outlives us
Export a run and it verifies on a laptop that has never talked to Krynix.

Getting evidence in

Five environment variables.

We read OpenTelemetry, which your stack probably already speaks.

Monitoring a Claude Code sessionfive vars · no code change
export CLAUDE_CODE_ENABLE_TELEMETRY=1
export OTEL_LOGS_EXPORTER=otlp
export OTEL_EXPORTER_OTLP_PROTOCOL=http/json
export OTEL_EXPORTER_OTLP_LOGS_ENDPOINT=https://api.krynix.dev/v1/otlp/v1/logs
export OTEL_EXPORTER_OTLP_HEADERS="Authorization=Bearer kx_live_..."

That is the connection, and it needs no code change. It is not yet the whole picture: prompts, tool arguments and tool output are four further opt-ins, off by default, because each one is a decision about what leaves your machine rather than a setting we should make for you. We hand you the block for your agent and tell you what every line exposes.

About that key

Today it is a long-lived organisation key, pasted into your shell, and it can read your evidence as well as write to it. We are not comfortable with that either. It is being replaced by a browser sign-in that issues a short-lived, write-scoped credential, with no secret pasted anywhere.

Until that ships, mint a separate key per machine — they are revocable individually, so losing one costs you one.

What we understand today

  • Claude Code

    Cost, tool calls, prompts, plugins and hooks.

  • Codex CLI

    Tool decisions and results, with real output.

  • OpenTelemetry GenAI

    The convention, not one adapter per framework.

  • Microsoft AGT

    Audit entries, held in independent custody.

  • OpenAI · Anthropic · LangChain · LlamaIndex · Vercel AI

    One-line wrap() adapters.

  • Anything else over OTLP

    Stored and marked unread, so nothing gets a confident wrong answer.

Coming soon

What we are building next.

What we do not do yet, so you can find out here rather than three weeks into a pilot.

Framework control mappings

soon

Evidence mapped to SOC 2, EU AI Act and FINRA control IDs. Today your auditor does that step.

Tool output on every call

soon

We capture tool arguments today, and output for file reads. Output for the rest is in flight.

OpenTelemetry metrics

soon

We read traces and logs today. Metrics are next.

EU data residency

soon

Single US region today. Ask, and EU moves up the list.

Self-serve sign-up

soon

Connect an agent without talking to anyone. Guided for now.

Need one of these sooner?Tell us which

Early access

Put your fleet on the list.

What you run, and what you have to answer for. About two minutes.

We are onboarding design partners by hand rather than opening sign-up, so the answers decide the order. If you would rather ask questions first, talk to a person instead.

Bring one agent. See what we can already tell you about it.

Point one agent at us. We will tell you plainly what we can see and what we cannot.

How onboarding works

  1. 01We give you a key for that machine — one per machine, so you can revoke one without touching the rest — and you set five environment variables. No code change.
  2. 02Your next agent run appears, and we walk through it together.
  3. 03We name the gaps in what your setup reports before you commit to anything.

Guided, not self-serve — for now.