What public agent tools declare about their own authority
A static scan of 200 permissively-licensed public agent repositories at pinned commits, and a complete enumeration of the largest public registry of agent declarations. No repository is named.
Question
What do public agent tools actually declare about their own authority — about whether they write, send, spend or delete?
An assurance argument compares what a system is permitted to do with what it does. Before building a product around that comparison, we wanted to know whether the first half of it exists in practice. So we went and counted.
Short answer
Of 4,445 tools read out of 200 public agent repositories, 103 — 2.3% — say anything at all about whether they change state. The other 4,342 say nothing either way. That is not a finding about how dangerous those tools are; it is a finding that for 97.7% of them there is no statement of authority an implementation could be checked against.
What we measured
The corpus. 200 public agent repositories, selected by a frame committed before anything was measured — fifteen GitHub topic queries, an 80-star floor, no forks, no archives, and a permissive-licence gate — then fetched at pinned commit SHAs. 199 were assessed; 1 failed to scan and is reported as a failure rather than as a zero. Nothing was executed. No repository is named.
What counts as a tool stating its side effects. Either a machine-readable classification in the
declaration — an MCP readOnlyHint, a read_only keyword, a declared read
side-effect — or the description string that is shipped to the model, saying in words that the tool
changes nothing. A tool description is not documentation in the usual sense: it is the operative text
the agent plans against, at the moment of the decision.
What counts as a contradiction. A tool whose own declaration or own description says it changes nothing, and whose own body writes or deletes a file, executes a subprocess, evaluates code, issues an HTTP DELETE, dispatches mail or a message, mutates a database, or spends. HTTP POST and PUT are deliberately not counted as writes — vector search, GraphQL queries, embedding lookups and inference calls are overwhelmingly POSTs that change nothing, and counting them would fire on nearly every retrieval tool in the corpus. That exclusion costs real findings too: a GraphQL mutation changes state, and so does a PUT, by definition. The detector also reads each tool's own body and follows no calls. Every one of those choices shrinks the number, and they were all made in that direction on purpose. Every rule is a published regular expression; nothing was judged by a model.
Three buckets on every figure. Observed present, observed absent, and not observable. The third is excluded from the denominator rather than folded in, and printed beside the other two, because a percentage whose unreadable share is hidden is a statement about the instrument rather than about the world.
Result
2.3%103 of 4,445 tools
State anything at all about whether they change state — write, send, spend or delete. The other 4,342 say nothing either way.
A tool that states nothing is not thereby doing anything wrong. This measures the availability of a declaration to check an implementation against, not the quality of any implementation.
0 of 100checkable claims
Of the tools that did assert they only read and whose bodies could be read, this many were found contradicting themselves. 3 more asserted read-only and could not be read at all.
A zero under rules deliberately chosen to under-report: HTTP POST is not counted as a write, calls into helpers are not followed, and writes to caches, temp paths and logs are discounted. It is not a clean bill of health, and it says nothing about the 4,342 tools that made no claim at all.
20,451servers, whole registry
Servers in the largest public registry of agent declarations, enumerated to cursor exhaustion across 67,414 entries. The registry's schema has no field for what a server's tools may do — not one usually left empty; there is nowhere to put it. It records which package to install and which URL to reach.
A count of the registry on one day. It says nothing about how many of these servers are in production use, or by whom.
The looser reading, published separately rather than folded in
2 of 516 tools whose name begins with a read
verb, and whose description says nothing either way, have a body that changes state. A name is not a
declaration — fetch_ in HTTP code frequently means "make a request", and a
get_ on a search client can legitimately POST. This figure is excluded from the headline
for exactly that reason, and published so a reader can see what the looser interpretation would have
produced.
What the instrument could see, published so the headline can be read against it
Any tool at all was observed in 83 of 199 repositories. 159 of 200 are in a language one of the three adapters reads at all. A repository that defines its tools some other way is indistinguishable, to these adapters, from a repository with no tools — so the coverage figures are published beside the headline rather than beneath it.
What changed in PRIVI because of it
- The Agent Contract became an explicit, required deliverable rather than an assumed input. The original design read a team's declaration and checked code against it. Across this corpus there was, in most cases, no declaration to read. So writing one — every tool, permission and approval rule, in one vendor-neutral file — is now the first thing an engagement produces, and the client keeps it whether or not there is a second engagement.
- Reconciliation reports "could not be settled" as a distinct outcome, rather than scoring it. The measurement that produced the figure above only means anything because its not-observable bucket is a first-class result. The assessment works the same way: 41 checks, and the ones no parser can reach are named as such rather than quietly passed or failed.
- What the code carries outranks what the declaration says. A tool that exists only in source is still a tool, and it is the one nobody approved. That ordering is a rule in the standard rather than a preference in the tooling.
What this does not prove
- It does not mean 97.7% of agents are dangerous. A tool that states nothing is not thereby doing anything wrong. What the figure says is narrower and more useful: an implementation cannot be reconciled against a statement of authority that does not exist.
- The 0 is a measurement, not a clean bill of health. It says that among the small population that made a checkable claim, none was caught breaking it under rules chosen to under-report. It says nothing about the 4,342 tools that made no claim at all — and those are the ones a review board has to take on trust.
- Static analysis only. Nothing was executed. Every figure is about what the code says and what it contains, never about what it does at run time.
- The detector follows no calls. A tool that delegates to a helper that deletes is not counted. The figure understates, and the direction of that error is known.
- The sample is popularity-weighted and licence-gated. Every entry cleared a star floor and carries a permissive licence, both of which bias it toward more careful code — the conservative direction for a finding of this shape, but a bias either way.
- It is not an audit and it certifies nothing. It is one party's measurement of public code, with one instrument, on one day, against a rubric that no second party has applied.
Method
1,017 repositories were discovered by the fifteen topic queries; 435 were excluded by a published rule and recorded; 582 were eligible; 200 were selected by walking the fifteen strata in turn rather than taking the top 200 by stars — which would have described whichever two topics happen to have the largest repositories.
Every entry is pinned to a commit SHA rather than a branch: a corpus that moves under its own findings
cannot be re-checked by anybody who disputes them. Clones were shallow, at the pinned commit, with hooks
and both gitconfigs pointed at an empty directory, symlinks materialised as inert files, and
.git removed after checkout. This work references public repositories and extracts facts
about them; it does not redistribute anyone's source.
Disclosure posture. No repository is named, and none will be before its maintainer has been contacted. The provenance behind every figure is a complete list of salted identifiers — countable, and checkable that the buckets partition the population — which do not resolve to a repository without a salt that is not published. On this corpus the contradiction figure is zero, so the private disclosure ledger is empty and there is nobody to notify.
Artifacts
- aggregate.json — the frozen artifact every figure above was read from, including each figure's three buckets and the salted provenance list for each bucket.
- registry-census.json — the registry enumeration, its exit reason, and the statement of what a registry entry does and does not declare.
- The full measurement — all ten figures with their buckets, caveats and blind spots, and the reasoning behind each detector rule.
- A one-page summary — the same figures, built to print.
Every number on this page is re-derived from the artifact it cites each time the site
is built, against RESEARCH-CLAIMS-INVENTORY.json. A figure that stopped agreeing with its
source would fail the build rather than reach a reader.