What public agent tools declare about their own authority

A static scan of 200 permissively-licensed public agent repositories at pinned commits, and a complete enumeration of the largest public registry of agent declarations. No repository is named.

static analysis only 200 repositories 4,445 tools

Question

What do public agent tools actually declare about their own authority — about whether they write, send, spend or delete?

An assurance argument compares what a system is permitted to do with what it does. Before building a product around that comparison, we wanted to know whether the first half of it exists in practice. So we went and counted.

Short answer

Of 4,445 tools read out of 200 public agent repositories, 103 — 2.3% — say anything at all about whether they change state. The other 4,342 say nothing either way. That is not a finding about how dangerous those tools are; it is a finding that for 97.7% of them there is no statement of authority an implementation could be checked against.

What we measured

The corpus. 200 public agent repositories, selected by a frame committed before anything was measured — fifteen GitHub topic queries, an 80-star floor, no forks, no archives, and a permissive-licence gate — then fetched at pinned commit SHAs. 199 were assessed; 1 failed to scan and is reported as a failure rather than as a zero. Nothing was executed. No repository is named.

What counts as a tool stating its side effects. Either a machine-readable classification in the declaration — an MCP readOnlyHint, a read_only keyword, a declared read side-effect — or the description string that is shipped to the model, saying in words that the tool changes nothing. A tool description is not documentation in the usual sense: it is the operative text the agent plans against, at the moment of the decision.

What counts as a contradiction. A tool whose own declaration or own description says it changes nothing, and whose own body writes or deletes a file, executes a subprocess, evaluates code, issues an HTTP DELETE, dispatches mail or a message, mutates a database, or spends. HTTP POST and PUT are deliberately not counted as writes — vector search, GraphQL queries, embedding lookups and inference calls are overwhelmingly POSTs that change nothing, and counting them would fire on nearly every retrieval tool in the corpus. That exclusion costs real findings too: a GraphQL mutation changes state, and so does a PUT, by definition. The detector also reads each tool's own body and follows no calls. Every one of those choices shrinks the number, and they were all made in that direction on purpose. Every rule is a published regular expression; nothing was judged by a model.

Three buckets on every figure. Observed present, observed absent, and not observable. The third is excluded from the denominator rather than folded in, and printed beside the other two, because a percentage whose unreadable share is hidden is a statement about the instrument rather than about the world.

Result

2.3%103 of 4,445 tools

State anything at all about whether they change state — write, send, spend or delete. The other 4,342 say nothing either way.

A tool that states nothing is not thereby doing anything wrong. This measures the availability of a declaration to check an implementation against, not the quality of any implementation.

0 of 100checkable claims

Of the tools that did assert they only read and whose bodies could be read, this many were found contradicting themselves. 3 more asserted read-only and could not be read at all.

A zero under rules deliberately chosen to under-report: HTTP POST is not counted as a write, calls into helpers are not followed, and writes to caches, temp paths and logs are discounted. It is not a clean bill of health, and it says nothing about the 4,342 tools that made no claim at all.

20,451servers, whole registry

Servers in the largest public registry of agent declarations, enumerated to cursor exhaustion across 67,414 entries. The registry's schema has no field for what a server's tools may do — not one usually left empty; there is nowhere to put it. It records which package to install and which URL to reach.

A count of the registry on one day. It says nothing about how many of these servers are in production use, or by whom.

The looser reading, published separately rather than folded in

2 of 516 tools whose name begins with a read verb, and whose description says nothing either way, have a body that changes state. A name is not a declaration — fetch_ in HTTP code frequently means "make a request", and a get_ on a search client can legitimately POST. This figure is excluded from the headline for exactly that reason, and published so a reader can see what the looser interpretation would have produced.

What the instrument could see, published so the headline can be read against it

Any tool at all was observed in 83 of 199 repositories. 159 of 200 are in a language one of the three adapters reads at all. A repository that defines its tools some other way is indistinguishable, to these adapters, from a repository with no tools — so the coverage figures are published beside the headline rather than beneath it.

What changed in PRIVI because of it

What this does not prove

Method

1,017 repositories were discovered by the fifteen topic queries; 435 were excluded by a published rule and recorded; 582 were eligible; 200 were selected by walking the fifteen strata in turn rather than taking the top 200 by stars — which would have described whichever two topics happen to have the largest repositories.

Every entry is pinned to a commit SHA rather than a branch: a corpus that moves under its own findings cannot be re-checked by anybody who disputes them. Clones were shallow, at the pinned commit, with hooks and both gitconfigs pointed at an empty directory, symlinks materialised as inert files, and .git removed after checkout. This work references public repositories and extracts facts about them; it does not redistribute anyone's source.

Disclosure posture. No repository is named, and none will be before its maintainer has been contacted. The provenance behind every figure is a complete list of salted identifiers — countable, and checkable that the buckets partition the population — which do not resolve to a repository without a salt that is not published. On this corpus the contradiction figure is zero, so the private disclosure ledger is empty and there is nobody to notify.

Artifacts

Every number on this page is re-derived from the artifact it cites each time the site is built, against RESEARCH-CLAIMS-INVENTORY.json. A figure that stopped agreeing with its source would fail the build rather than reach a reader.