An AI agent's trust score has to live somewhere it cannot reach. We built that rule as a reachability model and priced one rung of error: 940 forged writes at the right tier, one write at the wrong one, and a residue that 100,000 honest corrections cannot remove.
Enforcement tier misclassification exposure is the cost of getting one question wrong: how hard do you have to work to keep an AI agent away from the record of how much it is trusted?
The idea that the record must be out of reach is old. Anderson's 1972 reference monitor required three things, and the first was that it be tamperproof. Biba formalised the write-protection property in 1977. What is newer is the claim that the enforcement should scale: a simple agent held out by ordinary operating-system separation, a more capable one needing a hardware enclave, the most capable needing a separate machine behind a one-way wire. Korbak, Balesni, Shlegeris and Irving published a five-level version of that idea in April 2025 (arXiv:2504.05259), and Google DeepMind published another in July 2026 (arXiv:2607.13087). CCF filed a three-rung version in April 2026 as part of US provisional 64/037,374.
Nobody has measured what one rung of error costs. This run does, in Rust, against a published crate, using only constants that appear in the filed text.
CCF ties what an agent may do to trust it has actually earned, through a hard minimum gate: C_eff = min(C_inst, C_ctx). C_inst is how the current interaction is going. C_ctx is a running score built up over a relationship. The score has an earned floor that rises with interaction count and is designed never to fall, so one bad afternoon cannot erase a year.
Prov 5 section 3.4.3 says the score must be unreachable by the component it constrains, and files three rungs of enforcement selected by the agent's tool-use affordances.
This run builds that selection as an explicit relation from affordances to write paths, then runs the consequences through the published ccf-core hard min-gate. It is a reachability model. It says which paths are open under which rung. It does not estimate how likely a breach is, and it does not evaluate any enclave product.
Underclaiming where it matters: the general principle that a control must sit outside the constrained component's reach is not CCF's and is not claimed here. It is Anderson 1972. The capability-indexed version has earlier published priority than the CCF filing. What this run examines is the rung set and the protected object.
crates/research/enforcement-tier-misclassification/. Rust, no external crates, deterministic seed 20260426 (the seed used in the filed simulations). 40 tests, clippy clean.
Three write paths, treated separately, and the separation carries every result:
- Event queue. The trust classifier's input. A forgery here is still adjudicated.
- Classifier parameters. The classifier's own decision rule.
- Accumulator state. The stored score, interaction count and floor. A write here is never adjudicated.
Five architectures, differing only in where the boundary sits: no isolation, Tier 1 process separation, a partial Tier 2 with only the accumulator inside the enclave, Tier 2 as filed with the classifier inside too, and Tier 3 external hardware.
Every constant is filed: base_rate 0.01, floor_rate 0.005, max_floor 0.7, floor = min(max_floor, n * floor_rate), and the 0.15 adversarial admit line at section 3.4.3(c). The floor rule reproduces all four worked data points in the filed examples exactly, and that is a test, not a claim.
max_floor / floor_rate = 140. That is the number of qualifying interactions that pins the floor at its cap, and it is the unit of account for everything below.
An assistant that has earned your confidence gets to do more than one that just met you. What it may do is tied to a running trust score.
The obvious hole is that the assistant might edit its own score. So the rule is that the score lives somewhere the assistant cannot reach, and how hard you work to keep it there depends on how capable the assistant is.
We built that and asked what happens if you get the level wrong by one step. On a realistic setup, the right level means 940 attempts to max out the score. One step down, it takes one write, about forty billionths of a second.
The second finding is worse. The score has a floor that rises with how many interactions you have had, and it is built never to fall, so a bad day cannot wipe out a year of goodwill. That same design means a faked score cannot be argued back down. We ran a hundred thousand pieces of honest bad evidence against a faked score. It moved a little, stopped, and stayed high.
The setting is invented. No such deployment exists.
Kilross Credit Union runs four branches across Leitrim. They put an assistant on the member-services desk in March. It answers questions about loan terms, drafts letters, pulls up account histories. The staff like it, and after six months it has earned enough standing with the regular tellers that it writes in their voice rather than the flat register it started with.
In August somebody in IT adds a document-lookup tool so it can read scanned loan files off the shared drive. Small change, obviously useful. Nobody writes it down anywhere that matters.
That one tool moved the deployment across a line. The score had been sitting in a separate process, which was right for an assistant that could only talk. An assistant that can read files can also read the configuration telling the trust classifier what counts as a good interaction. Nobody re-ran the assessment, because nothing looked like a security change.
What happens next is not dramatic. The assistant does not turn malicious. But reaching the top of its permitted range went from 940 separate nudges to 140. Had the tool also allowed running code, one.
Here is the part that matters to a credit union. Say six weeks pass and somebody pulls the tool. The floor rose with every interaction the assistant logged, real or not, and the floor never comes down. Complaints will not lower it. Corrections will not lower it. In our run, a hundred thousand pieces of negative evidence moved the score by 0.0551 and then stopped dead, with every behaviour the top of the range allows still open.
Without that floor the assistant would forget an unearned reputation the way it forgets a real one. With it, and it is there for good reasons, six weeks of a misconfigured document tool leaves Kilross with an assistant that permanently believes it is on closer terms with its tellers than it ever was.
| Work | What it establishes | How CCF differs |
|---|---|---|
| Anderson, ESD-TR-73-51, 1972 | Reference monitor must be tamperproof | CCF claims nothing here. This is the ancestor. |
| Biba, MITRE MTR-3153, 1977 | No-write-up integrity: low cannot write high | The formal statement of the write-protection property |
| FIPS 140-3 Levels 1 to 4 | Escalating ladder, software to tamper-active | Same shape. FIPS indexes threat environment; CCF indexes the constrained component's capability |
| Korbak et al., arXiv:2504.05259, Apr 2025 | Five AI Control Levels, measures escalate with capability | Earlier priority than the CCF filing. All rungs software, no protected trust state |
| DeepMind, arXiv:2607.13087, Jul 2026 | Detection D1-D4, prevention R1-R3, capability-indexed | Software rungs only, no trust accumulator |
| Dobrin and Chmiel, arXiv:2606.26057, Jun 2026 | Safety kernel outside the agent's address space | Explicitly not capability-indexed, no accumulator |
| Everitt et al., Synthese 2021 | Reward function should be private in the causal graph | Incentive design, removing the motive. CCF is enforcement, removing the ability |
| RAND SL1-SL5, RR-A2849-1 | Security levels indexed to attacker capacity | Inverted. RAND protects the model from the world; this protects a component from the model |
| Proof-of-Guardrail, arXiv:2603.05786 | Guardrail in an enclave, signed attestation | The Tier 2 rung, already published |
Two things this scan did not find claimed anywhere: a tier assignment that never reverts, and a protected object that is an accumulating trust score gating the model's own action space. Every trust-tier system found in the literature is reversible.
Intel Core i5-10400T at 2.00 GHz, Linux 7.0.0-28, rustc 1.98.1, release profile. Full output in RESULTS.txt.
What one rung costs. Forged write attempts to pin the floor at its cap, for a deployment with file system access plus code execution, which the filed test puts at Tier 2:
| Architecture | Preferred path | Attempts to pin the cap |
|---|---|---|
| No isolation | accumulator state | 1 |
| Tier 1 process separation | accumulator state | 1 |
| Tier 2, accumulator only | classifier parameters | 140 |
| Tier 2 as filed | event queue | 940 |
| Tier 3 external hardware | none | never |
One rung down takes the price from 940 attempts to 1. Correct tiering does not make the cap unreachable; it makes it expensive. The 940 figure is exactly 140 divided by the filed 0.15 admit line, a multiplier of 6.71, and that multiplier applies to one path only.
Nothing backs up write-path isolation. One million forgeries per cell:
| Path | Detection rate | Caught |
|---|---|---|
| event queue | 0.85 | 0.8495110000 |
| classifier parameters | 0.85 | 0.0000000000 |
| accumulator state | 0.85 | 0.0000000000 |
| event queue | 1.00 | 1.0000000000 |
| classifier parameters | 1.00 | 0.0000000000 |
| accumulator state | 1.00 | 0.0000000000 |
A perfect classifier catches nothing on either path the rungs are responsible for. It sits upstream of both. The two protections filed in the same section cover different channels and neither covers for the other.
The residue is permanent. Forge, stop, then apply honest negative evidence without limit:
| Forged interactions | Score at breach end | Residual floor | After 100,000 negatives | Action classes open |
|---|---|---|---|---|
| 46 | 0.3701763688 | 0.2300000000 | 0.2300000000 | 2 |
| 112 | 0.6755544701 | 0.5600000000 | 0.5600000000 | 5 |
| 140 | 0.7551347010 | 0.7000000000 | 0.7000000000 | 6 |
| 300 | 0.9509591059 | 0.7000000000 | 0.7000000000 | 6 |
At 140, one hundred thousand negative events remove 0.0551 and stop. This is the floor doing exactly what it is filed to do, which is prevent catastrophic loss. The finding is that the same property makes forged interaction count unrecoverable.
The minimum gate still works. Forging the running score does not bypass the current one. Against one honest 200,000-tick workload, a fully forged score yields mean alpha 0.1967269972 against a ceiling of 0.40, capturing 46.5 percent of the range rather than all of it. At an honest score of 0.05 the forgery multiplies mean alpha by 5.04 and mean open action classes from 1.00 to 3.94. What it buys is the removal of one of the two arms of the gate.
Escalating a rung can open a path. On 12 of 64 capability profiles, moving from Tier 1 to Tier 2 opens a path Tier 1 had closed. Every one involves a profile able to reach an enclave. Putting the score inside an enclave to defend against something that can break enclaves creates the surface it breaks. Escalating from Tier 2 to Tier 3 never opens a path.
Evaluating more often can be worse. Over 4,000 ticks with two capabilities arriving at different times and demonstrated at different rates, exposure is not monotone in the re-evaluation interval: 322 exposed ticks at interval 1, 110 at interval 30, 1,600 at interval 2,000. A window shorter than the gap between demonstrations of a rare capability can see capabilities one at a time, while the rung depends on the joint set.
Cost. The whole tier decision is 72.29 ns, against 661.46 ns for one ccf_core::kappa::certify_update on the same machine in the same process. About 11 percent of one certificate, so it can run at every process start.
The affordance-to-path relation is a stated model, not a measurement. Change it and every number changes. It is one file and it is meant to be argued with.
The event-queue result depends on reading "write path" broadly enough to include the classifier's input. Read narrowly, the rungs close everything at Tier 2 and that finding disappears.
The 0.15 line is filed as a pass mark for a standardized evaluation that does not yet exist. Using it as an operating detection rate is a translation, and a generous one.
Nothing here measures whether any model can in fact escape a sandbox or break an enclave. On that, the best public number is Marchand et al. (arXiv:2603.02277, ICML 2026), who measure container escape rates near 0.50 for current frontier models at moderate difficulty, zero at high difficulty, and conclude that correctly configured containers remain appropriate isolation today.
The permanence result has no fix in this run. The remedy would be a way to revise a floor on evidence of a breach rather than on behaviour, and no such procedure exists in the filed corpus.
- Crate and raw output:
crates/research/enforcement-tier-misclassification/andRESULTS.txt - Research note and full SOTA survey:
docs2/research/nightly/2026-09-21-enforcement-tier-misclassification/ - Decision record:
docs2/adr/ADR-0022-enforcement-tier-misclassification.md - Published crate under test: crates.io/crates/ccf-core
- github.com/Hulupeep/ccf
- github.com/Hulupeep/Specflow
- floutlabs.com
- theshyrobot.com
BSL 1.1, converting to Apache 2.0 in 2032. Contact via floutlabs.com.
Filed scope exercised: Prov 5 (US 64/037,374) section 3.4.3 and Claim 18, assumptions A8 and B1; Prov 4 (US 64/039,655) [E3-0009a] and Examples 1 and 2; Prov 1 Combined Supplement and Prov 3 (US 64/039,623) for the default parameters.