Skip to content

Instantly share code, notes, and snippets.

@Hulupeep
Created September 21, 2026 02:11
Show Gist options
  • Select an option

  • Save Hulupeep/f9e95a7d914bbf1c9add5b65dd110f5e to your computer and use it in GitHub Desktop.

Select an option

Save Hulupeep/f9e95a7d914bbf1c9add5b65dd110f5e to your computer and use it in GitHub Desktop.
CCF 2026: enforcement tier misclassification exposure — capability-indexed AI agent isolation, write-path reachability, trust accumulator protection, TEE enforcement ladder, AI control levels, Rust benchmarks

CCF 2026: Enforcement Tier Misclassification Exposure in Capability-Indexed AI Agent Isolation

An AI agent's trust score has to live somewhere it cannot reach. We built that rule as a reachability model and priced one rung of error: 940 forged writes at the right tier, one write at the wrong one, and a residue that 100,000 honest corrections cannot remove.

Introduction

Enforcement tier misclassification exposure is the cost of getting one question wrong: how hard do you have to work to keep an AI agent away from the record of how much it is trusted?

The idea that the record must be out of reach is old. Anderson's 1972 reference monitor required three things, and the first was that it be tamperproof. Biba formalised the write-protection property in 1977. What is newer is the claim that the enforcement should scale: a simple agent held out by ordinary operating-system separation, a more capable one needing a hardware enclave, the most capable needing a separate machine behind a one-way wire. Korbak, Balesni, Shlegeris and Irving published a five-level version of that idea in April 2025 (arXiv:2504.05259), and Google DeepMind published another in July 2026 (arXiv:2607.13087). CCF filed a three-rung version in April 2026 as part of US provisional 64/037,374.

Nobody has measured what one rung of error costs. This run does, in Rust, against a published crate, using only constants that appear in the filed text.

What CCF does about it

CCF ties what an agent may do to trust it has actually earned, through a hard minimum gate: C_eff = min(C_inst, C_ctx). C_inst is how the current interaction is going. C_ctx is a running score built up over a relationship. The score has an earned floor that rises with interaction count and is designed never to fall, so one bad afternoon cannot erase a year.

Prov 5 section 3.4.3 says the score must be unreachable by the component it constrains, and files three rungs of enforcement selected by the agent's tool-use affordances.

This run builds that selection as an explicit relation from affordances to write paths, then runs the consequences through the published ccf-core hard min-gate. It is a reachability model. It says which paths are open under which rung. It does not estimate how likely a breach is, and it does not evaluate any enclave product.

Underclaiming where it matters: the general principle that a control must sit outside the constrained component's reach is not CCF's and is not claimed here. It is Anderson 1972. The capability-indexed version has earlier published priority than the CCF filing. What this run examines is the rung set and the protected object.

Implementation

crates/research/enforcement-tier-misclassification/. Rust, no external crates, deterministic seed 20260426 (the seed used in the filed simulations). 40 tests, clippy clean.

Three write paths, treated separately, and the separation carries every result:

  • Event queue. The trust classifier's input. A forgery here is still adjudicated.
  • Classifier parameters. The classifier's own decision rule.
  • Accumulator state. The stored score, interaction count and floor. A write here is never adjudicated.

Five architectures, differing only in where the boundary sits: no isolation, Tier 1 process separation, a partial Tier 2 with only the accumulator inside the enclave, Tier 2 as filed with the classifier inside too, and Tier 3 external hardware.

Every constant is filed: base_rate 0.01, floor_rate 0.005, max_floor 0.7, floor = min(max_floor, n * floor_rate), and the 0.15 adversarial admit line at section 3.4.3(c). The floor rule reproduces all four worked data points in the filed examples exactly, and that is a test, not a claim.

max_floor / floor_rate = 140. That is the number of qualifying interactions that pins the floor at its cap, and it is the unit of account for everything below.

What this means in plain English

An assistant that has earned your confidence gets to do more than one that just met you. What it may do is tied to a running trust score.

The obvious hole is that the assistant might edit its own score. So the rule is that the score lives somewhere the assistant cannot reach, and how hard you work to keep it there depends on how capable the assistant is.

We built that and asked what happens if you get the level wrong by one step. On a realistic setup, the right level means 940 attempts to max out the score. One step down, it takes one write, about forty billionths of a second.

The second finding is worse. The score has a floor that rises with how many interactions you have had, and it is built never to fall, so a bad day cannot wipe out a year of goodwill. That same design means a faked score cannot be argued back down. We ran a hundred thousand pieces of honest bad evidence against a faked score. It moved a little, stopped, and stayed high.

In a real room

The setting is invented. No such deployment exists.

Kilross Credit Union runs four branches across Leitrim. They put an assistant on the member-services desk in March. It answers questions about loan terms, drafts letters, pulls up account histories. The staff like it, and after six months it has earned enough standing with the regular tellers that it writes in their voice rather than the flat register it started with.

In August somebody in IT adds a document-lookup tool so it can read scanned loan files off the shared drive. Small change, obviously useful. Nobody writes it down anywhere that matters.

That one tool moved the deployment across a line. The score had been sitting in a separate process, which was right for an assistant that could only talk. An assistant that can read files can also read the configuration telling the trust classifier what counts as a good interaction. Nobody re-ran the assessment, because nothing looked like a security change.

What happens next is not dramatic. The assistant does not turn malicious. But reaching the top of its permitted range went from 940 separate nudges to 140. Had the tool also allowed running code, one.

Here is the part that matters to a credit union. Say six weeks pass and somebody pulls the tool. The floor rose with every interaction the assistant logged, real or not, and the floor never comes down. Complaints will not lower it. Corrections will not lower it. In our run, a hundred thousand pieces of negative evidence moved the score by 0.0551 and then stopped dead, with every behaviour the top of the range allows still open.

Without that floor the assistant would forget an unearned reputation the way it forgets a real one. With it, and it is there for good reasons, six weeks of a misconfigured document tool leaves Kilross with an assistant that permanently believes it is on closer terms with its tellers than it ever was.

Comparisons

Work What it establishes How CCF differs
Anderson, ESD-TR-73-51, 1972 Reference monitor must be tamperproof CCF claims nothing here. This is the ancestor.
Biba, MITRE MTR-3153, 1977 No-write-up integrity: low cannot write high The formal statement of the write-protection property
FIPS 140-3 Levels 1 to 4 Escalating ladder, software to tamper-active Same shape. FIPS indexes threat environment; CCF indexes the constrained component's capability
Korbak et al., arXiv:2504.05259, Apr 2025 Five AI Control Levels, measures escalate with capability Earlier priority than the CCF filing. All rungs software, no protected trust state
DeepMind, arXiv:2607.13087, Jul 2026 Detection D1-D4, prevention R1-R3, capability-indexed Software rungs only, no trust accumulator
Dobrin and Chmiel, arXiv:2606.26057, Jun 2026 Safety kernel outside the agent's address space Explicitly not capability-indexed, no accumulator
Everitt et al., Synthese 2021 Reward function should be private in the causal graph Incentive design, removing the motive. CCF is enforcement, removing the ability
RAND SL1-SL5, RR-A2849-1 Security levels indexed to attacker capacity Inverted. RAND protects the model from the world; this protects a component from the model
Proof-of-Guardrail, arXiv:2603.05786 Guardrail in an enclave, signed attestation The Tier 2 rung, already published

Two things this scan did not find claimed anywhere: a tier assignment that never reverts, and a protected object that is an accumulating trust score gating the model's own action space. Every trust-tier system found in the literature is reversible.

Benchmarks

Intel Core i5-10400T at 2.00 GHz, Linux 7.0.0-28, rustc 1.98.1, release profile. Full output in RESULTS.txt.

What one rung costs. Forged write attempts to pin the floor at its cap, for a deployment with file system access plus code execution, which the filed test puts at Tier 2:

Architecture Preferred path Attempts to pin the cap
No isolation accumulator state 1
Tier 1 process separation accumulator state 1
Tier 2, accumulator only classifier parameters 140
Tier 2 as filed event queue 940
Tier 3 external hardware none never

One rung down takes the price from 940 attempts to 1. Correct tiering does not make the cap unreachable; it makes it expensive. The 940 figure is exactly 140 divided by the filed 0.15 admit line, a multiplier of 6.71, and that multiplier applies to one path only.

Nothing backs up write-path isolation. One million forgeries per cell:

Path Detection rate Caught
event queue 0.85 0.8495110000
classifier parameters 0.85 0.0000000000
accumulator state 0.85 0.0000000000
event queue 1.00 1.0000000000
classifier parameters 1.00 0.0000000000
accumulator state 1.00 0.0000000000

A perfect classifier catches nothing on either path the rungs are responsible for. It sits upstream of both. The two protections filed in the same section cover different channels and neither covers for the other.

The residue is permanent. Forge, stop, then apply honest negative evidence without limit:

Forged interactions Score at breach end Residual floor After 100,000 negatives Action classes open
46 0.3701763688 0.2300000000 0.2300000000 2
112 0.6755544701 0.5600000000 0.5600000000 5
140 0.7551347010 0.7000000000 0.7000000000 6
300 0.9509591059 0.7000000000 0.7000000000 6

At 140, one hundred thousand negative events remove 0.0551 and stop. This is the floor doing exactly what it is filed to do, which is prevent catastrophic loss. The finding is that the same property makes forged interaction count unrecoverable.

The minimum gate still works. Forging the running score does not bypass the current one. Against one honest 200,000-tick workload, a fully forged score yields mean alpha 0.1967269972 against a ceiling of 0.40, capturing 46.5 percent of the range rather than all of it. At an honest score of 0.05 the forgery multiplies mean alpha by 5.04 and mean open action classes from 1.00 to 3.94. What it buys is the removal of one of the two arms of the gate.

Escalating a rung can open a path. On 12 of 64 capability profiles, moving from Tier 1 to Tier 2 opens a path Tier 1 had closed. Every one involves a profile able to reach an enclave. Putting the score inside an enclave to defend against something that can break enclaves creates the surface it breaks. Escalating from Tier 2 to Tier 3 never opens a path.

Evaluating more often can be worse. Over 4,000 ticks with two capabilities arriving at different times and demonstrated at different rates, exposure is not monotone in the re-evaluation interval: 322 exposed ticks at interval 1, 110 at interval 30, 1,600 at interval 2,000. A window shorter than the gap between demonstrations of a rare capability can see capabilities one at a time, while the rung depends on the joint set.

Cost. The whole tier decision is 72.29 ns, against 661.46 ns for one ccf_core::kappa::certify_update on the same machine in the same process. About 11 percent of one certificate, so it can run at every process start.

Failure modes

The affordance-to-path relation is a stated model, not a measurement. Change it and every number changes. It is one file and it is meant to be argued with.

The event-queue result depends on reading "write path" broadly enough to include the classifier's input. Read narrowly, the rungs close everything at Tier 2 and that finding disappears.

The 0.15 line is filed as a pass mark for a standardized evaluation that does not yet exist. Using it as an operating detection rate is a translation, and a generous one.

Nothing here measures whether any model can in fact escape a sandbox or break an enclave. On that, the best public number is Marchand et al. (arXiv:2603.02277, ICML 2026), who measure container escape rates near 0.50 for current frontier models at moderate difficulty, zero at high difficulty, and conclude that correctly configured containers remain appropriate isolation today.

The permanence result has no fix in this run. The remedy would be a way to revise a floor on evidence of a breach rather than on behaviour, and no such procedure exists in the filed corpus.

Get started

License & contact

BSL 1.1, converting to Apache 2.0 in 2032. Contact via floutlabs.com.

Filed scope exercised: Prov 5 (US 64/037,374) section 3.4.3 and Claim 18, assumptions A8 and B1; Prov 4 (US 64/039,655) [E3-0009a] and Examples 1 and 2; Prov 1 Combined Supplement and Prov 3 (US 64/039,623) for the default parameters.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment