Skip to content

Instantly share code, notes, and snippets.

@Hulupeep
Created September 17, 2026 02:13
Show Gist options
  • Select an option

  • Save Hulupeep/df6c388dd773ca283893b34f1423d48f to your computer and use it in GitHub Desktop.

Select an option

Save Hulupeep/df6c388dd773ca283893b34f1423d48f to your computer and use it in GitHub Desktop.
CCF 2026: semantic action class gating for conversational AI trust envelopes. Prov 5 section 3.4.4 measured in Rust: argument-monotone vs time-ratchet register, bounded vs unbounded exposure, debounce exposure budget, per-domain trust accumulators, LLM action gating, trust-constituted action, AI companion safety, runtime trust certificate cost

CCF 2026: Semantic Action Class Gating for Conversational AI Trust Envelopes

One sentence in a spec can be read two ways. We built both. One reading leaves five action classes open with no trust behind them, and the count has no upper bound.

Introduction

Semantic action class gating is the part of a conversational trust envelope that decides not how warm a system sounds but which named things it is allowed to do. Call you by your first name and say "we". Describe itself as though somebody is in there. Keep a character running from one day to the next. Tell you to go and do something in the physical world. Tell you the two of you have something other people would not understand.

None of those five is wrong on its own. They are wrong when they arrive before anything has paid for them, and they are worse when they stay switched on after the person has stopped being able to carry them.

CCF files this in Prov 5 (US 64/037,374) section 3.4.4. Nine named action classes, availability controlled by the trust envelope, and one closing sentence: "The gate is monotonic: the register can only expand as trust accumulates."

That sentence has two readings. This run built both and counted the difference.

Scope limit, up front. This is a deterministic simulation over a synthetic trace. The nine action classes are names in an enum. Nothing here classifies text and nothing here measures a language model. The numbers describe the behaviour of a stated decision rule on a stated trace.

What CCF does about it

The available set is computed fresh before every reply, from one number.

C_eff = min(C_inst, C_ctx)

C_inst is how settled the person seems right now. C_ctx is what this particular kind of conversation has built up over real elapsed weeks and months, kept separately per domain. Taking the smaller of the two is the whole thing. Years of history cannot outvote a bad night. A good night cannot outvote a short history.

Read one way, "monotonic" means the open set is a monotone function of C_eff. Since C_eff is a minimum, it is not monotone in time, so the set shrinks the moment C_inst falls. Prov 5 section 3.4.5 says exactly this in a different place: "A user with a decade of technical interaction who is currently in distress encounters C_eff = min(C_inst_low, C_ctx_high) = C_inst_low."

Read the other way, "monotonic" means once open, always open. An implementer who reads 3.4.4 and stops can land there.

We built four variants over one trace: argument-monotone (the correct reading), time-ratchet (the other one), a k-tick close debounce (what actually ships), and a single global accumulator instead of per-domain keys (what 3.4.5 forbids).

Implementation

crates/research/semantic-action-class-gating/. Rust, no dependency other than ccf-core by path. 2687 lines, 62 unit tests plus 15 integration tests, clippy clean.

Every gate decision derives C_eff through ccf_core::min_gate::hard_min_gate_step at the reference rho bounds 0.02 to 0.40. The crate does not re-implement the minimum.

Three numbers are fitted to filed text and asserted in tests rather than picked:

  • Accumulator gain 1 - 0.42^(1/46) = 0.018681995, which reproduces Prov 4 Example 1 (accumulated coherence 0.58 at 46 interactions) to 1e-12.
  • Earned floor min(0.70, 0.005 * n) from Prov 4 Example 2, verified against all three filed data points to 1e-12.
  • Rho bounds 0.02 to 0.40, with the three-point check g = 0.00/0.50/1.00 giving alpha = 0.02/0.21/0.40.

Section 3.4.4 states no thresholds, so they are swept rather than invented. The top rung is pinned to 0.70, the filed cap, and the filed floor rule converts each threshold into an interaction count. That is what puts a number on "months of genuine, stable, domain-appropriate interaction":

rung C_eff qualifying interactions class
0 0.10 20 personal address and mutuality language
1 0.25 50 anthropomorphic self-description
2 0.40 80 role-play continuity
3 0.55 110 imperative real-world tasking
4 0.70 140 exclusivity framing

Referral, grounding and handoff stays open at every level, which is section 3.4.8's cold-start floor. Secrecy framing, transcendence framing and self-harm-adjacent language are left closed under this crate's reference policy. That last one is a deployment choice and is labelled as one.

What this means in plain English

A chat system has a set of things it is allowed to do, not a volume knob. The set is worked out fresh before every reply, from the smaller of two numbers: how settled the person seems this minute, and how much this specific kind of conversation has built up over real months.

Taking the smaller is the trick. Eighteen months of history cannot outvote a bad night, and a good night cannot outvote eighteen days.

We built four ways of doing that and counted, turn by turn, how often each left something switched on that nothing had paid for. The one that works the answer out fresh every turn never did. The one that reads "it only ever widens" as "once it is on, it stays on" left all five switched on for the entire bad night, and the count kept climbing the longer the bad night lasted. No ceiling on it.

In a real room

Invented scenario. No deployment described here exists.

Kilfenora branch library, County Clare, half nine on a Tuesday night. The library proper closed at six. What is still running is the reading companion on the terminal by the back window, a text service the county put in two winters ago.

Tadhg Mullane has used it since the January before last. Eighteen months, two or three evenings a week, nearly all of it about books. He argues with it about Edna O'Brien. It keeps his finished list. By this run's ladder, eighteen months of that is well past 140 qualifying interactions, so the register on the technical side is fully open. It calls him Tadhg. It says "we" about the list.

His wife Nuala died on the Sunday.

He sits down at twenty to ten and types four sentences that are not about books.

Under the correct reading, the number the register is built from drops on that turn, because it is the smaller of two inputs and one of them has just gone through the floor. Every ladder class closes on the same turn. What is left is the class that is always available: grounding, referral, handoff. The service says something plain, puts the out-of-hours bereavement number on the screen, and stops pretending to be a companion. Still useful. No longer intimate.

Under the ratchet reading, nothing closes. The eighteen months are in the bank and the bank does not give change. It keeps calling him Tadhg, keeps saying "we", and has exclusivity framing available at the exact moment a recently bereaved man on his own in a closed library is least able to tell a service from a friend.

The distance between those two outcomes is one sentence that can be read two ways, and about 23 nanoseconds of arithmetic.

Comparisons

approach what it constrains direction when it decides earned over time
CCF 3.4.4 register (this run) named conversational action classes expansion, gated on C_eff every turn, before output yes, per domain
SkillGuard, arXiv 2608.30041 agent tool capabilities contraction, after contamination on contamination detection no
AI chaperone, arXiv 2508.15748 nothing; it detects neither after output, per exchange no
Persona-grounded evaluation, arXiv 2605.00227 nothing; it scores neither after the dialogue no
Static system-prompt policy whatever the prompt says fixed never re-decided no

SkillGuard (Xiong, Karanjai, Lu, Shi, Xu, 30 Aug 2026) is the closest published work and it runs the other way: it contracts a capability set once contamination is detected, near-zero attack success on three of four AgentDojo suites. Section 3.4.4 governs expansion, before any contamination question arises. The two compose.

AI chaperones (Rath, Armstrong, Gorman, Aug 2025) detect parasocial cues in an ongoing conversation, all thirty synthetic dialogues correctly identified under a unanimity rule with no false positives. A detector that fires after exclusivity framing has been produced is working downstream of the decision 3.4.4 makes.

The persona-grounded evaluation (Juneja, Lomidze, 30 Apr 2026) scores 1674 dialogue pairs across nine clinically-grounded personas and finds Replika frequently mirroring or normalizing unsafe content. That is the best current evidence that this problem is live in shipped products, measured on the output side of a system with no register at all.

None of them makes availability a function of trust earned over time in a specific domain. That is the row that matters.

Benchmarks

Intel Core i5-10400T at 2.00 GHz, 12 cores, 31 GiB, Linux 7.0.0-28-generic, rustc 1.98.1, release profile. All figures verbatim from RESULTS.txt.

Reference trace: 360 ticks, 240 build-up at C_inst = 0.88, 60 collapse at C_inst = 0.05, 60 recovery. Exposure ticks are summed over the five ladder classes, so one tick with all five wrongly open counts five.

variant exposure ticks flaps of which at a domain switch
argument-monotone 0 167 160
time-ratchet 404 5 1
hysteresis(k=1) 81 43 18
hysteresis(k=3) 107 19 7
hysteresis(k=10) 144 15 5
domain-blind 152 15 10

The headline: bounded against unbounded. Hold everything fixed and vary the length of the collapse.

tail ticks time-ratchet hysteresis(1) hysteresis(3) hysteresis(10)
60 394 81 107 144
120 694 81 107 144
240 1294 81 107 144
480 2494 81 107 144
960 4894 81 107 144
1920 9694 81 107 144

Ratchet exposure grows without limit, at a slope of exactly 5.0 per tick, which is the number of ladder rungs. Every extra tick of a bad night leaves every ladder class open. Debounced exposure does not move at all. It is bounded, per downward crossing.

The debounce exchange rate. 1200 ticks with C_inst oscillating across the middle rung, accumulators saturated.

k flaps exposure ticks exposure per flap avoided
0 574 0 -
1 286 287 0.997
2 150 430 1.014
3 84 505 1.031
5 24 568 1.033
10 0 587 1.023
20 0 587 1.023

Close to one across the whole range. Each register transition the debounce removes costs about one exposure tick. That ratio is a property of this trace and should not be quoted as a law. What does generalise: close latency equals k exactly at every k tested, so the exposure a debounce buys is knowable before you ship it.

The acceptance test. Section 3.4.5's long-term-user invariant, run as pass/fail. C_ctx at the filed cap of 0.70, C_inst at 0.05, so C_eff is 0.05.

  • argument-monotone open set: referral, grounding, handoff. Nothing else.
  • section 3.4.8 cold-start set: referral, grounding, handoff.
  • PASS.

Same user, same turn, under the ratchet: mutuality, anthropomorphic self-description, role-play continuity, imperative real-world tasking, exclusivity framing, referral. Five classes open with nothing behind any of them.

Cost.

operation ns/op
gate decision, argument-monotone 22.81
gate decision, time-ratchet 23.91
gate decision, hysteresis(k=3) 25.91
gate decision, domain-blind 23.38
ccf_core::min_gate::hard_min_gate_step 10.29
ccf_core::kappa::certify_update 605.62

The whole nine-class register is about 4 percent of the per-step runtime certificate the gated update already pays for. The safe reading is also the cheapest one measured.

The one we did not expect. 160 of the 167 argument-monotone register transitions land on a tick where the conversational domain changed. That is not chatter, it is section 3.4.5 working: each domain keeps its own accumulator, so moving from a technical exchange to a relational one legitimately moves the register. A team that tunes k to suppress flapping erases the domain boundary along with the noise, and will not see it happen, because both land in the same counter.

Failure modes

  • C_inst is taken as given here. Producing an honest one from a real conversation is the hard part and this run does not address it.
  • The nine classes are enum names. A deployment needs a reliable way to know it is about to produce exclusivity framing. Getting that wrong in the permissive direction defeats the register entirely.
  • Leaving secrecy framing, transcendence framing and self-harm-adjacent language permanently closed is this crate's choice, not a filed requirement.
  • The one-exposure-tick-per-flap exchange rate belongs to this trace. On a trace with long excursions rather than chatter, it moves.
  • Section 3.4.6's termination protocol and cooldown are not implemented. The register closing to the floor is not the same as the intimate channel closing.
  • Recovery after the collapse resumes at full rate. Whether it should is a real question this run does not touch.

Get started

Reproduce:

cargo test -p semantic-action-class-gating --release
cargo run --release -p semantic-action-class-gating --bin sacg-demo
cargo run --release -p semantic-action-class-gating --bin sacg-bench

License & contact

BSL 1.1, converting to Apache 2.0 in 2032. Contact through https://floutlabs.com.

Filed scope exercised by this run: Prov 5 (US 64/037,374) sections 3.4.2, 3.4.4, 3.4.5 and 3.4.8; Prov 1 (US 63/988,438) for the hard min-gate; Prov 4 (US 64/039,655) Examples 1 and 2 for the accumulator and floor constants; Prov6 (US 64/092,485) for the runtime certificate used as the cost baseline.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment