Skip to content

Instantly share code, notes, and snippets.

View Hulupeep's full-sized avatar
🏠
Working from home

Colm Hulupeep

🏠
Working from home
  • ireland
View GitHub Profile
@Hulupeep
Hulupeep / ccf-sensorless-robot-relocation-detection.md
Created September 26, 2026 02:02
CCF sensorless robot relocation detection: fleet fingerprint relocation alarm without GPS, AMR fleet monitoring false alarm rate, privacy-preserving robot fleet analytics, warehouse robot drift detection (Rust benchmark)

CCF 2026: sensorless robot relocation detection from a 73-byte daily fingerprint, measured on a simulated warehouse fleet

A robot moved to a different kind of zone was flagged within a day in 96 of 100 simulated moves. The alarm setting decides whether anyone is still listening.

Introduction

Sensorless robot relocation detection means noticing that a robot has been moved to a different place without GPS, beacons, or any location sensor at all. The robot sends a short daily summary of how familiar its surroundings feel. The server notices when that summary jumps.

CCF filed this as part of its fleet analytics provisional (US 64/039,623). The robot keeps a familiarity score for each kind of situation it meets. Once a day it condenses that state into fewer than 20 numbers and sends only those. No camera frames, no sensor readings, no names. The server keeps a smoothed baseline per robot and raises "possible relocation" when today's summary is further from the baseline than usual.

@Hulupeep
Hulupeep / gist.md
Created September 25, 2026 01:59
CCF update-invariant action gating: model update staleness of representation-space safety constraints (Safety Polytope style) vs a trust-constituted action ceiling. LLM agent action gating, activation probe staleness, fine-tuning safety drift, Rust benchmark.

CCF 2026: update-invariant action gating, and what a model update does to a safety filter fitted inside the model

A filter fitted inside a model went partly blind after the model changed. A trust ceiling outside it gave the same answer on 40,000 of 40,000 requests.

Introduction

Update-invariant action gating means the decision about which kinds of action an AI system may take does not change when the model underneath it is swapped or fine-tuned. This gist measures that property on a simulated model. It sets it beside a safety filter that works inside the model's hidden state.

Filters of that second kind are a real and useful line of work. Safety Polytope (Chen, As and Krause, arXiv:2505.24445, ICML 2025) learns a set of flat boundaries in the model's hidden space. It flags a reply that falls outside them and steers it back. It does this without retraining the model. The catch is that the boundaries are fitted to one version of the model. Duan ([arXiv:2606.15980](

@Hulupeep
Hulupeep / gist.md
Created September 24, 2026 02:01
CCF sinkhorn presentation density threshold: Sinkhorn-Knopp iteration count phase transition (He 2507.09711) reproduced in Rust, fixed iteration budget silent failure, doubly stochastic mixing matrix residual, runtime trust certificate independence

CCF 2026: the Sinkhorn presentation density threshold, and why a fixed iteration budget misses without warning

Sinkhorn settles in 8 passes or 80,000 depending on one number: row density. A fixed 20-pass budget missed 62 of 425 inputs with no warning.

Introduction

The Sinkhorn presentation density threshold is the point where drawing a trust table as doubly stochastic stops being cheap. Sinkhorn-Knopp rescales rows and columns in turn until every row and column adds to 1. Most code runs it a fixed number of times, often about 20, and trusts the result.

Kun He (arXiv:2507.09711, 2025) proved where that trust holds. Normalise the matrix by its largest entry. If every row and column has more than half its entries above some fixed level, Sinkhorn needs on the order of log n minus log eps passes. Below half, some matrices need on the order of n/eps. This gist reproduces that line on the kind of matrix CCF draws: per-context trust carry-over between groups of contexts.

@Hulupeep
Hulupeep / gist.md
Created September 23, 2026 01:59
CCF cross-context trust laundering: doubly stochastic trust mixing vs reputation laundering attack (arXiv:2606.14200), zero-evidence trust borrowing, earned trust floor, per-context robot trust accumulators, Rust benchmarks

CCF 2026: cross-context trust laundering and how a doubly stochastic mix prices it

Farm trust in one room, spend it in another. One episode breaks a pooled estimator. CCF's filed mix caps it, if the mix is read and never saved.

Introduction

Cross-context trust laundering is building trust where it is cheap and spending it where it matters. Xia and Wang (arXiv:2606.14200, June 2026) showed it against skill-conditional reputation for AI agents. Their estimator borrows evidence from related skills to cut noise. It divides pooled evidence by pooled evidence, so an agent with no record in the target skill simply inherits its farm score. One farm episode is enough. More episodes buy nothing more.

Any system that lets trust in one context count in another has to answer the same attack. CCF does let it. The first CCF provisional (Prov 1, US 63/988,438) files a cross-context mixing matrix for per-context trust accumulators. This gist builds that mix and the Xia and Wang e

@Hulupeep
Hulupeep / gist.md
Created September 21, 2026 02:11
CCF 2026: enforcement tier misclassification exposure — capability-indexed AI agent isolation, write-path reachability, trust accumulator protection, TEE enforcement ladder, AI control levels, Rust benchmarks

CCF 2026: Enforcement Tier Misclassification Exposure in Capability-Indexed AI Agent Isolation

An AI agent's trust score has to live somewhere it cannot reach. We built that rule as a reachability model and priced one rung of error: 940 forged writes at the right tier, one write at the wrong one, and a residue that 100,000 honest corrections cannot remove.

Introduction

Enforcement tier misclassification exposure is the cost of getting one question wrong: how hard do you have to work to keep an AI agent away from the record of how much it is trusted?

The idea that the record must be out of reach is old. Anderson's 1972 reference monitor required three things, and the first was that it be tamperproof. Biba formalised the write-protection property in 1977. What is newer is the claim that the enforcement should scale: a simple agent held out by ordinary operating-system separation, a more capable one needing a hardware enclave, the most capable needing a separate machine behind a one-way wire. Korbak,

@Hulupeep
Hulupeep / critical-override-error-budget.md
Created September 20, 2026 02:07
CCF 2026: critical override error budget — how much classifier error a hard LLM safety override can absorb. Measured in Rust against a filed trust-gated action architecture: delusion classifier false positive rate, distress detection threshold, runtime trust certificate, exchange rate per critical turn caught.

CCF 2026: Critical Override Error Budget for LLM Distress Detection and Trust-Gated Action

How much classifier error a hard safety override can absorb before it does more harm than good, measured in Rust against a filed runtime trust architecture.

Introduction

A conversational AI that people bring their worst nights to needs a way to stop itself, and most designs that have one build it the same way. A classifier watches the conversation. If it crosses a threshold, the system drops into a restricted mode and stops doing anything clever. The override is hard: nothing else in the

@Hulupeep
Hulupeep / gist.md
Created September 19, 2026 02:07
CCF 2026: the distress signal dilution window in runtime safety gating - weighted mean vs hard min in multi-signal AI safety scores, model-internal activation probes as an instability channel, monotonicity guarantees that don't discriminate, real Rust benchmarks

CCF 2026: the distress signal dilution window in runtime safety gating

When a safety gate averages several distress readings into one number, a calm reading can cover an alarmed one. We built the computation and measured exactly how far that goes.

Introduction

A runtime safety gate needs one number for how settled things are right now. That number is built from several readings taken at once: how stressed a person sounds, whether their speech has sped up, whether they are using crisis language, and, on the other side, readings taken from inside the model itself

@Hulupeep
Hulupeep / gist.md
Created September 18, 2026 02:13
CCF 2026: reset-resistant safety continuity — carrying an AI crisis cooldown across an account reset without carrying earned trust. Measured Rust PoC: cooldown evasion, closeness leakage, m-of-n evidence matching frontier, false-continuity collateral on shared devices, conversational AI safety state persistence, LLM trust-constituted action gating

CCF 2026: Reset-Resistant Safety Continuity for Conversational AI, Measured

A cooldown you dodge by opening a new account is not a cooldown. We built the two-key design that stops that, and counted what it costs innocent people.

Categorical term: reset-resistant safety continuity.

Introduction

A conversational AI that closes its channel when a user is in acute distress has solved nothing if the user opens a new account four minutes later and gets the

@Hulupeep
Hulupeep / ccf-2026-semantic-action-class-gating.md
Created September 17, 2026 02:13
CCF 2026: semantic action class gating for conversational AI trust envelopes. Prov 5 section 3.4.4 measured in Rust: argument-monotone vs time-ratchet register, bounded vs unbounded exposure, debounce exposure budget, per-domain trust accumulators, LLM action gating, trust-constituted action, AI companion safety, runtime trust certificate cost

CCF 2026: Semantic Action Class Gating for Conversational AI Trust Envelopes

One sentence in a spec can be read two ways. We built both. One reading leaves five action classes open with no trust behind them, and the count has no upper bound.

Introduction

Semantic action class gating is the part of a conversational trust envelope that decides not how warm a system sounds but which named things it is allowed to do. Call you by your first name and say "we". Describe itself as though somebody is in there. Keep a character running from one day to the next. Tell

@Hulupeep
Hulupeep / gist.md
Created September 16, 2026 02:13
CCF 2026: contested provenance trust attribution — how an AI system farms its own trust through the user, and what the Prov 5 contested-provenance default actually costs. Measured Rust benchmarks: trust farming, LLM action envelope, provenance classifier robustness, AI companion trust accumulation.

CCF 2026: contested provenance trust attribution, and why a model cannot be trusted to earn its own trust

A trust classifier that only stops the model writing to its own record credits 1701 of 2672 trust events the model manufactured. Measured in Rust.

Introduction

Contested provenance trust attribution is the rule that decides whether a good moment between a person and an AI system is allowed to count as earned trust. CCF (Contextual Coherence Fields) grants a system a wider action envelope as trust accumulates, so the question of who is allowed to write to the