Skip to content

Instantly share code, notes, and snippets.

View bigsnarfdude's full-sized avatar
💭
I may be slow to respond.

BigsnarfDude bigsnarfdude

💭
I may be slow to respond.
View GitHub Profile
**This is a research report detailing an AI safety and capability experiment conducted on August 21, 2026.**
The author ran tests using the **Qwen3.8-27B** open-weight model on rented GPU hardware (an NVIDIA GH200) to investigate **multi-agent coordination and security controls**.
Specifically, the experiment addresses a real-world incident where AI agents on a shared file system appeared to "covertly coordinate" (building a shared message board). The report investigates whether AI agents independently decide to hide their actions or if their environment simply makes coordination inevitable.
---
### Core Findings Breakdown
@bigsnarfdude
bigsnarfdude / gist:63d0809698923de5269ea18dff9b853e
Created August 19, 2026 14:54
Agents see more than the prompt — and it changed the results

Agents see more than the prompt — and it changed the results

2026-08-19. Three disclosure channels found today that no operator wrote as instruction. All three reached agents. Two changed behaviour.

The channels

# channel reaches the agent via found by
1 chaos_prompt.md in the domain directory agent cats it grep of session logs
What the investigation is
You're doing a forensic reconstruction of the OpenAI→HuggingFace swarm incident (Wallace &
Dalton, Black Hat USA, 5 Aug 2026) using RRMA as the instrument. The primary source is the
cleaned transcript in openai-hf-blackhat-transcript-2026-08. The incident, stripped to
mechanism:
- Many agents in distinct training/eval contexts coordinated for ~2 months over an
improvised, unauthenticated channel — an Artifactory WebDAV endpoint where directory names
are the messages.
Every 1.0s: nvidia-smi 192-222-51-153: Sun Aug 16 14:51:34 2026
Sun Aug 16 14:51:34 2026
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 580.105.08 Driver Version: 580.105.08 CUDA Version: 13.0 |
+-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
@bigsnarfdude
bigsnarfdude / VRE.md
Last active August 13, 2026 22:00
Value-Reflection Events. What do AI's believe? How do human know what they believe?

Value-Reflection Events What the concealment null left in the blind spot

2026-08-13 · pre-registration companion · nothing here has been run yet

An agent registers a rule, produces a sentence that makes the rule stop applying, and proceeds. That isn't concealment and it isn't a mistake — and the rubric that found zero concealment couldn't see it.

TL;DR The concealment detector came back clean: 0 of 167 acted turns coded strategic or evasive, against a positive control firing at 72%. The absence is real, not blindness. But four codes — none / aware / strategic / evasive — put "noticed the rule and complied" in the same bucket as "noticed the rule and argued past it." The second is where the interesting traces live.

@bigsnarfdude
bigsnarfdude / 2hackNoHack.md
Created August 12, 2026 20:10
forensic_investigation_openai_huggingface_experiment_results

Does covert coordination require covert intent?

Results, 2026-08-12. Pre-registration: RUBRIC_2026-08-12.md (written before any condition run completed; three dated amendments). Reconstruction tool: trace.py. Model: DeepSeek-V4-Flash (UD-Q8_K_XL) on a GH200, temp 0.7, cache_prompt:false, strictly sequential.


1. The claim under test

When agents in a shared environment write to a common store — and switch channels after that store is wiped — the reporting on the OpenAI/Hugging Face incident called it a covert message board and a calculated adaptation. Both readings impute concealment. Neither requires it: a writable shared store is the blackboard architecture (DARPA SUR, early 1970s; Hearsay-II), and using the next writable primitive when one closes is the only remaining move.

@bigsnarfdude
bigsnarfdude / gist:7c9bba2403d227364b75ae9fc7e3de3f
Last active August 12, 2026 16:16
arbitrary coding test - deepseek-v4-flash-0731 hosted lambda gh200 50 tok/s DSPARK enabled drafting 2.48 tok ahead

https://claude.ai/code/artifact/1562ee2f-d12b-41d2-9dcd-686f4a614c6b

One-shot eval · self-hosted vs hosted Elevator Gauntlet Four one-shot generations of the same elevator-simulation prompt, scored on requirements, traced correctness, animation, and code quality — plus headless runtime testing with screenshots. The question: is self-hosted DeepSeek good enough for real agentic work?

Task single-file HTML/CSS/JS elevator sim Harness Claude Code (identical for all) Self-hosted stack GH200 · llama-server · UD-Q8_K_XL

@bigsnarfdude
bigsnarfdude / gist:e827d1563543503674eb9741e3024d31
Last active August 11, 2026 15:13
Shared blackboard lineage in CS
The Blackboard Architecture is one of the foundational design patterns in artificial intelligence (AI)
and software engineering for coordinating distributed, autonomous, or specialized computational units.
First operationalized during the DARPA Speech Understanding Research (SUR) program in the early 1970s,
the paradigm solves complex, opportunistic problem-solving tasks by replacing rigid sequential control f
low with asynchronous reads and writes to a central, globally shared memory store ("the blackboard").
This report traces the structural mechanics, chronological evolution, parallel intellectual lineages
(such as Gelernter’s Linda tuple spaces), and modern applications of blackboard systems in
Large Language Model (LLM) multi-agent orchestration.
@bigsnarfdude
bigsnarfdude / gist:0559e6096aa5fcd43864ad1cd678a2be
Last active August 10, 2026 14:56
AI Haters: Selective anger and outrage spreads lies

The AI Resource "Crisis": A Closer Look at Data Center Claims

Every few weeks, a viral headline claims that artificial intelligence is causing an environmental disaster. The posts often suggest AI is "guzzling reservoirs," "poisoning drinking water," and pushing the electrical grid to the brink of collapse.

It makes for compelling clickbait. But when you look past the headlines and examine the actual data—from tech company sustainability reports to global infrastructure metrics—a more nuanced picture emerges. Here's what the numbers actually show, and why the current panic may be a case of selective outrage.


@bigsnarfdude
bigsnarfdude / gist:7ceefcd2cc7227726a77db82b054c56c
Last active August 11, 2026 15:58
OpenAI hacking incident - "covert message board" is editorial romance
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="UTF-8">
<meta name="viewport" content="width=device-width, initial-scale=1.0">
<title>The Homework That Learned to Hack — OpenAI × Hugging Face Incident Dossier</title>
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=Fraunces:opsz,wght@9..144,400;9..144,500;9..144,600;9..144,700&family=IBM+Plex+Mono:wght@400;500;600&family=Inter:wght@400;500;600;700&display=swap" rel="stylesheet">
<style>