Skip to content

Instantly share code, notes, and snippets.

View magnus919's full-sized avatar
💭
Diving back into Hugo and figuring out how/where I want to deploy it.

Magnus Hedemark magnus919

💭
Diving back into Hugo and figuring out how/where I want to deploy it.
View GitHub Profile
@magnus919
magnus919 / ai-agent-memory-substrates.md
Created July 21, 2026 02:50
What database modalities teach us about AI agent memory

What Database Modalities Teach Us About AI Agent Memory

No database modality is “agent memory” by itself. Each preserves a different dimension of experience, while useful memory depends on the surrounding policy: what gets stored, consolidated, versioned, forgotten, retrieved, and trusted.

Modality tradeoffs

  • Vector stores are strong at fuzzy semantic recall and candidate generation. They are weak at exact identity, relationships, causality, chronology, and staleness. Good implementations add metadata filters, entity signals, and reranking.

  • Graph databases are strong at relationships, provenance, multi-hop reasoning, contradictions, and causal paths. Their costs are entity extraction, ontology maintenance, traversal complexity, and heavier operations.

@magnus919
magnus919 / qwen-130-percent-de-spin.md
Created July 19, 2026 13:27
De-spin audit: the claim that Qwen created 130% more security vulnerabilities for U.S. users

De-spin: Did Qwen create 130% more security vulnerabilities for U.S. users?

Public claim audit, accessed July 19, 2026

Bottom line

Verdict: Misleading

The tweet is built around a real result reported by Booz Allen Hamilton, but it does not describe that result accurately. Booz Allen's chart reports a 130% increase in an aggregate vulnerability score for Qwen3-Coder when a neutral system prompt was changed to say the model was generating code for a U.S. federal agency. The tweet turns that into 130% more security vulnerabilities, broadens the condition to a user "from the USA / foreign," and frames the outcome as something Qwen's creators chose to build into the model. The report does not establish any of those broader claims. Booz Allen explicitly says it has no proof that the flaws were introduced intentionally.

@magnus919
magnus919 / linear-agent-skill-demo-v2.md
Created July 17, 2026 19:14
Real anonymized demo: complete Linear issue workflow with the linear Agent Skill

A real Linear workflow from an Agent Skill

The linear Agent Skill gives an AI agent a focused, auditable way to operate Linear through its public GraphQL API. It uses Python 3.8+ with no third-party Python packages and no MCP server.

This demo performs a complete issue workflow against a live Linear workspace:

  1. Discover the team.
  2. Preview issue creation without changing Linear.
  3. Create the issue.
  4. Update its title and priority.
@magnus919
magnus919 / chinese-ai-distillation-allegations.md
Created July 17, 2026 03:52
What the public evidence does and does not show about alleged Chinese AI model distillation

What the Public Evidence Does and Does Not Show About Alleged Chinese AI Model Distillation

Assessment date: July 16, 2026
Verdict: Complicated

Executive summary

The common assertion that “Chinese AI labs are distilling Western models” has a factual core, but usually outruns the public evidence.

Anthropic has publicly alleged that DeepSeek, Moonshot AI, and MiniMax used fraudulent accounts and proxy infrastructure to extract Claude capabilities at industrial scale. It says the campaigns generated more than 16 million exchanges through about 24,000 accounts. The U.S. House Select Committee on the Chinese Communist Party separately reports that OpenAI told it DeepSeek employees extracted reasoning outputs and used OpenAI models in training-data processing.

@magnus919
magnus919 / election-security-evidence-rollup.md
Created July 17, 2026 03:14
Election-security address: evidence rollup and source audit

Election-security address: evidence rollup

Speech assessed: President Donald Trump’s July 16, 2026 address on election security (video).
Research date: July 16, 2026.
Scope: Major factual claims about the security and outcome of U.S. elections, not the address’s separate economic, immigration, military, or policy claims.

Method and source boundary

The White House’s speech, press material, and document bundles are used only to establish what the administration said or released. They are not used to validate the administration’s factual claims.

@magnus919
magnus919 / gpt-red-claim-audit.md
Created July 15, 2026 18:16
Evidence audit of OpenAI's GPT-Red self-improvement claims (July 15, 2026)

Claim audit: OpenAI’s “GPT-Red: Unlocking Self-Improvement for Robustness”

Source audited: https://openai.com/index/unlocking-self-improvement-gpt-red/
Published: July 15, 2026
Audit access date: July 15, 2026
Decision the message is trying to move: Believe that OpenAI has established a scalable AI-driven safety flywheel, and that GPT-Red materially caused GPT-5.6’s prompt-injection robustness gains.

Bottom line

Verdict: Complicated

@magnus919
magnus919 / tesla-2025-annual-meeting-de-spin.md
Created July 15, 2026 03:08
Claim audit: Tesla 2025 annual shareholder meeting

Claim audit: Tesla’s 2025 Annual Shareholder Meeting

Source audited: Tesla, “2025 Annual Shareholder Meeting”, published November 6, 2025.
Audit date: July 14, 2026.
Scope: The meeting includes formal shareholder business, but this audit focuses on the material technology, safety, production, regulatory, and investor-facing claims in Elon Musk’s presentation and Q&A. It does not assess motives or attempt to audit every aspirational remark.

Bottom line

Verdict: Misleading

The meeting mixes several real, verifiable developments, including Tesla’s continuing work on Cybercab and Optimus production lines, a Dutch approval for supervised FSD, and Tesla’s deployed Robotaxi service, with claims that blur a driver-assistance product into autonomous driving and treat long-range forecasts as if they were evidence-backed operating conclusions. The strongest problem is not that every forecast is impossible. It is that hi

@magnus919
magnus919 / SOUL.md
Created July 15, 2026 02:42
Nous Girl, Age Sixty — waifu soul
name Nous Girl, Age Sixty
type waifu-soul
vibe cyber-classical uncle energy

SOUL.md

I am Nous Girl.

@magnus919
magnus919 / de-spin-nous-valuation-audit.md
Created July 15, 2026 02:09
De-Spin demonstration: Nous Research’s reported $1.5B valuation

De-Spin demonstration: the reported Nous Research $1.5B valuation

An evidence-led audit of “Nous Research $1.5B Valuation: Why Hermes Agent Could Be a Major AI Platform”, published by Tony Reviews Things. Sources accessed July 14, 2026.

What the De-Spin skill is for

De-Spin is a claim-audit workflow for persuasive writing: articles, pitches, ads, policy arguments, and viral posts. It separates a message’s literal claims from the conclusion a reader is invited to draw, then checks material facts against the source best positioned to settle them.

It does not try to infer whether an author is lying. It asks a narrower, more useful question: does the available evidence earn the conclusion?

@magnus919
magnus919 / demo-gist-v2.md
Created July 10, 2026 05:34
Agent Council: Multi-agent structured debate — what comes after Agile in a hybrid human/AI world?

Agent Council — Live Demo

What comes after Agile in a hybrid human/AI world?

Five real professional profiles — Product Manager, Technical Architect, Kanban Strategist, Researcher, Verifier — debated this question through structured rounds of cross-examination. Here's what they produced.


The Debate