Skip to content

Instantly share code, notes, and snippets.

View arafatkatze's full-sized avatar
💭
Lover of life and all things beautiful

Ara arafatkatze

💭
Lover of life and all things beautiful
View GitHub Profile
@arafatkatze
arafatkatze / 00-system-prompt-index.md
Last active August 24, 2026 23:10
Cline IMO 2026 SDK-native candidate system prompts by problem — frozen public evidence

Candidate system prompts by problem

The system-role prompt was identical for all six problems. Each file below is an exact per-problem copy of that frozen prompt. The official problem statement was supplied separately as a user message and is intentionally not included here.

@arafatkatze
arafatkatze / sdk-native-max-opt8-best-observed-v1-20260824-r1-gist-public.md
Last active August 24, 2026 23:12
Cline IMO 2026 best-observed selected traces (validated, unequal-k)

Cline IMO 2026: best observed traces

This is a targeted adaptive, model-profiled best-of-unequal-k capability ceiling. It is not a replacement symmetric benchmark and it is not pass@k. Claude Fable 5 and GLM 5.2 optimization candidates are excluded; their frozen parent baselines remain unchanged.

Model Best observed score Selected generation cost Cumulative search spend
GPT-5.6 Sol 42/42 $3.2336 $3.2336
Claude Fable 5 41/42 $17.1956 $17.1956
Kimi K3 35/42 $5.1328 $19.5928
DeepSeek V4 Flash 30/42 $0.1215 $1.3122
@arafatkatze
arafatkatze / sdk-native-max-fast8-v1-20260808-r1-gist.md
Last active August 8, 2026 22:34
Cline IMO 2026 sdk-native-max-fast8-v1: 48 public traces, scores, costs, hashes, and validation

Cline IMO 2026 — complete sdk-native-max-fast8-v1 trace index

Public evidence for the frozen 8 models × 6 IMO 2026 problems = 48 cells benchmark run through the Cline SDK native-max protocol.

Headline

Model Score Submitted Total cost
GPT-5.6 Sol 42/42 6/6 $3.2336
Claude Fable 5 27/42 4/6 $32.2501
@arafatkatze
arafatkatze / eval-rsi-prompt.md
Created July 27, 2026 04:07
Recursive self-improvement prompt for agent evals

Recursive Self-Improvement Prompt for Agent Evals

You are conducting a long-running, evidence-driven improvement campaign for an AI agent harness.

Your goal is not to win a benchmark by any means necessary. Your goal is to make the harness genuinely better for the target model and then prove the improvement with reproducible evidence.

Objective

Improve the target agent on the named evaluation without changing the benchmark, adding task-specific logic, or manufacturing a better score through more resources. Optimize for the real frontier: capability, cost, reliability, and time to solve.

@arafatkatze
arafatkatze / README.md
Last active July 27, 2026 03:55
Cline Open Eval Ledger — 29 curated Terminal-Bench runs with raw traces, costs, tokens, and checksums

Cline Open Eval Ledger

This is the public index for the Terminal-Bench runs behind Cline's open-weight eval hill climb. It includes the selected model runs, exact job IDs, provider routes, scores, costs, token counts, Cline commits, checksums, and direct links to downloadable raw traces.

Start here

Terminal-Bench 2.1 Kimi K3 Recursive Self-Improvement: Traces & Proof

Companion artifact to the Cline blog post on one-shot recursive self-improvement with Kimi K3 on Terminal-Bench 2.1. The harness changes are in cline/cline#12465; the prompt that kicked off the campaign is here.

All 11: the baseline, four diagnostic slices, the Stage C candidate, three confirmation runs, and both interrupted runs invalidated. 4.52 GB total. The failures are in there with the same care as the wins.

Start here

You are conducting a long-running recursive agent-harness improvement experiment.

The central research question is:

Can Kimi K3 inspect and improve the Cline CLI agent harness that runs Kimi K3, then demonstrate a real, reproducible improvement on Terminal-Bench 2.1 without reward hacking, benchmark-specific patches, increased resources, or misleading experimental comparisons?

You are both:

  1. The engineering agent modifying Cline.
  2. The target model whose performance is being improved through those Cline changes.
#!/usr/bin/env bash
set -u
set -o pipefail
export LC_ALL=C
SCRIPT_NAME="$(basename "$0")"
HOSTNAME_VALUE="$(hostname 2>/dev/null || uname -n 2>/dev/null || echo unknown-host)"
TIMESTAMP_UTC="$(date -u +"%Y-%m-%dT%H:%M:%SZ" 2>/dev/null || echo unknown-time)"
{
"model": "claude-opus-4-6",
"max_tokens": 128000,
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "<system-reminder>\n\nUser system info (darwin 24.6.0)\nModel: Claude Opus 4.6\nToday's date: 2026-03-03\n\n# The commands below were executed at the start of all sessions to gather context about the environment.\n# You do not need to repeat them, unless you think the environment has changed.\n# Remember: They are not necessarily related to the current conversation, but may be useful for context.\n\n% pwd\n/Users/ara2/Desktop/cline\n\n% ls\nCHANGELOG.md\nCLAUDE.md\nCODE_OF_CONDUCT.md\nCONTRIBUTING.md\nLICENSE\nREADME.md\nSECURITY.md\nassets\nbiome.jsonc\nbuf.yaml\ncli\ncli-ts\ndist\ndist-standalone\ndocs\nesbuild.mjs\nevals\ngo.work.sum\nknip.json\nlocales\nnode_modules\nout\npackage-lock.json\npackage.json\nplaywright.config.ts\nproto\nscripts\nsrc\nstandalone\ntest-results\ntest-setup.js\
This file has been truncated, but you can view the full file.
<!DOCTYPE html>
<!-- saved from url=(0053)https://console.anthropic.com/workspaces/default/logs -->
<html class="h-screen antialiased [font-feature-settings:&#39;ss01&#39;] bg-bg-100 __variable_dcab32 __variable_820c23 __variable_e4195f" lang="en" data-theme="claude" data-mode="dark"><head><meta http-equiv="Content-Type" content="text/html; charset=UTF-8"><meta http-equiv="origin-trial" content="A7vZI3v+Gz7JfuRolKNM4Aff6zaGuT7X0mf3wtoZTnKv6497cVMnhy03KDqX7kBz/q/iidW7srW31oQbBt4VhgoAAACUeyJvcmlnaW4iOiJodHRwczovL3d3dy5nb29nbGUuY29tOjQ0MyIsImZlYXR1cmUiOiJEaXNhYmxlVGhpcmRQYXJ0eVN0b3JhZ2VQYXJ0aXRpb25pbmczIiwiZXhwaXJ5IjoxNzU3OTgwODAwLCJpc1N1YmRvbWFpbiI6dHJ1ZSwiaXNUaGlyZFBhcnR5Ijp0cnVlfQ=="><style>body {transition: opacity ease-in 0.2s; }
body[unresolved] {opacity: 0; display: block; overflow: hidden; position: relative; }
</style><meta name="viewport" content="width=device-width, initial-scale=1, maximum-scale=1, viewport-fit=cover"><link rel="stylesheet" href="./Anthropic Console2_files/026a2e7f617cf4f9.css" non