Skip to content

Instantly share code, notes, and snippets.

View tonydzi's full-sized avatar

Anton Dziatkovskii tonydzi

View GitHub Profile
@tonydzi
tonydzi / bu_both.py
Created September 27, 2026 20:30
browser-use 0.13.10: read_file content lives in <read_state> for exactly one step; only the 1000-char long_term_memory persists. A/B harness + a custom action that survives.
"""browser-use 0.13.10 — где живёт содержимое файла: <read_state> (1 шаг) vs history (навсегда).
Harness fix: history items are only appended when model_output is not None
(message_manager/service.py "Build the history item"), so we pass a real AgentOutput.
"""
import asyncio, csv, os, tempfile
from browser_use import Tools
from browser_use.agent.message_manager.service import MessageManager
from browser_use.agent.views import ActionResult, AgentBrain, AgentOutput, AgentStepInfo
from browser_use.filesystem.file_system import FileSystem
@tonydzi
tonydzi / headroom_1099_repro.py
Created September 27, 2026 20:21
headroom #1099 repro: technique=preserve reports (1-N)*100% 'saved' because original_tokens counts the FIRST image and compressed_tokens sums ALL images
"""Repro for headroom #1099 — "negative savings" under technique=preserve.
Root cause (headroom 0.39.1, commit d13e196), headroom/image/compressor.py:
line 254 _extract_image_data() -> docstring: "Extract FIRST image data from messages"
line 530 _apply_compression() -> for "preserve", returns messages UNCHANGED
line 730 original_tokens = _estimate_tokens(<FIRST image>) + tile_saved
line 738 compressed_tokens = _count_result_tokens(<ALL images in ALL messages>)
The two counters measure different sets, so with N images in the transcript the
@tonydzi
tonydzi / probe.py
Created September 23, 2026 16:50
pydantic-ai #8551: where the 'no usage object' signal dies (models/openai.py L5171-5173), and whether RequestUsage.details can carry it
"""Where does pydantic-ai lose the "provider sent no usage object" signal?
Context: pydantic/pydantic-ai#8551. Measured on pydantic-ai-slim 2.48.0,
openai 3.19.0, opentelemetry-sdk 1.44.0, Python 3.12, macOS.
Part 1 hits the real production mapper (pydantic_ai.models.openai._map_usage)
with three responses; C is the positive control and must differ from A.
Part 2 asks whether RequestUsage.details -- a field that already exists --
survives accumulation into RunUsage.
Part 3 asks whether it reaches the OpenTelemetry span attributes unaided.
@tonydzi
tonydzi / probe_mem0.py
Created September 22, 2026 22:51
mem0 PR #6859: Azure indexing-result guard driven with nine result shapes (real SDK objects + mappings)
"""Run the guard from mem0 PR #6859 against edge-shaped indexing results.
The two methods are lifted verbatim out of the PR's own file (fetched from the branch),
not re-typed, so nothing about this harness can flatter or punish the result.
"""
import ast, subprocess, sys, textwrap
RAW = ("https://raw.githubusercontent.com/iroiro147/mem0/"
"PR_BRANCH/mem0/vector_stores/azure_ai_search.py")
@tonydzi
tonydzi / probe.py
Created September 22, 2026 22:47
pydantic-ai 2.47.0: three shapes of run-usage attribution (unspanned run / direct model request / no usage object)
"""Probe: three shapes of usage that a run span may or may not credit.
A (the issue): an UNSPANNED nested RUN -> parent span over-credits.
B (adjacent case raised in the thread): a model request made with NO run at all
(a tool calling the provider layer directly) -> who gets it?
C: a provider that returns no usage object -> what survives as a signal?
"""
import asyncio
@tonydzi
tonydzi / jcs_key_order_probe.py
Created September 21, 2026 22:46
RFC 8785 (JCS) key-ordering parity vector for payment claim keys - stripe/ai#402: UTF-16 vs code-point sort splits naive Python/Rust from JS and from the spec
"""Cross-language parity vector for RFC 8785 (JCS) as used as a payment claim key.
Context: stripe/ai#402 proposes deriving a duplicate-charge guard's identity from
the JCS canonicalization of the request, implemented independently in Python and
Rust and pinned to parity by cross-language tests.
RFC 8785 section 3.2.3 requires object members to be sorted on their **UTF-16
code units**. Python's `sorted()` on `str` sorts by **code point**. Those two
orders are identical for the whole BMP and then invert above it, because a
supplementary character encodes as a surrogate pair starting at U+D800, which is
@tonydzi
tonydzi / fork_parity_probe.py
Created September 21, 2026 22:44
claude-agent-sdk-python#1274 / #1208: live fork_session() title probes at head f7547d7 — ordering flaw is user-visible, and fork_session() vs fork_session_via_store() disagree
"""Parity probe: fork_session() (disk) vs fork_session_via_store() (store).
``_derive_title_from_entries`` documents itself as
"Precedence matches ``_extract_last_json_string_field`` semantics"
but it iterates PARSED dicts, so it is spelling-blind, while the disk path
scans RAW TEXT with a two-pattern loop that prefers the spaced spelling.
If the documented parity holds, both fork entry points must return the same
title for the same transcript. Same entries, both paths, one comparison.
@tonydzi
tonydzi / harness2.py
Created September 21, 2026 16:51
openai-agents 0.22.3: deterministic offline probe for prompt_cache_key + prefix on the nested Agent.as_tool() path (openai/openai-agents-python#5085)
"""Probe 2 for openai/openai-agents-python#5085.
Adds what probe 1 was missing: prompt caching hits on the shared PREFIX
(system instructions + leading input items), not on the key alone. So measure,
for consecutive calls to the same sub-agent:
- is system_instructions byte-identical?
- how many leading input items are byte-identical?
across four shapes:
@tonydzi
tonydzi / mem0_7211_bench.py
Created September 19, 2026 21:12
mem0ai/mem0#7211 — shared battery comparing candidate fixes for Azure AI Search indexing-result validation
"""GIT-S28 bench: mem0ai/mem0#7211 - compare candidate fixes for Azure AI Search
indexing-result validation on one shared battery. Run once per variant."""
import json, sys, traceback
from unittest.mock import MagicMock, Mock, patch
VARIANT = sys.argv[1]
from azure.search.documents._generated.models import IndexingResult as _IR
def make_obj(key, succeeded, status_code, error_message=None):
@tonydzi
tonydzi / surface_prompt_error.diff
Created September 15, 2026 21:18
claude-agent-sdk-python#1108: deterministic fake-transport repro (main 763922b vs PR #1109 e6cab50) + 4-line sketch
diff --git a/src/claude_agent_sdk/_internal/query.py b/src/claude_agent_sdk/_internal/query.py
index 2841adf..84e619a 100644
--- a/src/claude_agent_sdk/_internal/query.py
+++ b/src/claude_agent_sdk/_internal/query.py
@@ -171,6 +171,7 @@ class Query:
self._request_counter = 0
# Message stream
+ self._prompt_error: Exception | None = None
self._message_send, self._message_receive = anyio.create_memory_object_stream[