How the Grok Bot harness works - @gabefletcher
The Grok Bot host builds model requests, runs tools and saves the agent's state. Hosted services also handle identity, scheduling, sync and other features.
A message starts in the desktop app. The desktop sends it to the box host and shows the response. It also manages account connections and can run tools on the user's computer. The host builds the model's context, calls the provider, runs requested tools and saves the results. It then decides whether another model call is needed. The visible chat, model conversation and learned memory use separate stores. Changing the model provider leaves the other hosted service calls in place.
The original architecture sections cover public host 17e335e and desktop 0.52. The later comparison adds five-host source analysis, desktop .57.1 findings and isolated runtime measurements. Host and desktop have separate version numbers, and their exact pairing remains unproven. The available source is emitted JavaScript with source labels, not the original TypeScript or source maps.
S means the code shows the behavior. I means an inference. U means unresolved. Source IDs link to tables with source labels, functions, line and byte positions, and SHA-256 hashes. The original evidence tables describe source behavior. Runtime measurements are marked and scoped separately in the later comparison.
| Jump to | What it explains |
|---|---|
| Architecture and control flow | Six diagrams: components, user turn/model-tool loop, context/restart, peers/scheduling, startup/auth and cancellation; state and edge tables |
| Prompt behavior | Injection order, catalogs, completion, summaries, memory scopes, peer chatter and voice boundaries |
| Changes from .47 through .57.1 | Native request comparison, tool offloading, prompts, latency, retries, hidden work and accounting limits |
| Embedded source evidence | Artifact hashes, source labels, functions and exact byte positions |
A turn is one accepted unit of work. A step is one model-and-tool cycle within a turn. A checkpoint saves conversation and run state for later use. The transcript is the record shown in chat. Learned memory stores facts for future conversations. A summary carrier is the message that puts a summary into the model's next context. A gate is a runtime feature switch.
The box host runs the agent. The desktop handles the interface, communication between processes, account and gateway connections, and tools on the user's machine. A standalone eval program uses the same agent code through a separate entry point. The execution daemon runs tools. Blob-store workers handle saved conversation data. These are separate jobs. Their presence doesn't establish that each bot has its own VM.
The host behavior below refers to the source for 17e335e. Desktop nodes are labeled .52. Compatibility between those captures remains unresolved. Arrows show calls or contracts found in the code. They don't mark paths tested in a running system. [H10] and the other source IDs identify exact labels, functions, fragment hashes and byte positions. The captured host-main.cjs has SHA-256 f0e16cfa469d4b134735b020362d360aced9a3eaf729291c73f55111218cd1e4. the source index
flowchart TB
subgraph Desktop["Desktop 0.52 - separate captured product"]
UI["Renderer: chat, roster, permissions, context"]
IPC["Preload IPC and coordinator"]
Main["Electron main: accounts, gateways, OS integration"]
Local["Local computer execution daemon"]
UI --> IPC --> Main
Main --> Local
end
subgraph Host["Box host 17e335e"]
GW["Authenticated gateway and events"]
TM["Transcript manager and admission"]
RS["Per-agent scheduling and runner registry"]
Runner["SandAgentRunner and turn shell"]
Agent["AnysphereAgent model-tool loop"]
Ports["Injected inference, tools, memory and policies"]
GW --> TM --> RS --> Runner --> Agent
Runner --> Ports
end
subgraph Storage["Host storage - separate state domains"]
DB["Roster, transcript, admission and run records"]
Blobs["Conversation blobs, root checkpoints, summary archives"]
Mem["Agent memory and user/project shards"]
Files["Profiles, skills, automation definitions"]
end
subgraph Execution["Execution surfaces"]
Router["Loopback box adapter and window router"]
Daemon["Box execution daemon and tools"]
OS["Workspace, shell, browser and desktop"]
Router --> Daemon --> OS
end
subgraph Eval["Standalone eval executable"]
Request["Request file and eval main"]
EvalLoop["Shared runner with eval composition"]
EvalBox["LocalExecSandBox"]
Request --> EvalLoop --> EvalBox
end
subgraph Providers["Inference service boundary"]
Model["Shipped inference RPC transport"]
end
subgraph Hosted["Other hosted dependencies"]
Backend["Identity, migration, room membership"]
Cloud["Automation sync and event polling"]
Services["Search, media, voice, MCP, catalogs, labeling"]
end
Main -->|"gateway protocol; exact version pairing unproven"| GW
TM --> DB
Runner --> Blobs
Ports --> Mem
Ports --> Files
Ports --> Router
Ports -->|"registered machine routing"| Local
Ports --> Model
EvalLoop --> Model
Host --> Backend
Host --> Cloud
Ports --> Services
The host can run on a separate machine and still change data through hosted services. SAND_BOX_SESSION=1 disables one automation-sync path. It doesn't disconnect every hosted client. Selecting a user's computer changes where a tool runs. The main loop stays in the host. The chart assigns jobs to components. It doesn't establish enforced access controls between them. H01–H11 T01/T04 V01/V02 A03
| Component | Job and state it owns | Prompt behavior | Evidence / certainty |
|---|---|---|---|
| Renderer + preload/coordinator | Displays the interface. Checks and routes commands. Tracks roster and permission-screen state. The host owns conversation checkpoints. | Voice and UI channel instructions are separate from the host prompt. Main-bot selection assigns a roster role. | D03/D04; S. Compatibility between these exact versions is U. |
| Electron main | Manages account and backend context, the gateway connection, and optional OS, voice and network-routing adapters. | Limits voice transcript recall. A Temporal package can route send_task to the backend. |
D01; S |
| Gateway + send pipeline | Accepts API requests. Tracks retry IDs, input hashes, the displayed send and acceptance. The transcript manager dispatches the work. | The host adds context to the user message later. The displayed send isn't a model reply and may not have been saved durably. | H04/H05; S |
| Run scheduler + lifecycle | Tracks each agent's active run and user, agent and background queues. Handles cancellation and recovery. | Peer messages and routines get their own prompt context. They aren't presented as new human instructions. | H06/H07; S |
| Host composition + runner | Connects the profile, stores, inference, tools, permissions, MCP and transport. Run-generation checks reject old callbacks. | Builds the base, profile and memory prompt sections. Tracks parent and child request relationships. | H08–H11; S |
| Native agent loop | Manages history, step setup, model output, tool results, queued messages and compaction state. | Supplies tool schemas and reminders to the next model call. Post-step callbacks run after results are appended. | H12, T01; S |
| Checkpoint and transcript stores | Stores conversation roots, blobs and summary archives separately from the transcript mirror and journal. Metadata records completion and waiting-for-user state. | Saved summaries enter later model context. The UI transcript doesn't contain the whole prompt. | H13–H15, S01/S02; S |
| Blob-store worker pool | Runs store operations through Node worker threads and messages. Has explicit close and terminate operations. | The inspected pool doesn't own a bot prompt or model session. | H14; S |
| Learned memory | Stores agent files, user and project memory shards, background evidence and synthesis state. | Adds compact facts with their scope and source. privateMain wording doesn't enforce access control. |
M01–M06; S. Complete isolation is U. |
| Tool dispatch and permission adapters | Lists tools, selects their machine, applies local execution scope, handles approval and delivers results. | Some core tools stay in the static catalog. A gate enables the dynamic catalog. Task subagents run client-side. | T01–T04; S. Complete enforcement is U. |
| Background subagents + peer messaging | Tracks child runs and completion separately from peer bots' own contexts and incoming queues. | Sends completion and wake context to the parent. Peer prompts identify the assistant sender. The box gate disables conservative replies. | B01–B05; S |
| Inference extension | Checks credentials. Creates model sessions. Supplies separate web, audio and labeling helpers. | Uses separate purposes for main and summary sessions. Provider transport must preserve structured tool calls and results. | I01/I02, H10; S |
| Eval main | Reads a request file, creates the eval runner and writes results. Records updates, tools, deliveries and usage. | Accepts a system override, disabled tool IDs and simulated time. Its request context returns no rules. | V01/V02; S |
| Hosted identity/automation clients | Links bots to server identities. Runs backfill and migration. Manages remote shadow routines and consumes events. | Available services and gate settings affect catalogs, profiles and wake prompts. | I03/I04/I06, A01–A03; S |
flowchart TD
Input["UI send: agent, prompt, attachments, nonce"] --> Gateway["Gateway sendPrompt"]
Gateway --> Admit["Send pipeline: validate, dedupe, echo, acceptance"]
Admit --> Ack["Accepted response and events"]
Admit --> Queue["Run lifecycle and per-agent queue"]
Queue --> Compose["Create runner; load state; assemble action and tools"]
Compose --> Sessions["Main inference session plus summary session"]
Sessions --> Step["Agent step and model stream"]
Step --> ToolQ{"Tool calls?"}
ToolQ -->|"yes"| Gate["Resolve tool, validate and apply permission scope"]
Gate --> Exec["Execute local, box, MCP or hosted adapter"]
Exec --> Results["Append actual tool results and reminders"]
Results --> Checkpoint["Post-step state update and checkpoint promise"]
ToolQ -->|"no"| Checkpoint
Checkpoint --> More{"More work and no stop condition?"}
More -->|"yes"| Step
More -->|"no"| Settle["Drain checkpoints; final state and settlement"]
Exec -->|"SendToUser delivery"| Out["Transcript and client-visible message"]
Out -->|"end_turn when offered"| Stop["Request completion at checkpoint boundary"]
Stop --> More
Settle --> Memory["Memory evidence or extraction; labeling hooks"]
sendPrompt passes the request to the manager. The manager records acceptance and shows the sent message before the turn finishes. If a locked database drops the write, the pipeline can still accept the send. It logs that the input wasn't saved durably. An HTTP success or a visible message therefore doesn't always mean the input will survive a crash. The acceptance response and the durable write are separate events. H04/H05
createTurnRunShell copies the saved binary conversation structure. It prepares profile updates and updates to pinned prompt sections. It creates main and summary sessions, then calls agent.runStream. After a step, the native handler appends response and tool messages. It applies history changes, computes the next state and calls onStateUpdate, then afterStepCheckpoint.
In 17e335e, applyPostStepProcessing no longer takes isLastIteration. onStateUpdate no longer receives the third {betweenSteps} argument. The outer loop still calculates isLastIteration. The checkpoint persistence promise remains. That doesn't prove every caller waits for it or receives its errors. The fire-and-forget path catches and logs persistence errors. H10–H13
The message tool first calls the delivery callback. It then requests turn completion if end_turn:true is present and that feature is wired in. The shell handles completion, waiting for the user and upgrade pauses around checkpoint processing. The persistence path also checks pending tools. Reaching the step limit, having no tools or queued messages, cancellation and delivering a message are separate events. T02 H12 H10
Claude Code joke: a summary is where “one small change” goes to become the project history.
flowchart TB
subgraph Assembly["Model context assembly"]
Config["Base policy, profile, gates and capabilities"]
Frozen["Pinned sections and compaction epoch"]
UserInfo["Rules, skills and tool/subagent catalogs"]
Turn["New user action, attachments and update notes"]
History["Prior messages and tool results"]
Context["Ordered model messages and tool schemas"]
Config --> Frozen --> Context
UserInfo --> Context
Turn --> Context
History --> Context
end
Context --> Usage["Usage or overflow / summary trigger"]
Usage --> Pending["Summary candidate and original-prefix tracking"]
Pending --> Native["Generic self / external / provider-specific strategy"]
Native --> Valid{"Candidate valid and applicable?"}
Valid -->|"yes"| Carrier["Summary carrier, preserved tail and archive"]
Valid -->|"no"| Keep["Failure/retry path; preserve-context requirement"]
Carrier --> CP["Checkpoint root and blobs"]
CP --> Mirror["Transcript journal commit"]
CP --> Restart["Reload state; reconcile pending or stopped turn"]
Restart --> History
Carrier -->|"epoch advances"| Frozen
subgraph Learned["Learned memory is a separate store"]
Evidence["Turn evidence or legacy extraction"]
Synthesis["Synthesis plus verification when enabled"]
Facts["Agent facts and user/project shards"]
Evidence --> Synthesis --> Facts
end
Turn --> Evidence
Facts --> Frozen
Preserving context after a failed summary is a requirement. It hasn't been proved across every 17e335e runtime path. Before reusing a summary, the orchestrator compares its original message prefix with the current messages. Pending-summary code saves information about candidates for later use when they qualify. Generic self-summary keeps the system and user-info prefix and the latest user query. It adds the summary as a user-role message. The external summarizer also uses a user-role message, marked with cursor.isSummary. The selected strategy and extra saved context can differ. Provider-specific compaction uses separate implementations. The generic path doesn't establish their behavior. C01–C05 prompt details
turn-settle prepares the transcript mirror, writes the agent checkpoint and then commits the mirror with the root ID. If the checkpoint succeeds but the later transcript step fails, it raises TranscriptAppendAfterCheckpointError. Another branch skips transcript persistence under its privacy policy. The code orders these operations. It doesn't prove they all share one transaction.
The shell queues checkpoint callbacks through checkpointChain. It rejects callbacks from old run generations, records progress and commits automation-completion state. After a restart, saved completion or awaiting-user IDs can prevent a stopped turn from running again when no tool calls remain pending. Startup also resumes ownership. Accepted-input recovery and routine scheduling are separate jobs. H13/H15 H10/H16
| State domain | Owner / storage interface | Restart implication and limit |
|---|---|---|
| Bot identity, profile and settings | Session/profile services and agent directories | The same bot ID must survive worker replacement. Hosted identity links are separate state. |
| Accepted inputs, transcript, run/ack records | Transcript manager, admission ledger and local store | The displayed send, acceptance, settlement and model context have separate records. A branch accepts input without saving it durably. |
| Native conversation | Agent store root metadata, blob store and summary archives | Recovery needs history, tool state and summary messages. A provider conversation ID alone can't restore them. |
| Transcript mirror | Prepare/commit journal around checkpoint persistence | Partial failures need reconciliation. One transaction across all databases hasn't been proved. |
| Memory | FileMemoryStore, UserMemoryStore, ProjectMemoryStore, synthesis evidence | Stored separately from compaction. Shared recall needs explicit scope and permissions. |
| In-flight summary | Native pending/background state and adoption interfaces | Summary eligibility, cancellation and lifetime need tests on the exact combined build. |
| Peer queue and dedupe | Upstream pendingAgentInbound Map and received-ID Set in the inspected dispatcher |
These in-memory fields don't prove messages survive restart. |
| Routine definitions versus active queue | Automation store/cloud-sync versus transcript scheduler | Disabling the transcript scheduler to suppress hosted cron can break recovery of accepted input. |
SandSessionConversationState reads transcript entries from session.db. It builds the outline from session.agentStore.getFullConversation. The chat transcript and native conversation structure use different read paths. S01–S04 H05/H06 H13–H16 M02/M04/M05 C04 B02 A01/A03
flowchart TD
Human["Human input"] --> UQ["User lane"]
Peer["Another bot: SendToAgent"] --> Route{"Local target or remote peer?"}
Route -->|"local"| Inbox["Inbound queue, priority partition, dedupe"]
Route -->|"remote"| Remote["Hosted remote-agent messaging client"]
Inbox --> AQ["Agent lane and wake prompt"]
Routines["Routine definition and trigger"] --> Authority{"Scheduling authority"}
Authority -->|"captured hosted paths"| CloudSync["Shadow sync and event polling"]
CloudSync --> BQ["Background lane"]
UQ --> Scheduler["Per-agent scheduler"]
AQ --> Scheduler
BQ --> Scheduler
Scheduler --> Parent["Target bot runner with its own context"]
Parent --> Task["Task or background subagent dispatch"]
Task --> Child["Child runner, lineage and scoped tools"]
Child --> Completion["Completion metadata and parent revival"]
Completion --> Parent
Inbox -.->|"priority can interrupt background lane"| Parent
SandRunScheduler has user, agent and background queues. Within the user queue, it prefers items that aren't group-member work. Hosted automation uses another scheduling path. AgentToAgentMessaging routes messages to local bots, groups or remote peers. It stores a bounded set of received IDs to detect duplicates. It orders priority messages, batches queued messages and wakes the target runner. A priority peer message can interrupt background work. The prompt identifies the sender as another assistant. Priority doesn't grant authority over the user. Group posting has separate membership checks and differs from one-to-one priority delivery. H06/H07 B02 I05
Task configuration sets requireServerSideSubagent:false and useClientSideSubagent:true. The runtime tracks subagent starts, stalls, steering, cancellation, settlement and completion back to the parent. Cursor cloud agents are a separate optional feature.
Peer prompts have a conservativeReplies option. The dispatcher gets its value from reducePeerChatter. The box host sets that gate to false. The option exists in shared code, but this capture doesn't enable it for the box. Copying the prompt alone wouldn't enable it. T01 B01–B05
Hosted automation still reconciles local definitions with remote schedules. isBoxHostedForAutomationSync excludes session boxes from that path only. When the local definitions directory is missing, remote shadow schedules that aren't wanted locally remain untouched. Empty or invalid definitions take different paths. Deleting local files doesn't establish that remote schedules were deleted. A01–A03
flowchart TD
Main["runHostMain: lock, environment and crash handling"] --> Host["SandHost.start"]
Host --> Ext["Dependency-ordered extension startup"]
Ext --> Auth["Auth service and credential readiness"]
Ext --> Stores["Stores, sessions and transcript recovery"]
Ext --> Bind["Bind runner composition into turn execution"]
Bind --> Resume["Resume ownership at startup"]
Host --> Gateway["Gateway auth token, APIs and event bridges"]
Auth --> Ready{"Mock enabled or access token present?"}
Ready --> Infer["Shipped inference readiness"]
Ext --> Identity["Hosted mint, backfill and resweep"]
Identity --> Migration["Server binding, group adoption and room migration"]
Ext --> Automations["Hosted automation polling and sync"]
Ext --> Other["Capabilities, telemetry and other service clients"]
Starting the host can resume saved work and contact services. The registry has 54 entries. Extension startup follows dependencies. The entry count doesn't establish which features are active. Inference readiness checks for the shipped auth token unless mock mode is configured. Web, audio and post-turn helpers have separate backend calls.
Identity registration also runs separately. noteAgentMinted returns its serialized operation. Failures can request another sweep. A sweep marker older than 24 hours allows another sweep. That freshness check doesn't prove a timer runs every 24 hours. Group adoption and room migration still depend on server identity bindings. H01–H03 I01–I04/I06 H04
flowchart LR
UI["UI interrupt with agent/session identity"] --> Run["Run lifecycle and runner cancellation"]
Run --> Stream["Inference request context / stream abort"]
Run --> Tool["Active tool context"]
Tool --> Router["Window router upstream request"]
Router --> Daemon["Execution daemon handler child context"]
Daemon -.->|"must be measured"| Child["Subprocess termination and no late side effects"]
Run --> Drain["Checkpoint drain and settlement"]
The 17e335e window router destroys its upstream request if the caller disconnects before the response finishes. It destroys the caller response if the upstream response errors. Separately, the .52 desktop daemon cancels its ControlledExec child context during cleanup and when no handler exists. These changes belong to two different captures.
Cancellation crosses several layers: the visible turn, provider request, daemon context and OS process tree. Socket closure alone does not prove that a child process stopped. D02/D05
The evidence table gives an artifact or fragment hash for each source ID below. Each certainty label applies to the inspected call. Behavior under failure needs separate tests.
| Edge | Actual caller / callee or contract | Data passed | Certainty and unresolved part |
|---|---|---|---|
| UI → desktop bridge | Preload commands, source/roster APIs and interruptAgentRun; D03/D04 |
Agent/session IDs, UI requests | S; exact desktop↔17e335e pairing U |
| Gateway → manager | sendPrompt: async (args) → manager.sendPrompt; H04 |
Prompt, attachments, nonce, routing options | S; successful acceptance isn't turn completion |
| Admission → ledger | sendPromptOnce → recordPending, emitSendAck; H05 |
Nonce/digest and echo ID, durable flag | S; locked-database non-durable branch |
| Lifecycle → scheduler | Run lifecycle + SandRunScheduler.enqueue; H06/H07 |
Per-agent tasks and lane | S; queue state alone isn't durable recovery |
| Composition → runner | createRunner → new SandAgentRunner; H08/H09 |
Injected stores, permissions and service ports | S |
| Shell → model/summary sessions | runTurn → inference.createSession; H10 |
Main lineage versus summarization purpose | S; provider-account context isolation U for this build |
| Composition → agent | buildAgentForRun → new AnysphereAgent; H11 |
System generator, tool generator, summary handlers, hooks | S |
| Step → checkpoint hook | applyPostStepProcessing → onStateUpdate(ctx,currentState) → afterStepCheckpoint; H12 |
Updated native state and persistence promise | S; no old third metadata argument |
| Shell → serialized persistence | checkpointChain → persistStepCheckpoint; H10/H13 |
State checked against the current run generation | S; no runtime crash/cancel proof here |
| Settle → mirror/store | prepareCheckpoint → handleCheckpoint → commitCheckpoint; H13/H15 |
Blobs/root ID and transcript journal | S. A single transaction across stores isn't proved. |
| Store → worker | AgentWorkerPool.ensure → Node Worker, blob request messages; H14 |
Blob-store operations | S; thread isn't a VM or model worker |
| Assembly → system prompt | renderSystemPrompt, named push calls; P01 |
Ordered sections and SHA observations | S; gates determine active sections |
| Action → new user message | assembleTurnAction; P02 |
Address, reply, attachments, update and reminder notes | S; XML-style reminder tags don't change API role |
| Catalog → later turn | deferUserInfoCatalogRerender, update helpers; H11/P03/P04 |
Dynamic tool and subagent catalog changes | S. Catalog freshness after compaction is U. |
| Tools → machine executor | createMachineRoutedTool, buildTurnTools; T01/T04 |
Explicit machineId versus box target |
S. Not every permission path was tested. |
| Send tool → completion | onSendMessage then completeTurnAfterSend; T02 |
Delivered message and stop request | S; summary-drain lifecycle not proven here |
| Turn → memory | runTurnMemory: evidence branch or extraction; M03/M06 |
Bounded exchange evidence | S; not always legacy extraction |
| Synthesis → memory apply | runAgent → proposal, verification, applySynthesis; M04 |
Cited evidence, snapshot, changes | S. The code handles stale snapshots. Complete security is U. |
| Explicit memory → shard | writeMemory → memoryShardFor; M05 |
Agent/user/project target; project membership | S; privateMain alone no isolation proof |
| Summary → context carrier | buildSummaryMessage/assembleFinalMessages; C01/C02 |
User-role summary, retained recent messages and extra saved context | S; provider-specific paths distinct |
| Background result → adoption | handleSummarization, takePendingSummaryForAdoption; C03/C04 |
Whether the message prefix still matches, and the candidate's state | S. Lifetime and cancellation tests on this exact build are U. |
| Peer → wake | sendToAgent/receiveFromPeer → queueInboundAndWake → runAgentInboundWake; B02 |
Sender identity, text/images, priority | S. In-memory duplicate detection doesn't prove mailbox durability. |
| Subagent → parent | settleBackgroundSubagentTurn, revival prompt; B03/B05 |
Completion/status and return context | S; optional capabilities/gates vary |
| Identity → hosted backend | noteAgentMinted, syncDownAgent, backfill; I03/I04/I06 |
Server identity, group binding and sweep state | S; hosted service dependency |
| Automation → hosted authority | Cloud-sync reconcile, fire consumer, session-box exclusion; A01–A03 | Definitions, shadow state and events | S; session-box flag is narrow |
| Voice → hosted harness | Imt generated session/tool RPCs; D01 |
Backend/account context and voice request | S. Separate from the main inference interface. |
| Egress → backend policy | PIe → getGrokBotEgressPolicy; D01 |
Domains/CIDR/ports | S; optional network policy authority, not inference |
| Eval → shared runner | main → runSandEvalRequest → createSandEvalRunner; V01/V02 |
Request file, override, tool disables, result file | S. No normal-host launch call was found. The host bundle lacks the literal sand-eval-runner.cjs. |
| Open question | Current evidence limit |
|---|---|
Desktop .52 ↔ public host 17e335e |
Independently captured versions; no proven pairing or interoperable boot |
| Actual tool cancellation | Router/socket and daemon-context cleanup observed; stopping the OS child process tree and preventing late writes remain unproven |
| Startup and hosted services | Minting, migration, polling, sync and capability discovery include hosted calls |
| Summary settlement/restart | Callback signatures and promise order are visible in source; complete settlement and restart behavior remains unverified |
| Memory API and privacy | Legacy UI command absent in new host; privateMain text can't certify enforced scope/isolation |
| Tool freshness and limits | Catalog deferral is gated; MCP clamp fragments are initialization only, so no numeric cap inferred |
| Peer durability and fairness | The dispatcher uses an in-memory queue and Set; cross-host delivery and restart durability remain unresolved |
| Provider semantics | Source alone does not establish service-side context transformation, limits, auth renewal or billing classification |
The following sections paraphrase what the prompt code asks the model to do. Model compliance, active gates and production wiring need separate evidence. The source reference is 17e335e unless a passage names desktop behavior. The full prompts aren't reproduced. hashes and exact anchors
createSystemPromptAssembly.renderSystemPrompt appends each non-null section in the order below. The table uses the section names from the code. P01
| Order | Sections | Behavior / assumptions |
|---|---|---|
| 1 | base, optional spotlight, optional voice_channel |
Base variant depends on gates, host capabilities and role. A supplied system override replaces the base; it doesn't automatically erase all later sections. Voice is excluded for ordinary subagent runners. |
| 2 | profile, user_identity, team_bot, multitask |
Includes saved name and profile text and role instructions. Subagents and parent-mediated automation children receive different amounts of the parent's context. The less-fanout policy can favor reusing an executor. |
| 3 | internal_details_boundary, mcp_multi_account, user_form, draft_external_message, timezone |
Gates select interaction instructions. Internal-details guidance can discourage disclosure of architecture. The prompt doesn't enforce access control. |
| 4 | memory, current_session, active_sessions |
Includes learned facts and session context. Memory appears in this order: user facts, project facts, agent facts. Each has its own recall limit and source information. |
| 5 | automations, skills, stripe_link_purchasing, channels |
Describes available features and their use. The skill-based path can replace long instructions with skill and tool discovery. Content varies by role and available features. |
| 6 | agent_directory, group_chat_turns, mcp_instructions, mcp_status, remote_box, bot_secrets, computer |
Includes the peer roster, connector guidance and status, and execution environment. A fresh bot can see roster names in its prompt even when conversation histories are separate. |
The assembler hashes each section to track changes to the prompt prefix. It pins some sections until the next compaction epoch, when the conversation is summarized again. A stable prefix helps preserve the provider's prompt cache. Live changes can appear as notes in a new turn. Compaction then refreshes the pinned sections. If user memory becomes unavailable, the host invalidates a pin that still contains it. Profile changes are announced and later folded into the profile section. An edit doesn't necessarily rewrite the system message immediately. P01
assembleTurnAction builds the new user message separately. Address, machine and reply context come before the user's text. Attachment notes follow. The code then attaches automation-status, profile and instruction updates, followed by reply and unfinished-task reminders. With isSilenceAllowed, each attachNote call prepends its update. That reverses the update order compared with appending. Hidden wakes get hidden and untrusted handling. Images and video can use structured context. Recent messages, shell-watch material and unanswered or dismissed-question context are prepended separately. P02
The native agent starts with a user-info message for rules, skills and catalogs. In 17e335e, deferUserInfoCatalogRerender uses the stableDynamicToolCatalog gate. When enabled, changes to tool namespaces, subagent types and subagent models arrive on a new user turn. This preserves the initial user-info prefix for caching. The update helper checks the latest update for each catalog before checking the original version. It tells the model which sections the new catalog replaces. Request-context recovery, summaries and team-rule changes retain separate rerender flags. A system_reminder tag doesn't change the message's API role. H11 P03/P04/P05
The model must use the new note when a pinned section is stale. The required tests include tool removal, permission changes during a turn, catalog changes after compaction and both gate settings. A cache hit doesn't prove the model received the current catalog.
| Component | Noteworthy behavior | What still needs checking | Source / status |
|---|---|---|---|
| Base host persona | Concise desktop-assistant behavior, real work and factual reporting; uses tool-based delivery, task continuity, approvals, security and environment guidance | Text can't enforce tool scope; visible behavior depends on tool availability and stop policy | P06; S |
| Profile and settings | Name guidance; profile changes through state tools preserve unspecified fields; hiding a sidebar row doesn't stop routines or erase context | UI visibility isn't lifecycle control; cached identity needs explicit updates | P01; S |
| User-info / catalogs | Describes available tools, skills, rules, subagent types/models; gated updates replace named catalog sections | A catalog entry isn't evidence of implemented or authorized execution | P03–P05, T01; S |
| Toolset | State update, memory recall and reactions are forced static; dynamic registry depends on parent parity, box scope and gate. machineId routes to a registered computer; absent it, tools target the box |
Model must not invent a target. Execute-time validation/permissions must reject stale or unadvertised requests | T01; S |
| Search/fetch descriptions | Missing/stale search results and a blocked fetch aren't proof a fact/page is absent; suggests alternate retrieval paths | Fallback browser or shell access still depends on available tools and permissions | T01; S |
| Local machine / daemon | Tool descriptions explain execution surface, approval and missing capability; .52 missing computer/recording handlers give nonretryable guidance |
Useful error text doesn't prove a tool succeeded or process cancellation works. | T01/T04, D02; S |
| Message delivery and completion | The prompt directs user-visible replies through SendToUser. When available, end_turn signals completion. Another base variant asks for a short assistant completion after the final send. |
A successful send, finished step, completed native run and fulfilled business task are different states | P06, T02, H10/H12; S |
| Checkpoint/transcript | No independent persona prompt. Saved summaries and history enter later context. The transcript mirror supplies saved chat content and extra context. | UI history isn't a faithful reconstruction of every hidden/tool/system message | H13/H15; S |
| Eval composition | Can override the system prompt, disable tool IDs, supply selected images/history and simulated time; resolves empty rules and collects tool/update/delivery events | An eval success would validate only the configured fixture. Eval main consumes a request file and requires its own auth environment; it isn't an offline safe import | V01/V02; S |
| Providers | Receive main or helper requests through the executor interface. Summarization has a separate purpose. | Separate logical sessions must map to separate provider transport identities where stateful. A new provider string isn't sufficient isolation | H10, I01/I02; S |
| Startup/auth/hosted extensions | No unified standalone prompt. They decide readiness, capabilities, metadata and catalogs that the prompt sees | Their service calls aren't all inference. Turning off one catalog or sync flag doesn't remove hosted authority | H03, I01–I04, A03; S |
runTurnMemory first checks for recordMemoryEvidence. If available, it records the exchange, clears pending episode turns and returns. Otherwise it runs the older extraction helper. That path records episode turns and sometimes generates an episode narrative. Extraction uses a separate system/user request with relevant existing facts. It parses and applies the returned changes. That request updates learned facts. Native compaction separately reduces the conversation sent to the model. M01/M03
MemoryService.createAgentStore can record evidence for background synthesis. The sand_memory_dreaming gate enables synthesis. A separate gate, grok_bot_server_memories, enables hosted shard synchronization. These switches control different behavior. The synthesis extension creates summarization sessions with labeling skipped. M02/M06
Synthesis asks for small changes with cited evidence. It treats existing state and conversation evidence as untrusted data. It separates lasting profile facts from time-bound logs, rejects speculation and protects explicit user-authored entries. Superseding a fact requires supporting evidence. A second model step reviews the proposal. Code checks schemas and evidence references, then applies changes against the saved snapshot. A stale snapshot causes another pass. Protection of explicit entries and allowed scopes requires code enforcement. Model review alone doesn't enforce those rules. M04
Agent, shared user and project memory have separate interfaces. writeMemory chooses a shard. User scope writes to that bot's user-memory shard. Project scope requires an existing project and bot membership. Readers can merge shards while retaining attribution and applying size limits. When scope is omitted, sand-state-tool uses the conversation default, then falls back to agent scope. M02/M05 T03
privateMain changes the prompt's ownership wording. In the relevant renderer, it changes the default user-tier wording to owner. The state-tool description cache also includes it. These sites change text and cached descriptions. Storage isolation requires separate checks of shard readers and writers, read authorization, shared-memory access and backend sync. The word private in a prompt doesn't enforce access control. M01 T03
Desktop .52 removes legacy memory commands. getAgentMemories is absent from both e6e1fb3 and 17e335e, while learned-memory code remains. An absent UI command does not mean the memory subsystem is gone.
Generic self-summary and the external summarizer request a technical handoff. They cover user intent, relevant facts and code, errors and fixes, solved and open problems, pending tasks, current work and a supported next step. Their instructions include chronological analysis of the supplied history. The task is to compress that history. Quoted tool output doesn't gain authority to issue instructions. Native code can also retain task, transcript and skill context alongside the summary. C01/C02/C05
Generic self-summary sends the system and user-info prefix, the history to summarize and a summary request appended as a user message. The next context keeps that prefix and the latest user query. It adds a user-role summary message with continuation guidance and a summary count. The external path also uses a user-role summary. It can preserve the latest image when configured. These carriers retain their user role. Anthropic, OpenAI and xAI native compaction use separate implementations. Their carriers need separate inspection and tests. C01/C02
The orchestrator chooses background or foreground summarization. It checks the original message prefix and whether a pending result can be used. This work interacts with checkpoint persistence and turn settlement. C03/C04 H10–H13
Claude Code joke: give every agent a mailbox and eventually one of them replies-all.
A task subagent gets model and type options from its configuration. It can have a role-specific prompt, inherit parent request lineage and receive fewer tools or profile sections. Parent-mediated automation children omit selected user-facing instructions. The runtime separately records completion, errors, stalls and the information used to wake the parent. A persistent peer bot has its own identity and conversation, so it needs a separate delivery contract. T01 H11 B03/B05 P06
Peer wake prompts identify the sender as another assistant. They say the message arrived asynchronously and is visible in chat. A batch includes sender identities and priority markings. Priority wording asks for earlier action. The conservative branch explains that replies arrive on later turns and discourages needless replies. The captured box host sets reducePeerChatter to false, so that branch isn't an active chatter fix here. B01/B02/B04
Limits on bot reply chains need enforcement in the mailbox and scheduler. Required controls include duplicate detection, deadlines, hop and reply budgets, and visible status when an outcome is uncertain. Separate model calls can still create separate messages with repeated meaning. A sender's name doesn't grant permission.
Desktop .52 voice reads recent human and sent-message entries through transcript-tail pages. It can read up to 16 pages of 256 entries to find eight relevant chat messages. Its instructions forbid invented work and inbox items. UI guidance routes replies to the call channel, avoids duplicate acknowledgements and returns unfinished work to chat after hangup. These voice changes don't establish changes to ordinary compaction or learned-memory renewal. D01/D04
Imt sends voice session and tool requests through generated GrokBotService RPCs with backend and account context. A package harness using a Temporal loop sends send_task to the backend. Egress separately fetches destination policy. The main-bot selector reads and writes hosted settings. Replacing the main inference transport leaves those calls in place. D01
Generated email, room and voice schemas define API messages. They don't contain the server implementation or prove the UI calls every operation. clamp-advertised-tools.js is 333 bytes. clamp-direct-mcp-tools.js is 668 bytes. They retain initialization, logging, counters and provider priorities. Neither fragment contains an active truncation function. Another module could still enforce limits. Voice, hosted email and rooms, egress and connector catalogs have service dependencies beyond the ordinary inference stream. T05/T06
Added September 19, 2026. This section combines source comparisons with isolated runtime measurements. R marks measured runtime behavior under the stated fixture; S remains source evidence and U remains unresolved production behavior. The earlier architecture sections retain their original .52/17e335e scope.
We investigated the report that newer agents feel faster while consuming much more usage. We compared five recovered host artifacts, inspected desktop .57.1, executed extracted production functions, and ran actual stock old/new hosts against an isolated protocol-compatible backend.
The strongest correction to our initial theory: the newer host sent less data for an identical minimal task. Larger prompts exist, but “everything got bigger” does not explain this system.
Each host received . twice in a fresh synthetic conversation. Its real gateway, runner, serializer, SendToUser tool, persistence and nonce settlement executed. A private backend returned deterministic ok and the memory extractor's NONE sentinel. No live model or real account was used.
| Boundary | .47 cold | Latest cold | .47 warm | Latest warm |
|---|---|---|---|---|
| Main framed-protobuf bytes | 167,854 | 118,769 | 168,638 | 119,555 |
| Messages | 3 | 3 | 6 | 6 |
| Tool schemas | 40 | 11 | 40 | 11 |
| Decoded message JSON bytes | 49,811 | 59,056 | 50,914 | 60,189 |
| Decoded tool JSON bytes | 116,115 | 57,770 | 116,115 | 57,770 |
| Separate memory-request bytes | 1,909 | 1,909 | 1,909 | 1,909 |
The latest cold system text was larger: 51,180 versus 48,437 UTF-8 bytes. User-info/context grew to 6,849 from 378 bytes. But fewer upfront tool schemas outweighed that increase. Two complete turns sent 242,142 bytes latest versus 340,310 old, about 28.8% less.
Both selected grok-4.5, max mode, high effort and fast mode. Both made four inference requests for two inputs: answer, memory extraction, answer, memory extraction. Memory inference occurred after the visible reply and before native settlement. That hidden work is real, but already existed in .47.
The tool reduction has a concrete explanation: dynamic-tool offloading changes from compiled default-off to default-on by 95b13b6. The architecture already existed. Thirty formerly advertised tools are still accepted as unadvertised native tools; GetDynamicTools/CallDynamicTool expose discovery and invocation. They were not removed. The old 378-byte context is an identical prefix of the new one; the 6,471-byte addition is subagent/tool catalogs and separators, not more personal user information. The broad registry audit also found 117 false→true defaults between 17e335e and 95b13b6 across a shared multiproduct registry. These are not 117 proven active Grok Bot features. Caller and effective account configuration still matter.
Bytes are not tokenizer counts. Decoded JSON and encoded protobuf are different representations and must not be added together. The backend may transform inputs or apply caching and tariffs we cannot see.
We then changed only the supported dynamic-tools flag on each same binary, retaining all other fixture responses and explicit inputs:
| Mode | .47 cold main bytes / tools | Latest cold main bytes / tools |
|---|---|---|
| Static tool schemas | 167,854 /40 | 209,810 /49 |
| Dynamic tool schemas | 108,933 /10 | 118,769 /11 |
This intervention proves that offloading causes the initial schema reduction in the fixture. Latest carries roughly 25% more bytes than old in static mode and 9% more in dynamic mode, but its default switches to the cheaper initial mode. The default flip masks same-mode growth. No discovery ran in the minimal turns, so later discovery overhead and whole-task cost remain unmeasured. Dynamic tools do not necessarily require an extra round for every invocation.
The recovered .47 closure lacked piscina. Its successful control borrows captured newer dependencies read-only after the original path. It is a mixed-dependency laboratory control, not the complete original shipped distribution. Both use synthetic present-empty capabilities and default experiment responses, not the user's account gates. Public host d0de163 has not been proven paired with installed desktop .57.1.
The actual base-prompt builder was executed with matched options and optional cloud, voice, canvas and device sections disabled. These values measure one base section, not the complete model input.
| Recovered host | Full base, UTF-8 bytes | Skillified base, UTF-8 bytes |
|---|---|---|
| .47 | 26,626 | 22,319 |
| e6e1fb3 | 28,255 | 23,887 |
| 17e335e | 28,409 | 24,041 |
| 95b13b6 | 30,165 | 25,652 |
| d0de163 | 30,192 | 25,679 |
The base grows about 13.4% or 15.1% across the full comparison, but just 27 bytes between the last two captures under these options. Whole-system text grows differently because the assembled request includes other sections, some relocated or removed. This is why the base-prompt table and the native request table answer different questions. Neither supplies a tokenizer or billing measurement.
The changes point toward more persistent, delegated and event-driven agents: better replay, dynamic tools/context, cloud-agent watches, task tracking, browser recovery and work resumed without another typed message. This is an architectural inference, not a developer statement about their roadmap.
| First observed artifact | Concrete changes |
|---|---|
| Already in .47 | Parallel workers; post-turn memory extraction; stable prompt/schema gates; MCP discovery deadline; browser prewarming; retry/continuation policies |
| e6e1fb3 | xAI compaction includes original system/tools; summary/compaction wire tags; streamed-tool omission recovery; awaited rule refresh; browser readiness and stale-seat recovery; email review surface; image memoization |
| 17e335e | Conditional catalog deferral preserves earlier user-info prefix and appends changes |
| 95b13b6 | Dynamic tools/stable catalog compiled defaults become true; transcript recall; incoming provider ID/phase retention; forced-manual review can skip classification; failed-secret hidden wake; unsupported-image conversion option |
| d0de163 | Gated cloud watches; eligible voice-card copy inference; broader snapshot fallback; stronger delegation wording relative to .47 |
| Desktop .57.1 | Gated incremental task-list request, watch/task-card surfaces and conditional SCM routing |
First observed means the first capture containing the change, not an exact release/deployment date. Public host revisions are not verified desktop release pairings.
Compaction. Old XaiCompactionHandler.generateSummary sends a small dedicated system message and no tools. From e6e1fb3, it includes the original system and catalog, with compactor instructions appended as a user message. Executing the actual method with a synthetic 20 KB system and ten schemas grew captured JSON from 4,585 to 27,398 bytes. This proves construction changed; it is not a sixfold real-world billing estimate. History, selected handler and cache treatment matter.
Transcript recall. ReadTranscript, first observed in 95b13b6, nominally budgets 48,000 bytes but counts JavaScript string length and exempts the first line. Part truncation does not bound the number of parts. Actual production formatting accepted 402,644 characters in one multipart message; a Unicode case admitted 132,725 UTF-8 bytes. The communicate renderer retained it. This is conditional amplification when recall is invoked, not automatic replay. A recursive single-result test plateaued at 7,587 characters rather than growing indefinitely.
Browser recovery. Latest scoped snapshot fallback can remove both target and depth before retrying. A narrow request can become a whole-page snapshot. Existing downstream spill remains; no universal cap bypass was demonstrated.
Delegation and events. Instructions now more explicitly suggest delegation beyond two tool rounds. Parallelism itself is old. More workers could finish faster while spending more across all workers, but actual user-workload fanout is missing. Enabled cloud watches can revive parents; failed secret storage can queue a hidden turn; eligible voice receipts can invoke a Gemini copy-generation stream. New email actions expand classification surfaces. These need separate attribution.
Task tracking. Desktop .57.1's task keeper sends incremental bounded transcript text through preload/main to a service. Its gate defaults off. Backend inference and billing are unknown. A method or endpoint name alone is not another proven model call.
Later code preserves streamed tool calls omitted from final responses, retains provider IDs/phases on replay, checks browser readiness, and recovers stale desktop assignments. Conditional catalog handling can preserve an earlier prompt prefix rather than rebuilding it. That may improve cache reuse while appending more context. Stable hashes alone do not prove actual cache hits.
Meanwhile, team-rule caching became synchronous on more requests. .47 could return an existing snapshot immediately and refresh in the background after five minutes. Newer code awaits loading even with a snapshot. Each request-context executor reuses its own promise; another newly built executor can wait again. Membership and rule RPCs have individual ten-second timeouts, not one universal end-to-end ceiling. Actual function tests reproduce this blocking change; its live latency contribution remains unknown.
Two more defaults matter. Auto-review moves from shadow to enforce: default-settings users already caused classification in the old mode, but the new mode waits for its result. This can add critical-path latency without adding a classifier call. Disabled settings require an actual backend team lock plus admin enforcement to be overridden; a default flag alone is insufficient.
Post-output idle recovery becomes enabled by the compiled fallback. Its existing 90-second deadline suspends during a pending tool. A silent-stream expiry cancels the attempt; ordinary policy resumes after visible output only with a persisted checkpoint. The retry budget itself did not expand. This may make recovery happen more often even though the retry function is unchanged. Actual account flags and stall frequency remain unknown.
flowchart TD
Input[Desktop input] --> Route[Preload and coordinator route]
Route --> Queue[Host session and turn queue]
Queue --> Context[Rules skills identity memory catalogs]
Context --> Prompt[System user information and history]
Prompt --> Wire[Serializer and inference backend]
Wire --> Action[Text or tool call]
Action --> Tool[Approval and execution]
Tool --> History[Retained model history]
History --> Wire
Action --> Reply[Visible SendToUser reply]
Reply --> Memory[Observed memory extraction]
Memory --> Settle[Native settlement and persistence]
Settle --> Events[Conditional later events or worker revival]
Events --> Queue
Native control timings, milliseconds from driver dispatch:
| Stage | .47 cold/warm | Latest cold/warm |
|---|---|---|
| First model request | 292 / 147 | 280 / 122 |
| Native visible reply | 318 / 190 | 303 / 138 |
| Native settlement | 512 / 262 | 458 / 278 |
Startup to observed gateway health was 1,264 ms old and 1,053 ms latest. One pair on an already running WSL system is not a benchmark distribution. There is no actual model latency here. The measurements separate local work and show why a visible answer is not the end of processing. Native transcript timestamps corrected misleading polling lag in early observations.
- Larger retry budget/higher default effort: local retry allowances, fallback grok-4.5/high/fast/max-mode and step ceiling did not expand. Idle-trigger enablement did change, so unchanged budgets do not mean identical operational retry frequency. Account overrides remain possible.
- Cached input added twice: actual stream processing keeps the supplied 500 input plus 20 output at 520, not 920 after adding cache-read 400 again. The desktop percentage/cents transforms are structurally equivalent; server accounting is unavailable.
- Huge classifier prompt on every call: the cooked literal grew 61,737→73,581 bytes, but its identifier occurs only at declaration in each inspected host. Stock classification sends a backend RPC without it. The backend's own prompt remains unknown.
- Every hidden call is new: memory extraction and the former voice-close hidden reply already invoked models. The newer voice-close path replaces existing work.
- All output caps grew: shell formatting and generic MCP spilling are byte-identical across five captures. Some old no-writer/first-item exceptions admit large data, but are not new universal growth.
- Idle sampling means periodic inference: the helper preempts an existing silent sample; the recovered Grok Bot provider supplies no enabling policy on the inspected path.
Exact source hashes and byte anchors accompany differential tests. Independent reruns reproduced 49 loop/transport checks, 35 census assertions, 13 context result groups and 15 cold-start cases. Separate suites cover 46 prompt/compaction/memory checks and 48 catalog checks. These are different scopes, not one whole-product acceptance score.
The final enablement audit added 95 independently rerun review-mode controls and 31 watchdog/gate groups. Same-binary native static/dynamic interventions were audited against raw reports, request events, exact environment differences and immutable fixture hashes.
Twelve bounded TypeSafe calls reviewed 36 questions twice with reversed order, zero retries: 97,996 research input tokens and 3,179 output. Source-support probabilities were 0.93–0.95 for xAI construction and 0.98 for team-rule control flow. Numerical attribution of the user's bill received 1.0 insufficient evidence. Two deep-review expectations received stable abstentions rather than anticipated contradiction; two final gate questions changed verdict with order. All packets, frozen labels and outcomes are preserved. No re-query manufactured agreement. These are one model's source judgments, not probabilities that a change caused the user's increase. Native interventions, not model agreement, establish the scoped tool-placement mechanism.
Original captures remain immutable. Failed startup and fixture errors remain labeled alongside clean results. Full synthetic payloads remain private local evidence; this draft reports mechanisms, sizes and hashes rather than dumping prompts. These measurements do not establish deployment or daily-use acceptance.
We have demonstrated cost mechanisms and rejected several tempting explanations. Attribution still needs the actual hosted revision and gates, matched real tool-heavy workload, parent/child/summary/memory/review/keeper IDs, provider fresh/cache/output/reasoning usage, and the displayed meter for the same interval. Disable one contributor at a time in an isolated profile and compare that workload.
Useful candidate fixes are a whole-result UTF-8 recall budget, bounded snapshot fallback, measured rule-cache freshness, and per-purpose cost/timing attribution. Compaction needs compatibility and summary-quality testing before removing context or schemas.
The defensible result: newer code changes orchestration, replay and conditional context handling; it does not simply multiply every request. The minimal native control got smaller. Particular long or tool-heavy paths can still become much more expensive. Account/backend evidence is needed to identify which explains this user's experience.
All native hosts ran in fresh isolated network namespaces with private protocol fixtures. The model loop, prompt assembly, tool execution and settlement came from the pinned stock host. Generated codecs were exported only in a separate fixture process. The fixtures emitted no provider usage totals. Present-empty identity capabilities, empty experiment responses and fresh state exposed compiled defaults; these inputs differ from a real account. The flag intervention used the supported development override and changed no host bytes.
Latest attempts with a missing invocation ID or an incorrect memory response were preserved and excluded from the clean comparison. The first .47 attempt failed during codec initialization before host startup. The corrected .47 control uses read-only borrowed dependencies, including piscina 4.9.3. The clean comparison and two flag interventions all settled successfully, preserved source closures, reaped their child processes and left empty namespace listener tables. Each measured cell contains one cold/warm pair, so no population-level timing estimate is claimed.
The first scope inventory missed the generated inference-client package. A broader inventory corrected that gap. It covers 2,052 source labels across five hosts and records long string candidates, callers and gates. A long literal may be a schema or unused code; it is not automatically an active prompt. Source presence, reachable caller, configured default and live account enablement remain separate claims.
| Artifact | Full SHA-256 |
|---|---|
| 047 host-main.cjs | 12f0ad99e5f5c729fe848db20087c8c7056ac6182d5c72bef56795c351a9164d |
| 051-public host-main.cjs | 3eef7508f9b30722128ecb5ad4e2739177e69466c57c2d983a95303f630610d9 |
| 052-public host-main.cjs | f0e16cfa469d4b134735b020362d360aced9a3eaf729291c73f55111218cd1e4 |
| 055-public host-main.cjs | bb69e8c7eb0d8d54f1f14dabf115e595aed4ed19c5036fe0f5fd9533b8d96b37 |
| latest-public host-main.cjs | 3ec376e8f6cf68746bcdf3b1bfe22dd71c037ae6b1ea65865f08e9c7ffa9c425 |
| Desktop .57.1 app.asar | 9308f4b1972ad4a8c5dfec27f7b49acf0537a5eba9f1ba5c6d4ed51483feb38f |
Public-host labels identify captures, not verified desktop release pairings. Byte intervals below are UTF-8, zero-based and half-open. Fragment SHA-256 values identify the exact source; they do not attest runtime enablement.
| Evidence ID | Host capture | Byte interval | Fragment SHA-256 |
|---|---|---|---|
R57-01: CG-readSandTranscriptWindow |
latest-public | [25702885, 25704450) | 5a7bd17256dbd822295f8987bbbcf4f13a254f7cdab857195ab58c514a17aad8 |
R57-02: CG-makeRender |
latest-public | [24024005, 24024471) | 258db8b500e2682f6d1c238164de061a7f7cc84edc031d2ac68a4039fef40e58 |
R57-03: CG-withPlaywrightSnapshotFallback |
latest-public | [25307358, 25308523) | 2760ba56698f094e884c461a803ce99a8db03a09ab7356a81e7cd3bd1bb531fa |
R57-04: LT-latest-public-stream-request |
latest-public | [20592729, 20595214) | 2190d6b9436a2c8eb3e975249e9b19c7b020b3620393fe58c7883a139f4a9be6 |
R57-05: LT-latest-public-apply-summary-metadata |
latest-public | [20581142, 20581548) | e1b6000a48c75e76dd6e89f2e1b5ee1a2adf5db773a58f76e07bedc165b37f4e |
R57-06: LT-latest-public-proto-response |
latest-public | [20614187, 20616182) | aa1d7ad535ed3f9976a5b794c8a7c9efc2002b7127742a46bea9f189040a6450 |
R57-07: IC-latest-public-hidden-resume |
latest-public | [24708720, 24709947) | afaa6f24826984aabea9be53935b9eeaae3697db539936e552f81c04e5096ef4 |
R57-08: IC-latest-public-classifier-rpc |
latest-public | [20618801, 20619839) | 166db57c4dfab4f20287c623fd4949c65b968839c6674dba82c1503f9e35c4ee |
R57-09: 047.rules-resolver |
047 | [25993826, 25995164) | 179631024042421c828ed3641577273db3057c7f18ff3133fe14f373b2858bfb |
R57-10: latest-public.rules-resolver |
latest-public | [22985547, 22986533) | ff31ba4f231ab4716b8f301d3da2b27eeee0eaa5f29b29829693e8a65f72f2b9 |
R57-11: latest-public.request-context-executor |
latest-public | [25309580, 25311226) | 0dc688c86c1d6d28886803c4585b0ff7b1a6e727c6862b4a0a3a47583d8cd3cd |
R57-12: latest-public.rules-network |
latest-public | [22986533, 22989257) | 9f96460b88c44aed49f0e5a7a78edee701dc319a4c6b8c2c65e6e339ba8dd86b |
R57-13: 047.grok_bot_dynamic_tools |
047 | [25214713, 25214781) | 69073a2e53a571c0af92db126ef18563727728cada875f5c053f3fc3a5d515b8 |
R57-14: 055-public.grok_bot_dynamic_tools |
055-public | [22144951, 22145018) | 966a812272dedc371d8e461181d481a0c1f346f8aee21dd86e75a5bf57f1030a |
R57-15: latest-public.grok_bot_stable_dynamic_tool_catalog |
latest-public | [22274358, 22274439) | 0950ad4f501505b36afbac04af67bee951159163aad335a7700e34c85ba17cfa |
R57-16: latest-public.sand_stream_idle_deadline |
latest-public | [22299361, 22299431) | cc0d5d9bfe185849c6a48ed3edb6fbb703c6131a403395b3c076caf73e93d8b0 |
R57-17: latest-public.sand_auto_review |
latest-public | [22335606, 22335667) | 235e6e1d91daa121f10ba3601d0208ba1ee852c4024a487194d7fc3b666ed350 |
The original architecture evidence follows unchanged. It refers to its stated older capture, while the comparison above names the newer source and measured fixture explicitly.
Full artifact hashes appear below. Each H/P/T/M/C/B/A/I/V row also identifies the exact emitted fragment by hash. Labels ending in .ts are bundle comments, not original TypeScript. Byte offsets are zero-based UTF-8. Lines are one-based. Repeated or minified matches identify source locations; they do not by themselves establish calls between functions.
| Captured artifact / input | Bytes | SHA-256 |
|---|---|---|
harness-0.52/host-capture/17e335e/unpacked/sand-host/host-main.cjs (harness-0.52/host-capture/17e335e/unpacked/sand-host/host-main.cjs, source path) |
25993397 | f0e16cfa469d4b134735b020362d360aced9a3eaf729291c73f55111218cd1e4 |
harness-0.52/host-capture/17e335e/unpacked/sand-host/sand-eval-runner.cjs (harness-0.52/host-capture/17e335e/unpacked/sand-host/sand-eval-runner.cjs, source path) |
18168057 | 74212b2604c20631036f33359b4ddf3f187c81ea10e8ca1cf31c595d22db6dbd |
harness-0.52/desktop-app/unpacked/dist/electron-main/main-app.cjs (harness-0.52/desktop-app/unpacked/dist/electron-main/main-app.cjs, source path) |
2032310 | c77d0fa5d3aabe60ecb21d3bb8fc2196cc33dae21fe56c1597e9edcdff12f57b |
harness-0.52/desktop-app/unpacked/dist/local-exec-daemon/main.cjs (harness-0.52/desktop-app/unpacked/dist/local-exec-daemon/main.cjs, source path) |
3280342 | 04680fd64cabca3a5ee1cc688b7f55af487be18c7ad761e535229dc3618101b8 |
harness-0.52/desktop-app/unpacked/dist/electron-preload/preload.cjs (harness-0.52/desktop-app/unpacked/dist/electron-preload/preload.cjs, source path) |
75503 | 28ce0669e3c2a1ea67d82ddffdcaec5db8e437ef9f6096fbe6eef82b797ffa7e |
harness-0.52/desktop-app/unpacked/dist/renderer/assets/index-e6gX-VaT.js (harness-0.52/desktop-app/unpacked/dist/renderer/assets/index-e6gX-VaT.js, source path) |
3739221 | 34772bfe93f15b90382fb474b00e440a6620675f64f306d85cdf586768442a09 |
harness-0.52/host-capture/17e335e/unpacked/sand-host/box-scripts/sand-window-router.mjs (harness-0.52/host-capture/17e335e/unpacked/sand-host/box-scripts/sand-window-router.mjs, source path) |
4357 | 27ec65c33f58ad68a76d738aa373e23056387bd737c3abd027f7dfeec97b0f8b |