Skip to content

Instantly share code, notes, and snippets.

@dhaupin
Last active July 28, 2026 19:40
Show Gist options
  • Select an option

  • Save dhaupin/6b9d33ae054b6aa60a033fae36fd06e4 to your computer and use it in GitHub Desktop.

Select an option

Save dhaupin/6b9d33ae054b6aa60a033fae36fd06e4 to your computer and use it in GitHub Desktop.
Vant cache hit analysis
To effectively utilize prompt caching and prevent massive token inflation on MiniMax, you need to intentionally design your prompt architecture.
The underlying mechanisms are identical for both MiniMax 2.5 and MiniMax 2.7. Both models leverage an [Anthropic-compatible API endpoint](https://platform.minimax.io/docs/api-reference/anthropic-api-compatible-cache) (/v1/messages), meaning they rely on explicit prefix matching using strict byte-by-byte rules. [1, 2, 3, 4]
The structural rules to prevent breaking the cache across requests include the following:
## 1. Enforce a Strict Top-Down Structure
MiniMax constructs its cache strictly from left to right in a fixed sequence: Tool Definitions → System Prompts → Early User Messages → Dynamic Context. [5]
*
* The Rule: Any change to a single character at the top of the prompt invalidates everything below it.
* The Fix: Push all fluid variables—such as timestamps, rotating user names, conversation histories, and tool output logs—to the absolute bottom of the payload. Keep the system instructions entirely static. [3, 5, 6, 7]
*
## 2. Inject Explicit Cache Control Blocks
Unlike systems that auto-cache text based on sheer frequency, MiniMax requires explicit cache signaling if you are routing through their Anthropic-compatible protocol. [1, 2, 8]
*
* Use the "cache_control": {"type": "ephemeral"} parameter on large, stable text blocks. [9]
* Place a checkpoint right after your massive codebase files or reference documentation, but before the user's specific multi-turn code request. [9, 10, 11]
* Limit: MiniMax supports a maximum of 4 cache breakpoints per call. Do not over-segment your prompt; instead, stick to 1 or 2 high-value blocks for your large codebase files. [12]
*
## 3. Maintain Exact Byte-Level Consistency
Because MiniMax relies on exact string comparisons for cache hits, the slightest format shift will trigger a full, expensive cache miss. [3, 13]
*
* Tool Shuffling: Ensure tool lists and available function schemas are passed in the exact same array order every time.
* Whitespace & Newlines: Be careful with templating languages (like Jinja or f-strings) that automatically strip or add unexpected trailing spaces and double-newlines. [3, 6]
*
## 4. Meet the Threshold and TTL Rules
*
* Minimum Size: Caching only triggers if the prompt prefix contains 512 or more tokens. It will not save small, lightweight instructions.
* Time-to-Live (TTL): The ephemeral cache layer expires after 5 minutes of inactivity. For long-running background developer loops or high-frequency autonomous agent sessions, this works beautifully. However, if your agent waits 10 minutes between iterations, it will trigger a full, costly cache write every single time. [3, 5, 12, 14, 15]
*
## The Big 2.5 vs. 2.7 Difference
While prompt structuring is identical, MiniMax 2.7 introduced automatic fallback caching on specific base endpoints, meaning it handles standard chat history caching slightly more leniently than 2.5. However, developer teams on [GitHub](https://github.com/anomalyco/opencode/issues/15338) have noted a specific bug where using complex tools drops the cache hit rate to 15% on both models. For heavy multi-agent coding workloads, relying on explicit cache_control blocks in your system payload remains the safest mechanism for both generations. [1, 2, 16, 17]
Are you running these agents through an integration framework like LiteLLM or LangChain, or are you writing the raw API payloads yourself? I can give you a clean JSON block showing exactly where to place the cache breakpoints.
[1] [https://platform.minimax.io](https://platform.minimax.io/docs/api-reference/anthropic-api-compatible-cache)
[2] [https://github.com](https://github.com/NousResearch/hermes-agent/issues/17332)
[3] [https://medium.com](https://medium.com/@sthanikamsanthosh1994/the-memory-hack-why-prompt-caching-is-the-secret-to-scalable-ai-52d2538bff98)
[4] [https://medium.com](https://medium.com/@srivatsa4123/prompt-caching-is-a-layout-discipline-not-a-feature-flag-0684c48feed1)
[5] [https://platform.minimax.io](https://platform.minimax.io/docs/api-reference/text-prompt-caching)
[6] [https://blog.dailydoseofds.com](https://blog.dailydoseofds.com/p/prompt-caching-in-llms)
[7] [https://analyticsindiamag.com](https://analyticsindiamag.com/ai-news/claude-code-is-built-around-prompt-caching)
[8] [https://platform.minimax.io](https://platform.minimax.io/docs/api-reference/anthropic-api-compatible-cache)
[9] [https://platform.minimax.io](https://platform.minimax.io/docs/api-reference/anthropic-api-compatible-cache)
[10] [https://docs.litellm.ai](https://docs.litellm.ai/docs/tutorials/prompt_caching)
[11] [https://caylent.com](https://caylent.com/blog/prompt-caching-saving-time-and-money-in-llm-applications)
[12] [https://platform.minimax.io](https://platform.minimax.io/docs/api-reference/anthropic-api-compatible-cache)
[13] [https://redis.io](https://redis.io/docs/latest/develop/use-cases/semantic-cache/)
[14] [https://platform.minimax.io](https://platform.minimax.io/docs/api-reference/anthropic-api-compatible-cache)
[15] [https://caylent.com](https://caylent.com/blog/prompt-caching-saving-time-and-money-in-llm-applications)
[16] [https://thomas-wiegold.com](https://thomas-wiegold.com/blog/minimax-m-2-7-review-is-it-worth-the-hype/)
[17] [https://github.com](https://github.com/anomalyco/opencode/issues/15338)
------------------------------------------------
Running MiniMax through [dhaupin/vant](https://github.com/dhaupin/vant/issues) requires precise prompt engineering because Vant handles file-based persistence, passive memory, and decentralized storage through Git. [1, 2, 3]
Instead of traditional top-down string layouts, Vant leverages Vant Project Brain files (identity.md, lessons.md, etc.). To make MiniMax 2.5/2.7 cash-efficient without destroying Vant's decentralized architecture, apply these adjustments: [4]
## 1. Pin the identity.md and lessons.md Core Manifests First
Because MiniMax maps its prompt cache sequentially from left to right, your static model identities must form the foundation.
*
* The Strategy: Configure Vant's initialization to force the core files (identity.md, lessons.md) to append to the system prompt in the exact same array sequence every time.
* Why: If Vant dynamically changes file order during an index refresh, it will completely break the Anthropic-compatible byte-level prefix match required by MiniMax.
*
## 2. Isolate Vant's Entropy-Patch Compression Blocks
Vant uses an Entropy-Patch Protocol (vant compress) to optimize raw brain updates and index tracking passively. [5]
*
* The Strategy: Treat the Entropy-Patch compressed text as a separate user block or a tool definition segment. Wrap this block inside an explicit cache_control marker block:
{
"role": "user",
"content": [
{
"type": "text",
"text": "[Vant Brain Index Updates Log...]",
"cache_control": {"type": "ephemeral"}
}
]
}
* Why: This tells MiniMax's API to build a checkpoint immediately after processing the file system modifications. This prevents re-parsing the core project files on subsequent turns.
*
## 3. Move the Active Repository State to the Absolute Bottom
Vant operates as a Git-native framework to pull file state, diff logs, and file pointers. [6]
*
* The Strategy: Keep dynamic runtime text—like the exact line-by-line file changes (git diff), current local tasks, and active code executions—at the very end of your active array context.
* Why: Any character change at the top destroys everything downstream. By isolating the file manipulation logs at the end, your large system blocks can successfully achieve hit rates over 90%.
*
## 4. Align Vant's Runtime Expiry with the 5-Minute TTL Rule
MiniMax cache blocks automatically expire after 5 minutes of total inactivity.
*
* The Strategy: For long-running or periodic cron tasks managed by Vant's internal automation engine, implement a "warm-up" loop or keep active generation tasks grouped together within tight 4-minute intervals.
* Why: Spacing out your runtime loops to run every 10 minutes forces MiniMax to execute a full cache write from scratch on every run, rendering your optimization ineffective.
*
Are you planning to deploy Vant directly via its Docker environment or use its Model Context Protocol (MCP) server integration? I can provide the setup configurations needed to inject the cache_control headers into those specific connections. [4, 5]
[1] [https://github.com](https://github.com/topics/persistance?l=javascript&o=desc&s=)
[2] [https://github.com](https://github.com/dhaupin/vant/blob/main/README.md)
[3] [https://github.com](https://github.com/dhaupin/vant/releases)
[4] [https://github.com](https://github.com/dhaupin/vant/releases)
[5] [https://github.com](https://github.com/dhaupin/vant/releases)
[6] [https://www.reddit.com](https://www.reddit.com/r/aiagents/comments/1v15dhn/i_stopped_building_a_database_for_my_ai_agents/)
------------------------------------------------
Yes, theoretically [Vant](https://www.google.com/search?q=vant&kgmid=/g/11bwg7mhyh#sv=CBwS1gMKpQMSogMK4gJBSmlUNHRKUTBNUHpqUTlnMDJXdDlfaWFlYlZ0WDV6eEtTUzZqZDdCNzBxaXFlOGpnZUhFZHhYbUFndzBxMHpVbkQtLXFYMlpUMnhsQWtSVWxWUDdFVGJxZl9FTzFrVjIzVWpZZnRvc1RoXzVPVHczMzJWWll6aVhBaDlyZ0R5UmRaUDZtMk8zMDhOcUNZc3R4TnRMMXRxX1hFcjNkbnEycWxyQjVTUzJKTkdOaUNyOGlZTnptY243SEFIQkxLSEptdXR6ZFVjNjFGaFJaOHRhNHd6Nk9vR3d1VkZfbUhjWEItWnJCVGJxcHFmSUZkaGYxa3J6UmFMREtPYWRYbENNOW1yLUtGT0RDQmM2c2RlYlNTa2sxTVMyekpqMFBtSTNVY1l2SDlUcmFLZFpHX3RrRk90bkRwMHZ0WVlaa2J2c0RtUHhwVFZ4WHJhSkNoZk04NXZ4Vk80ZmVvbG9jbHFaVncSF2JRUnBhdXZEQ3Q3WTVOb1AxZi1jbUE4GiJBRHNyOWZRNWZEY3YwaTFRcEY1TEhQZDlnN2N2Ynl5TzlnEgQ3ODU0GgEzIgkKAXESBHZhbnQiFgoFa2dtaWQSDS9nLzExYndnN21oeWgoABhFIJeMhfcI) could handle this completely if you pull request (PR) the necessary modifications.
Because Vant is built around highly structured configuration arrays and clear repository boundaries, patching it to support MiniMax's strict byte-level prefix caching is entirely realistic. You don't need to rebuild Vant's core engine; you just need to alter how it maps arrays and interacts with upstream APIs.
If you were to open a PR to optimize Vant for MiniMax 2.5/2.7 token efficiency, your implementation would need to tackle three specific components:
## 1. The Context-Assembly Matrix (Array Sorting)
Vant natively builds its prompt payload dynamically. To maintain cache-friendliness, you would need to patch the prompt assembler to force strict left-to-right deterministic array sorting.
*
* The PR Change: Write an interceptor in Vant's LLM context composer that runs a rigid sorting function on incoming text fragments before sending them to the provider.
* The Code Logic: Ensure that the immutable core files (identity.md, lessons.md) always occupy the first indexes of the text array, followed by the tool schemas, followed by the historical log, and ending with the highly erratic git diff entropy fragments. If a tool definition or file index randomly changes order by a single position, your cache hit rate will drop to zero.
*
## 2. The Native cache_control Injector Block
If you are passing MiniMax through Vant's Anthropic-compatible routing or an MCP middleware layer, Vant needs to know where to drop ephemeral breakpoints.
*
* The PR Change: Add a parsing step in Vant's configuration file (e.g., inside vant.config or the model provider logic) that allows users to flag key project boundaries.
* The Code Logic: Have Vant look for a demarcation point—such as immediately after it appends the codebase context but right before it loads the immediate shell execution logs. The code should explicitly inject "cache_control": {"type": "ephemeral"} at that index. Because MiniMax caps this at 4 breakpoints, your code would need a simple safeguard counter to prevent Vant from injecting a checkpoint onto every micro-file it reads.
*
## 3. Automated "Heartbeat" Loops for the 5-Minute TTL
Because MiniMax drops cached data after 5 minutes of total user inactivity, background background agents running intermittent tasks will routinely trigger expensive cold starts.
*
* The PR Change: Create a low-overhead, optional daemon flag in Vant (e.g., vant up --keep-cached).
* The Code Logic: If the flag is active, Vant can execute a tiny, lightweight "ping" query every 4 minutes utilizing the exact same top-level static system block. This simple loop resets MiniMax's internal expiration timer, guaranteeing that when you finally execute a massive, multi-step development sequence, the codebase layer is already fully warm and instantly accessible.
*
## Where to Draft the PR
If you are diving into the repository, focus your changes on the Model Provider abstraction layer (where payloads are compiled into JSON string bodies) and the Context Window trunking logic (where Vant decides how to truncate old project memory).
Would you like to draft a raw Python/TypeScript middleware helper or an adapter that you can test locally in your environment before packaging it into a formal Vant PR?
------------------------------------------------
Integrating directly with [Vant's /lib](https://github.com/dhaupin/vant/issues) directory is exactly how you would have to build this. Because Vant architecture treats the repository as a prototype agent operating system and file-based database, caching optimization can't just be an API wrapper. It has to sit right inside the core file-parsing and memory-persistence system. [1, 2, 3]
Targeting Vant’s internal mechanics to hit that 90%+ MiniMax cache match requires addressing specific structural areas:
## 1. Intercepting the Git-Native Context Compiler
Vant heavily relies on Git mechanisms to manage state, track histories, and compute diffs for long-running agent sessions. [1, 4]
*
* The Problem: Git diffs and branch histories change dynamically based on file modifications. If Vant merges these changing logs directly into the system or early instruction arrays, it will instantly invalidate the MiniMax prefix cache. [4]
* The /lib Fix: Inside Vant's core model or context-generation libraries, you would need to modify the array-building function. You must explicitly separate static Git-tracked files (like identity.md or persistent memory states) from the active branch mutations and commit histories. The static snapshots must be structured left-to-right at the front of the text array, while dynamic code edits are pushed strictly to the final entries. [1]
*
## 2. Matching Vant's Entropy Compression with Ephemeral Checkpoints
Vant utilizes custom tracking and compression updates to optimize passive brain growth over time. [1]
*
* The Problem: The MiniMax Anthropic-compatible API requires explicit cache_control flags on its block arrays to register breakpoints. Standard strings won't trigger it automatically.
* The /lib Fix: You would add an option within Vant's internal model provider module. When Vant compiles the compressed passive memory arrays, it should automatically map those text blocks into an object schema containing the "cache_control": {"type": "ephemeral"} header. By embedding this inside Vant's provider logic, the system handles the byte-level tracking natively without the user writing custom JSON headers.
*
## 3. Leveraging Vant's Cron/Runtime for the 5-Minute Heartbeat
MiniMax drops its ephemeral cache if there is no activity for 5 minutes.
*
* The Problem: If a Vant autonomous session pauses to wait for disk I/O, slow network loops, or developer feedback, the cache will expire and force an expensive re-read on the next turn.
* The /lib Fix: Because Vant functions like a localized agent OS with active runtimes, you can create a background process directly inside its core daemon code. If a MiniMax flag is active, Vant can schedule an internal 4-minute timer block. If no active user or file query has fired, it can execute a featherweight ping containing just the static project definition array to keep MiniMax's server-side cache warm. [1]
*
## The Verdict on PR Viability
This approach aligns well with Vant’s design philosophy. Vant was specifically built because standard frameworks fail at managing branching histories and ad-hoc JSON memory states cleanly. By patching Vant's /lib array sorting to support exact byte-matching caching, you enhance its capability as a cost-efficient, long-running agent host. [1, 4]
Would you like to pinpoint the exact JavaScript/Node.js files within Vant's provider abstraction to write a prototype snippet for the array sorter?
[1] [https://github.com](https://github.com/dhaupin/vant/releases)
[2] [https://www.reddit.com](https://www.reddit.com/r/indie_startups/comments/1uph72j/lets_talk_projects/)
[3] [https://github.com](https://github.com/topics/persistance?l=javascript&o=desc&s=)
[4] [https://www.reddit.com](https://www.reddit.com/r/aiagents/comments/1v15dhn/i_stopped_building_a_database_for_my_ai_agents/)
------------------------------------------------
Integrating the changes into Vant’s nova branch is a highly strategic choice, especially given that its active development focuses on multi-brain routing, cross-brain state sync, and advanced provider abstraction.
Because the nova branch re-architects how Vant loads distinct memory modules ("islands" and separate brains), a generic prompt string wrapper would break. The cache handling logic must live inside Vant's underlying core system compiler.
To implement this on the nova branch, focus on modifying two key architectural layers within /lib and /models:
## 1. The Context Orchestrator (/lib/brain.js or /lib/context.js)
In the nova paradigm, this module builds the multi-brain payload array by reading identity.md, fetching passive memory logs, and loading the target file system data. You must intercept the builder array here to enforce byte-perfect ordering.
// Prototype array sorter to inject into Vant's context compilerfunction prepareDeterministicCachePayload(systemBlock, brainFiles, conversationalHistory, dynamicDiffs) {
const payloadContent = [];
// 1. Static Rules & Model Personas (Must always be absolute first index)
payloadContent.push({
type: "text",
text: `System Persona:\n${systemBlock}\nIdentity/Lessons:\n${brainFiles.identity}\n${brainFiles.lessons}`
});
// 2. Active Tools or Multi-Brain Router Configurations
// Ensure keys/tools are sorted alphabetically to prevent schema shifting
const sortedTools = stableAlphabeticalSort(brainFiles.availableTools || []);
payloadContent.push({
type: "text",
text: `Available Agent Tools:\n${JSON.stringify(sortedTools)}`
});
// 3. Set the First Cache Breakpoint (The 512+ token boundary rule)
// We mark the massive static memory payload block right before history drops in
payloadContent[payloadContent.length - 1].cache_control = { type: "ephemeral" };
// 4. Stable Chat History Array (Early history stays cached)
conversationalHistory.forEach((msg, idx) => {
const messageBlock = { type: "text", text: `${msg.role}: ${msg.content}` };
// Optionally cache long stable message histories up to MiniMax's 4-checkpoint limit
if (idx === conversationalHistory.length - 2 && idx > 5) {
messageBlock.cache_control = { type: "ephemeral" };
}
payloadContent.push(messageBlock);
});
// 5. High-Entropy Git Diffs, Active Shell Logs, and File Output Mutations
// This shifts constantly, so it is pushed to the bottom of the payload matrix
payloadContent.push({
type: "text",
text: `Active Runtime Mutations:\n${dynamicDiffs}`
});
return payloadContent;
}
## 2. The Model Driver Wrapper (/lib/models/private/ or /lib/providers/)
Vant routes its API calls via generalized formats. For MiniMax 2.5/2.7 execution, the payload compiler needs to format objects to support block arrays natively rather than standard string formatting.
*
* The Fix: Find where Vant formats payloads for the OpenAI API vs. the Anthropic API endpoint. You need to verify that when the model is recognized as a MiniMax variant (minimax/abab...), Vant maps the array structure using the Anthropic message schema configuration.
* The Dynamic Code Hook:
// Inject into Vant's model routing libraryif (config.model.startsWith('minimax')) {
// Direct call routing to MiniMax's Anthropic-compatible /v1/messages endpoint
apiPayload.messages = prepareDeterministicCachePayload(
systemPrompt,
currentBrainState,
sessionHistory,
activeGitDiff
);
}
*
## 3. Adapting to the Multi-Brain Routing
Because the nova updates allow Vant to step through different "brains" dynamically depending on the task, a single fixed cache won't work.
*
* If Vant swaps from a "Coding Brain" to a "Documentation Brain", the entire MiniMax prefix changes.
* The Strategy: The PR should store the cache status per Brain Target ID. When Vant switches target contexts, it must recognize that it is performing a cold-start cache write. It can then run a localized multi-step execution loop to prevent bouncing back and forth rapidly, which would continually wipe MiniMax's short 5-minute memory cache window.
*
If you want to review the exact logic handling payload structures in the nova branch, point me to the path handling context synthesis inside /lib (e.g., lib/brain.js), and we can write the precise patch for it.
------------------------------------------------
In Vant's nova branch structure, the files handling context compilation and model execution are clearly separated into distinct engine layers within /lib.
To apply these caching optimizations successfully, you will need to patch three primary targets in Vant's Node.js architecture:
## Where to Implement in nova/lib## 1. lib/brain.js (The Content Aggregator)
This is where Vant targets the files stored in your local repository snapshots (such as identity.md or lessons.md) and concatenates them into system prompts.
*
* The PR Focus: Modify the data-collection function where Vant aggregates the file tree. You must ensure that array elements are pushed in a rigid, deterministic sequence (e.g., [Identity] -> [Lessons] -> [Tools]).
* If Vant’s multi-brain logic dynamically imports files based on task relevance, implement an alphabetical Array.prototype.sort() right before the context object compiles to avoid cache invalidation.
*
## 2. lib/vant.js or lib/headless.js (The State Controller)
Because Vant features a Headless API and active runtime management, this file controls session initialization and ongoing state tracking.
*
* The PR Focus: This is the ideal home for the 5-minute heartbeat loop. You will want to inject a lightweight setInterval or cron-style task scheduler inside the startup initializer (startHeadless or init).
* If a session configuration has a caching flag enabled, this hook will run a minimal background request every 4 minutes to reset MiniMax’s short TTL clock.
*
## 3. lib/mcp/ or lib/providers/ (The Network Formatter)
Vant supports Model Context Protocol (MCP) endpoints and standard model integration routing. The network adapter translates generic Vant messaging loops into vendor-specific schemas.
*
* The PR Focus: Locate the file handling message arrays sent to the API. You must add a conditional parsing block for MiniMax.
* Instead of converting arrays into a flattened string payload, it must output a strict JSON block structure that appends the cache_control: {"type": "ephemeral"} parameter right after the large codebase or context block definitions.
*
------------------------------
## Do These Optimizations Benefit Other Models & Frontier Models?
Yes, massively. Structuring your prompts left-to-right (Static → Dynamic) and enforcing byte-perfect consistency is universally beneficial, though how much you save depends on the model's backend architecture:
## 1. Anthropic Claude (Sonnet / Opus)
Claude pioneered this explicit array system (cache_control: {"type": "ephemeral"}). Because MiniMax cloned this exact API layout, your PR for Vant will natively optimize Claude at the same time. It will drop your Claude input token bills by up to 90% during massive repository coding tasks.
## 2. DeepSeek (V3 / R1)
DeepSeek uses an advanced automatic prompt caching mechanism. It does not require you to write explicit cache_control JSON blocks, but it strictly demands left-to-right byte-level sequence match.
*
* By keeping Vant's system files, core memory files, and early history rigidly sorted at the top of the context window, DeepSeek will automatically cache the entire block.
*
## 3. Frontier Models (OpenAI GPT-4o)
OpenAI uses automated, passive prompt caching. Like DeepSeek, it automatically matches prefixes from left to right.
*
* If your Vant script randomly shuffles tool descriptions or appends fluid logs (like an active timestamp or rapid git diff outputs) to the top of the prompt, it instantly wipes OpenAI’s cache. Pushing those dynamic elements to the absolute bottom ensures that even high-tier frontier models run with optimal latency and lower costs.
*
Would you like help writing the precise JSON block transformations or JavaScript sorting hooks to test locally before pushing them into your nova/lib files?
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment