Sources:
- Kimi K3 technical report, Kimi Team, July 2026.
- Official Kimi K3 technology blog, Kimi.
Kimi K3 is a very large AI model built to do more than answer isolated questions. Its designers want it to work through long jobs: research, coding, using software tools, examining images and videos, creating websites, and recovering from mistakes.
The model contains about 2.8 trillion adjustable internal values, but it does not use all of them for every word. It is organized like a giant company with many specialist departments. For each piece of input, only a small selection of specialists is called. About 104 billion values are active at a time. This is still enormous, but much cheaper than activating the whole model.
It can accept up to roughly one million text-like units in one working context. Images and video are converted into these units too. One million units is a capacity limit, not a promise that the model will remember every detail perfectly.
The development process had three broad learning phases:
- Broad pre-training on text, code, mathematics, knowledge material, images, and video.
- Supervised training on examples of good instructions, tool use, and long tasks.
- Reinforcement learning, where the model tried tasks, received scores, and improved.
Kimi K3's most distinctive ideas are:
- a combination of fast, compressed memory and occasional full attention to the earlier context;
- access to useful outputs from earlier model layers instead of relying only on the immediately preceding layer;
- hundreds of specialists, with a balancing system that distributes work evenly;
- text and vision trained together from the beginning;
- training across different amounts of "thinking effort";
- long-running practice environments with tools, mock applications, isolated computers, and objective checkers;
- engineering that pauses, saves, and resumes extremely long training attempts without starting over.
The authors say Kimi K3 is the strongest open-weight model in their comparison and is competitive with leading closed models, but they also say it remains behind the two strongest proprietary systems overall.
Delta retention is the central memory mechanism inside Kimi Delta Attention, or KDA. It is about how the model updates its internal working memory while reading a sequence.
Imagine that the model maintains a compact whiteboard instead of keeping a separate full copy of every earlier word. For every new piece of information, KDA makes two decisions:
- How much of the existing whiteboard should remain? This is controlled by the retention factor, called
alpha. - How strongly should the new information correct the whiteboard? This is controlled by the write strength, called
beta.
Alpha is not one master forgetting percentage. Each attention head produces a separate retention value for each memory channel at every token. A value close to 1 means "keep almost all of this channel for the next step." A small value means "fade this channel strongly." Kimi K3 mathematically keeps each one-step value between about 0.0067 and 1. That lower bound is mainly a numerical-stability device; it does not guarantee that information will remain remembered across a long sequence.
"Delta" means correction or change. KDA does not merely add every new fact to memory. It first checks the part of memory associated with the new fact, subtracts the outdated or conflicting association, and writes the corrected association.
Suppose the working memory contains:
Nasir's preferred drink -> tea
Later, the sequence says:
Nasir now prefers coffee.
A simple additive memory might leave both "tea" and "coffee" competing with each other. The delta rule approximately does this:
- Find the memory direction associated with "Nasir's preferred drink."
- Reduce the old answer stored in that direction.
- Write "coffee" into that direction.
- Retain or fade other memory channels according to their own
alphavalues.
This is closer to correcting a whiteboard entry than endlessly adding sticky notes.
For each new token, KDA creates:
- a key, which identifies what kind of memory is being addressed;
- a value, which is the information to store;
- a query, which asks the memory for useful information;
- two gates:
alphafor retention andbetafor writing or correction.
In plain language, its update is:
fade selected old memory channels, erase the outdated association addressed by the key, write the new value with the chosen strength, and then answer the query from the updated memory.
This gives KDA three useful properties:
- its memory state remains a fixed size even when the input becomes extremely long;
- new facts can correct old facts instead of only accumulating beside them;
- different types of information can be forgotten at different rates.
Earlier versions converted retention scores in a way that could produce extremely tiny cumulative values. The computer then had to divide by those tiny numbers when processing chunks, producing enormous intermediate numbers and possible overflow.
Kimi K3 uses a scaled sigmoid to keep each log-retention value within a known range. The practical result is:
- the calculations stay within safe numerical limits;
- slow, special-case calculations are avoided;
- ordinary high-speed GPU matrix hardware can process all the small blocks.
This is an engineering improvement to make delta retention fast and stable, not a new kind of permanent memory.
Updating a whiteboard sounds strictly one-step-at-a-time. That would be too slow on GPUs. KDA therefore divides a long sequence into chunks:
- calculations inside a chunk are performed largely in parallel;
- a compact memory state passes from one chunk to the next;
- the result is mathematically equivalent to processing the sequence in order.
Across multiple GPUs, KDA Context Parallelism calculates how each segment would transform an incoming memory state. The segment transformations are then combined with a prefix-scan-style algorithm. This lets many devices work on different portions of the same long context.
Kimi K3 uses three KDA layers followed by one global Gated MLA attention layer. KDA is the efficient notebook. MLA is the occasional opportunity to compare distant tokens directly. A final MLA layer performs one more global review before the model finishes.
This hybrid arrangement compensates for the notebook's compression: KDA makes million-token processing affordable, while MLA periodically restores high-capacity direct attention.
The report's main algorithms and techniques are:
| Algorithm or technique | What it does |
|---|---|
| Kimi Delta Attention (KDA) | Maintains a fixed-size working memory, selectively fades old channels, and corrects stored associations with new information. |
| Gated Multi-head Latent Attention (Gated MLA) | Periodically performs a broader global review while storing a compressed version of earlier attention information. |
| Hybrid 3:1 attention | Uses three efficient KDA layers for every one global MLA layer, combining speed with direct long-range comparison. |
| Attention Residuals (AttnRes) | Lets later layers consult selected earlier processing stages instead of receiving only the latest combined result. |
| Stable LatentMoE | Routes each token through 16 of 896 compact specialists plus two shared generalists. |
| SiTU-GLU | Smoothly caps extreme internal values so they do not destabilize low-precision training. |
| Quantile Balancing | Adjusts specialist-selection thresholds so work is distributed evenly rather than overwhelming a few popular experts. |
| Histogram quantile estimation | Estimates those balancing thresholds over millions of tokens without gathering every routing score centrally. |
| Per-Head Muon | Balances the learning update separately for each attention head so large-gradient heads do not dominate smaller ones. |
| Progressive context extension | Teaches the model with 8K, 64K, 256K, and finally 1M-token contexts instead of beginning with the most expensive length. |
| Supervised fine-tuning (SFT) | Gives the model verified examples of instruction following, tool calling, and long-task execution. |
| Partial-rollout reinforcement learning | Pauses unusually long attempts, learns from completed attempts, and resumes the paused work later. |
| Reasoning-effort budget control | Trains low, high, and maximum thinking modes while penalizing needless overthinking. |
| Agentic Generative Reward Model | Creates a rubric, compares candidate answers, and penalizes verbosity when no simple automatic checker exists. |
| Multi-Teacher On-Policy Distillation | Combines nine specialist teachers - three domains at three effort levels - into one final model. |
| Quantization-aware training | Teaches the model to work reliably with compact 4-bit expert weights and 8-bit activations for cheaper deployment. |
| EAGLE-3-style speculative decoding | Uses a small draft layer to propose future tokens, which the full model verifies to generate text faster. |
| MoonEP | Copies busy experts where needed and balances routed work exactly across training devices. |
| Prefix caching | Reuses the already-processed beginning of a long conversation or task instead of computing it again. |
| AgentENV microVMs | Runs tool-using practice inside isolated virtual computers that can be paused, copied, saved, and resumed. |
For completeness, the basic training mixture included web text, code, mathematics, knowledge material, images, video, OCR, mixed image-text documents, and code paired with rendered visuals. The team used filtering, quality scoring, synthesis, and duplicate removal. The report provides broad categories rather than a complete source inventory.
The blog is a shorter launch-oriented explanation. It does not provide a deeper mathematical explanation of delta retention than the technical report, but it adds several practical details.
Kimi K3 was trained in a mode where its earlier thinking content is passed back to it on later turns. The blog warns that performance can become highly unstable if an agent framework removes that thinking history or if a conversation is switched from another model to Kimi K3 midway through the session.
K3 expects to receive its own complete working notes when the next turn begins. Giving it only the visible answers may be like returning a project to an employee after throwing away the private notebook they were using to track decisions.
This is related to K3's chat format, not the alpha retention gate in KDA:
- KDA retention controls the model's compressed mathematical memory while processing tokens.
- Preserved thinking history controls which earlier message content the surrounding application sends back on the next request.
The blog recommends using a verified compatible harness such as Kimi Code and avoiding a mid-session model switch.
Because reinforcement learning emphasized difficult, long-running tasks, K3 may continue acting, improvise around ambiguity, or make decisions for the user when a more cautious assistant would stop and ask.
The blog recommends explicit constraints in the system prompt or AGENTS.md when the application requires strict boundaries. This is a practical consequence of its long-horizon agent training: persistence is useful, but without boundaries it can become unwanted initiative.
The blog recommends supernode configurations with at least 64 accelerators for deployment. This reinforces an important distinction:
- the model is open-weight;
- running the complete 2.8-trillion-parameter system efficiently is still a major infrastructure project.
The blog says Kimi contributed KDA-aware prefix-caching work to the vLLM community. Prefix caching is especially important because repeatedly processing hundreds of thousands of previous tokens would otherwise be extremely expensive.
At the blog's publication, the listed API prices were:
- $0.30 per million cache-hit input tokens;
- $3.00 per million cache-miss input tokens;
- $15.00 per million output tokens.
The official service claimed a cache-hit rate above 90% for coding workloads. That figure concerns repeated prefixes in the official workload and should not be assumed for every application. Prices and measured cache-hit rates are publication-time figures and may change.
The launch blog said Kimi K3 was available through Kimi.com, Kimi Work, Kimi Code, and the Kimi API. Maximum thinking effort was the initial default, with low- and high-effort product modes planned for later updates. It also said the complete model weights would be released by July 27, 2026. These are publication-time announcements rather than independently verified current availability.
The blog adds a fusion-industry research demonstration and describes two Kimi Work product features:
- Widgets, which place generated interactive components inside a chat and can connect them to changing data;
- Dashboard, which collects selected widgets into a persistent project or topic view.
These are surrounding product features. They should not be confused with algorithms inside the base Kimi K3 model.
The blog explicitly acknowledges:
- sensitivity to missing thinking history;
- excessive proactiveness on minor or ambiguous tasks;
- a remaining user-experience gap compared with Claude Fable 5 and GPT-5.6 Sol.
These limitations are useful because benchmark scores alone do not show whether a model behaves predictably in an ordinary conversation or application.
Older AI progress came mainly from spending more computing power before release: larger models and more training data. Newer reasoning models add a second method: spending more computation while answering, allowing the model to think longer, use tools, and revise its work.
The Kimi team argues that open models have advanced quickly in the second area but have not increased the size of their basic foundation as aggressively as the best closed systems. Kimi K3 tries to advance both at once:
- a much larger basic model;
- more reinforcement learning;
- adjustable thinking effort;
- long sequences of tool use;
- a one-million-unit working context.
The intended behavior is a repeating loop: think, act, inspect the result, verify it, and adapt. Some training jobs lasted hundreds or thousands of tool calls.
The model uses two complementary ways of reading its context.
Most layers use Kimi Delta Attention. Imagine reading a huge book while maintaining a compact, continuously updated notebook. Old notes can fade by different amounts, and new facts can correct or overwrite older ones. This makes long sequences more affordable because the notebook stays a fixed size instead of storing a separate full record for every earlier word.
After every three of these fast-memory layers, the model uses a global-attention layer. This is like occasionally reopening the full relevant record and comparing distant passages directly. A final global-attention layer is also placed at the end.
The ratio is therefore three fast memory layers to one global review layer. The fast mechanism supplies efficiency and awareness of order and recency; the global mechanism supplies unrestricted comparison across the context.
The designers also put a lower limit on how sharply the fast memory may decay in a single step. This prevents numbers inside the computer from becoming so extreme that calculations overflow. They changed the output gate too, letting every input decide which parts of the attention result should pass onward.
The global layers use a compressed cache: instead of saving a large separate record for every attention head, they save a smaller shared description and reconstruct what is needed. They do not use a conventional explicit position-number system. Position is learned through Kimi Delta Attention's order-sensitive memory.
In an ordinary deep network, each layer mainly receives the combined result of the layer immediately before it. Useful earlier details may become blurred.
Kimi K3's Attention Residuals let a later layer choose among representations produced at earlier depths. It resembles an employee who can consult selected earlier drafts instead of receiving only the most recently edited document.
Saving every layer separately would consume too much memory, so the 93 layers are grouped into blocks. The model keeps block-level summaries and selectively combines them. This captures most of the benefit at a manageable cost.
Each attention layer is followed by a Mixture-of-Experts section. Think of this as a very large company:
- 896 specialist teams are available in each routing layer;
- 16 specialists are selected for each input unit;
- two generalist teams always provide common processing.
The specialists work in a narrower internal space, which makes calling 16 of them affordable. This is why the model can have about 2.8 trillion total values while activating roughly 104.2 billion for a particular input.
At this scale, two problems appear. First, internal numbers can explode. Kimi K3 normalizes the specialists' combined output and uses a smooth cap called SiTU-GLU. It behaves normally for moderate values but prevents unusually large values from destabilizing training.
Second, popular specialists can become overloaded while others receive almost no work. Quantile Balancing adjusts each specialist's admission threshold so that the workload approaches an equal target. A compact histogram estimates the needed thresholds over millions of training items without collecting every score in one place.
Kimi K3 processes text, images, and video in one shared sequence. It can write code, inspect the resulting screenshot or video frame, and revise the artifact without handing the job to a separate vision model.
Its visual encoder, MoonViT-V2, has about 401 million values and 27 layers. Unlike Kimi K2.5, it was trained from scratch as part of next-token prediction instead of beginning from an already trained image model. The team reports that this was more stable and matched the earlier approach on vision tests.
Images and videos share the same visual machinery. Large images are divided into patches; video is processed both within frames and across time. Visual information is compressed before entering the main model. The system accepts images up to 3584 by 3584 pixels within the reported setup.
The model uses an optimizer called Muon to update its internal values. For attention, Kimi K3 applies this balancing separately to each attention head. Each head gets its own carefully scaled coaching signal instead of letting the loudest heads dominate one shared adjustment.
Text and vision were trained together from the beginning. The basic task was next-token prediction: given what has appeared so far, predict what comes next. Because words, image features, and video features share the stream, the model gradually learns relationships among them.
The context curriculum had four main sizes:
- 8,000 units;
- 64,000 units;
- 256,000 units;
- one million units.
The longest and most expensive sequences were concentrated late in training. This resembles teaching someone with short chapters first, then reports, books, and finally an enormous case file.
The team ran smaller experiments to decide the model shape, batch size, learning rate, and amount of data relative to model size. It reports roughly 2.5 times better scaling efficiency than Kimi K2. This means K3's overall design reached a better validation result for a given amount of training computation. It does not mean every answer is 2.5 times better or that the model runs 2.5 times faster.
The report does not publish the full pre-training compute budget, duration, GPU count, energy use, or total data volume.
The team first gave the broadly trained model worked examples of good behavior. Earlier Kimi models generated many long agent trajectories; these were checked in several stages and annotated with human involvement. The examples taught instruction following, tool calling, adaptive reasoning, and execution over long jobs.
The model then practiced tasks and received rewards. Training covered three broad specialties:
- general questions, knowledge, vision, search, truthfulness, and professional work;
- general agents, including assistants, deep research, and writing;
- coding agents, including software engineering, GPU optimization, web development, and practical coding behavior.
Each specialty was trained at low, high, and maximum reasoning effort, creating nine teacher models.
Long attempts create a scheduling problem: a few very slow tasks can leave all computers waiting. The system pauses unfinished attempts once enough other attempts have completed, trains on the completed work, and resumes the paused attempts later. Because an attempt may continue across several model updates, the algorithm limits how far each update can move, reducing instability from slightly old experience.
The team also controlled overthinking. Each problem received an estimated answer budget. Attempts that exceeded the permitted multiple were penalized. The maximum-effort model got the largest budget, followed by high and low. For answers without an automatic right-or-wrong checker, a judge model read the work, created a rubric, scored the candidates, and recorded the scores. Excessively long answers were penalized to discourage winning merely by being verbose.
The nine specialist teachers were consolidated into one model through multi-teacher distillation. A useful analogy is a student rotating among nine coaches. For each type of task and effort level, the matching coach guides the student's choices. The final model can then switch domain and effort without requiring nine separately served models.
Most of the model's stored size lies in the specialist weights. During post-training, those weights were represented with very low-precision numbers, while more sensitive parts remained at higher precision. Training included the resulting rounding effects so that the model learned to tolerate them. This cuts memory and serving cost.
Kimi K3 also has a small draft component that guesses several upcoming tokens. The full model checks those guesses and accepts only valid ones, so output can be generated faster without changing the final probability distribution.
The team did not train Kimi K3 inside one fixed agent interface. It built a configurable practice framework that could vary the system prompt, tools, memory, skills, context management, subagents, and interaction style. It could imitate interfaces such as Kimi Code, Claude Code, Codex, OpenClaw, and Hermes. The aim was to teach general tool use rather than memorization of one interface.
Practice tasks included:
- web research whose final claims could be checked;
- professional workflows in finance, data analysis, investment banking, and law;
- visual problems where the model could write Python to crop, zoom, transform, calculate, and verify;
- GPU programs judged on both correctness and speed, with anti-cheating checks;
- mock Gmail, Notion, Slack, and Canvas applications;
- simulated multi-day assistant jobs with interdependent events;
- autonomous tasks where the model saw the goal and tools but not the correct procedure;
- website, game, 3D, visualization, SVG, and full-stack application creation.
Many tasks used objective verifiers that examined the final state rather than trusting the model's claim that it had finished. Public checkers offered feedback, while hidden checkers tested whether the solution generalized. This was designed to reduce "reward hacking," where a model finds a way to score well without genuinely completing the task.
The personal-assistant applications were mock systems, not live user Gmail, Slack, or Notion accounts. A single simulated run could span thousands of tool calls and millions of accumulated context units.
Training a model of this size requires spreading it across many computers. The report's infrastructure section is largely about preventing computers from waiting on one another.
Kimi Delta Attention was rewritten into specialized GPU operations. Within a device and across devices, long sequences are divided into segments that can be processed partly in parallel and then combined exactly.
For the hundreds of specialists, MoonEP moves temporary copies of busy specialists to where the work is. It guarantees every computer receives the same number of routed input units. Because workloads then have predictable shapes, the system avoids pauses and unnecessary memory copies.
Memory is saved in several ways:
- recompute cheap intermediate results instead of storing them;
- compress many saved activations to eight-bit form;
- move inactive material to another GPU, ordinary memory, or fast storage;
- divide gradients and optimizer state among devices;
- schedule visual processing during otherwise idle gaps.
For million-context reinforcement learning, reusable conversation prefixes can be moved from GPU memory to CPU memory and restored later. The scheduler watches memory pressure and automatically reduces the number of simultaneous jobs as their contexts grow.
The model acts inside isolated virtual computers called sandboxes. These can be paused, resumed, copied for judging, and periodically saved for recovery. The report says training and evaluation created 51,219,741 sandboxes from 1,505,678 images. These figures refer to virtual-computer environments, not training images or pictures.
For production serving, the system jointly manages the compact KDA state and the growing global-attention cache. Matching prefixes can be reused at 512-token boundaries, with especially useful KDA checkpoints kept at conversation-turn boundaries. Requests are routed toward the cluster that already has their cache. Long requests and short requests receive separate resource budgets so a burst of million-token jobs does not block everyone else.
The report evaluates reasoning, knowledge, coding, tool-using agents, and vision. Kimi K3 was generally run at maximum reasoning effort and temperature 1.0. Different models sometimes used different agent harnesses, and some proprietary-model results included fallback behavior or safety systems. These details make the comparison informative but not perfectly controlled.
The broad conclusion is:
- Kimi K3 usually trails Claude Fable 5 and GPT-5.6 Sol overall;
- it generally beats the other open and closed models included in the report;
- coding, long research, agentic work, spreadsheets, and visual tool use are particular strengths;
- the hardest research reasoning, some persistent multi-day behavior, some computer-use tasks, and difficult cybersecurity exploitation remain weaker areas.
Representative reported results include:
- 93.5% on GPQA Diamond, a difficult science reasoning test;
- 77.8% on ProgramBench;
- 88.3% on Terminal-Bench 2.1;
- 91.2% on BrowseComp;
- 95.0 F1 on DeepSearchQA;
- 91.1% on OmniDocBench;
- 97.8% on Math-Vision when Python tools were allowed.
The report also includes many in-house tests. Kimi K3 led its comparison on swarm coordination and deep research, performed strongly as a practical coding assistant, and was preferred over Claude Opus 4.8 on the internal web-development comparison. It was weaker on an internal agent-behavior test, multi-system collaboration, always-on assistant work, and some visual knowledge work. Internal benchmarks are useful development evidence, but the organization that built the model also designed and administered them.
Third-party results reported as of July 23, 2026 placed Kimi K3 fourth of 580 entries on Artificial Analysis's Intelligence Index, second of 39 on Vals AI's index, first on the WebDev Arena, eighth on the Text Arena, and fourth on the Agent Arena. Arena rankings can change as new votes arrive.
The cost section claims near-leading results at lower per-task cost than the strongest proprietary models on four coding and agent benchmarks. Cost comparisons depend on API prices, reasoning settings, harnesses, cache behavior, and task length, so they are snapshots rather than universal prices.
The team separately tested vulnerability discovery and exploit creation.
For vulnerability discovery, the model produced hundreds of candidates across widely used software. About 70% of the human-reviewed findings were confirmed genuine, including 16 previously unknown vulnerabilities in six projects. The report highlights two serious Linux-kernel findings.
For end-to-end exploitation, Kimi K3 solved 14 of 36 internal tasks, compared with 8 for GLM-5.2. Ten of Kimi K3's successes were in ordinary user-space software; kernel exploitation was much harder. Common failures included not completing the last step, choosing the wrong strategy, getting stuck in debugging loops, and failing to verify the final exploit.
The report also cites an independent UK AI Security Institute and NIST CAISI assessment. Kimi K3 beat GLM-5.2 on some exploit measures but remained behind the strongest cyber-capable models and achieved arbitrary code execution on 0 of 41 tasks in one evaluation.
The honest interpretation is that Kimi K3 has meaningful dual-use cyber capability, especially for finding vulnerabilities, but it is not reliably equivalent to an expert human exploit developer.
The report presents examples rather than controlled general guarantees:
- Kimi K3 optimized four GPU operations, including a 73.6% runtime reduction for KDA in the reported setup.
- It built MiniTriton, a compact compiler and tensor library, including automatic differentiation and multi-GPU support.
- In a 48-hour autonomous run, it designed and verified a small experimental inference-chip prototype using open-source tools.
- For an astrophysics research reproduction, it reviewed more than 20 papers, processed more than 300 equations of state, wrote more than 3,000 lines of Python, and produced an interactive dashboard in about two hours.
- It created a research website about 42 years of AI chips using 87 quarterly reports, 99 PDFs, more than 2,800 web searches, and more than 1,100 terminal queries.
- It analyzed 391 gravitational-wave events using more than 20 subagents.
- It created a motion-graphics explanation of its own architecture and edited a teaser video from 56 clips.
These cases show what the system can do under favorable, heavily instrumented conditions. They do not tell us the success rate across ordinary users' tasks.
The mathematical appendices justify four engineering choices:
- SiTU-GLU behaves like the familiar activation for ordinary values but smoothly caps extreme outputs.
- Quantile Balancing can be derived as the best equal-work assignment between input units and specialists.
- Histograms estimate the balancing thresholds accurately and cheaply across distributed training.
- MoonEP's proof shows how many temporary specialist copies are sufficient to guarantee an equal workload across machines.
The final appendix describes XTML, Kimi K3's chat format. It is XML-like but uses reserved tokens so message boundaries are unambiguous. Assistant messages can contain separate thinking, user-visible response, and tool-call channels. Parallel tool calls carry index numbers so returned results match the correct call. Tool declarations can be added during a session, and requested reasoning effort is expressed as an instruction. The authors designed the format to be extensible, easy for the model to learn, and easy for software to parse.
The strongest part of the report is not simply "2.8 trillion parameters." It is the complete system around the model: selective specialists, long-context memory, vision, multi-effort reinforcement learning, verifiable practice worlds, resumable virtual computers, and production cache management.
The report is also unusually candid that the model remains behind the strongest proprietary systems overall and that difficult cyber exploitation and research-level reasoning still have clear gaps.
Important unanswered questions include:
- exact training-data volume and provenance;
- use of personal or copyrighted material;
- full training cost, hardware, energy, and carbon impact;
- safety training beyond the reported cyber evaluation;
- rates of ordinary hallucination and failure in real deployment;
- how often the spectacular case studies can be reproduced;
- independent replication of the architecture and training claims.
Kimi K3 is best understood as a giant, tool-using apprentice organization. It has hundreds of specialist departments, a compact notebook for long histories, periodic access to the broader record, built-in eyesight, multiple levels of thinking effort, and extensive practice in realistic but controlled computer environments.
Its training recipe is broad pre-training followed by verified demonstrations, reward-based practice, and consolidation of nine specialist teachers into one model. Its efficiency comes from activating only part of the organization, compressing working memory, carefully balancing specialists, and saving and resuming long jobs.
The architectural centerpiece is delta retention: each new token can preserve some memory channels, fade others, remove an outdated key-value association, and write a correction into a compact fixed-size state. Kimi K3 surrounds that mechanism with periodic global attention, selective expert routing, cross-layer retrieval, progressive long-context training, reinforcement learning, multi-teacher distillation, low-precision deployment, and extensive systems engineering.