I've watched “million-token context” become one of those AI specifications that sounds almost magical until you try to understand what the machine has to remember.
The number of tokens is only half the story.
The real problem is what happens inside the hardware while all that information is being processed and remembered.
DeepSeek and Xiaomi are now attacking that problem from two different architectural directions.
DeepSeek V4.1-Flash dramatically reduces its key-value cache footprint while supporting up to one million tokens. Xiaomi's HySparse2, designed for its next-generation MiMo architecture, tries to reduce both the computation required to process huge contexts and the memory needed to retain them.
The result is a fascinating glimpse at where AI infrastructure is going next.
Editorial concept illustrating two different approaches to reducing long-context AI memory and computation.
Why AI Memory Is Becoming the Real Bottleneck
When an AI agent searches the web, reads a document, calls a tool and continues reasoning, the amount of information entering its context can grow much faster than the amount of text it actually generates.
That creates a problem called prefill.
Before the model can generate its next token, it has to process the incoming context and build the internal state required for attention.
As context grows, that work becomes increasingly expensive.
The model also needs to retain a KV cache — the working memory used to avoid recomputing every previous token during generation.
For long-running agents, that cache can become a major memory and storage burden.
DeepSeek's Answer: Shrink the KV Cache
DeepSeek V4.1-Flash is a 552-billion-parameter multimodal mixture-of-experts model with a context window of up to one million tokens.
But the striking part is not only the context window.
DeepSeek redesigned the architecture around long-context efficiency.
DeepSeek V4.1-Flash's Memory Strategy
- Causal Encoder-Decoder: Uses different active parameter counts for prompt processing and generation.
- CSA2: Enables cross-layer KV cache reuse.
- FP4 KV caching: Stores cached attention information at very low numerical precision.
- SWA Bounded Replay: Uses selective recomputation to reduce persistent cache storage.
- Long context: Supports up to one million tokens.
DeepSeek reports that the global KV cache is only about 890 bytes per token, roughly one quarter of the corresponding footprint in V4-Flash.
The persistent KV-cache footprint can be reduced further through bounded replay.
Why 890 Bytes Per Token Is Such a Big Deal
At one million tokens, 890 bytes per token represents roughly 890 megabytes of global KV-cache storage.
That is dramatically smaller than the raw context length might make you expect.
But it would be a mistake to conclude that a million-token DeepSeek model therefore fits into a typical desktop GPU.
It doesn't follow.
The KV cache is only one part of the memory equation.
Model weights can be hundreds of gigabytes before quantization, and the actual runtime also needs memory for active experts, temporary buffers, vision components and other data structures.
Xiaomi's Approach Is Different
Several days later, Xiaomi published HySparse2, an architecture designed for long-horizon, multi-turn agent workloads and identified as the core architecture for its next-generation MiMo-V3 direction.
The research paper focuses on the same fundamental problem: agents accumulate enormous histories from web pages, tool responses, documents, code and execution traces.
Xiaomi's answer is not simply to compress the cache.
It changes how the model processes the history in the first place.
HySparse2 Uses Two Levels of KV Sharing
HySparse2 divides the model backbone into a self-decoder and a cross-decoder.
The first stage combines full attention with sliding-window attention, while the second uses full attention together with sparse attention for global retrieval.
KV Bridging lets cross-decoder full-attention layers derive their cached keys and values from hidden states produced by the self-decoder.
That means prefill can exit after the self-decoder rather than processing the entire backbone.
Overlooked Detail: Xiaomi Is Saving Compute Before It Saves Memory
DeepSeek's V4.1-Flash story is famous for its tiny KV cache. HySparse2 attacks another cost first: the amount of computation required to process long histories before generation can begin.
Token-Level Retrieval Is the Other Big Change
Earlier sparse-attention systems could select groups or blocks of tokens.
HySparse2 moves to token-level selection.
That sounds like a tiny implementation detail.
It isn't.
If a relevant fact occupies only a small part of a large block, block-level selection can waste part of the attention budget on irrelevant information.
Token-level selection lets the model be more precise about which historical information deserves attention.
Xiaomi's paper reports that this improved long-context retrieval while also reducing resource requirements.
The Numbers Xiaomi Reports Are Interesting
| Metric at 1M Tokens | Hybrid SWA | HySparse2 |
|---|---|---|
| Prefill FLOPs | Baseline | 5.02× lower |
| KV Cache | About 12.09GB in the cited evaluation | About 2.69GB |
| Selection | Hybrid attention strategy | Token-level sparse selection |
| Architecture | Hybrid SWA | Two-level KV sharing |
The paper's evaluation used an 80B-A3B MoE research model.
That matters because these figures are architectural experiment results, not proof that a released MiMo-V3 model will deliver exactly the same performance.
The 80% Number Needs Context
You will see headlines describing HySparse2 as reducing compute costs by 80%.
The underlying paper gives the more precise comparison: at one million tokens, HySparse2 reduces prefill FLOPs by 5.02× relative to Hybrid SWA.
That corresponds to roughly an 80% reduction in the tested setup.
But FLOPs are not the same thing as the final electricity bill, API price or wall-clock latency of every production system.
Those outcomes depend on hardware, kernels, memory bandwidth, batching, network overhead and software implementation.
DeepSeek and Xiaomi Are Solving Different Parts of the Same Problem
DeepSeek V4.1-Flash
- 552B multimodal MoE backbone.
- 1M-token context support.
- 890-byte global KV cache per token.
- FP4 KV caching.
- Cross-layer KV reuse and bounded replay.
Xiaomi HySparse2
- Designed for long-horizon agents.
- Two-level KV sharing.
- Token-level sparse retrieval.
- Prefill can exit early.
- Research evaluation rather than a released MiMo-V3 model.
The important point is that there isn't a single “AI memory trick.”
Researchers can reduce precision, reuse cached information, select fewer tokens, replay information selectively or redesign the computational path itself.
Nvidia's Jensen Huang Has Been Warning About This
The industry's hardware leaders are increasingly treating AI memory as an architectural problem rather than a simple GPU specification.
“One of the hardest parts is memory.”
Nvidia has been building context-memory infrastructure designed to extend effective GPU memory and move AI working memory closer to the accelerator fabric.
That direction is consistent with what DeepSeek and Xiaomi are doing at the model-architecture level.
The software is learning to remember more efficiently while the hardware ecosystem is being redesigned to move and store that memory more efficiently.
Why This Matters for AI Agents
Long-context chat is only one use case.
Agents are the bigger reason the memory problem is accelerating.
An agent can search the web, inspect a repository, call an API, read the result, perform another action and repeat the cycle dozens or hundreds of times.
That creates a constantly expanding history.
For agents, memory is not just a convenience. It is part of the runtime architecture.
Where Better AI Memory Matters
- Coding agents: Large repositories and long execution histories.
- Research agents: Thousands of pages and repeated searches.
- Enterprise assistants: Large document and workflow histories.
- Scientific AI: Long experiments and accumulated observations.
- Autonomous workflows: Multi-step tasks that continue for hours.
Amazon: Hardware for Experimenting With Local AI
NVIDIA GeForce RTX 5090
The RTX 5090's 32GB of VRAM makes it a useful high-end consumer platform for smaller local AI models and experimentation, although memory-intensive models still require more capacity or offloading.
Check RTX 5090 on AmazonApple Mac Studio
Apple's unified-memory architecture makes high-memory Mac configurations interesting for local AI workloads where model capacity matters more than conventional discrete-GPU VRAM.
Check Mac Studio on AmazonHigh-Speed NVMe SSD
Fast NVMe storage can become important when AI runtimes need to stream model weights or persistent context instead of keeping everything resident in GPU memory.
Browse NVMe SSDs on AmazonWatch DeepSeek V4.1-Flash Explained
This English-language breakdown from The Stack focuses specifically on DeepSeek V4.1-Flash's KV-cache architecture, 890-byte figure and the memory techniques behind its million-token context.
What AI Developers Should Take Away
The biggest lesson is simple: do not confuse a larger context window with unlimited memory.
Context length tells you how much information the model can theoretically process in one request.
Memory architecture determines how efficiently the system can actually do that.
When evaluating a long-context model, look beyond the headline token count.
Overlooked Checklist
Check KV-cache size, weight memory, precision, prefill compute, memory bandwidth, retrieval accuracy, runtime support and whether the published results came from the actual released model or a research configuration.
The Bottom Line on DeepSeek, Xiaomi and AI Memory
DeepSeek V4.1-Flash and Xiaomi HySparse2 point toward the same future from different directions.
AI systems are becoming more dependent on long histories, tool results and persistent context.
That makes memory one of the defining engineering problems of the agent era.
DeepSeek attacks the problem with aggressive KV-cache compression, cross-layer reuse, FP4 caching and selective replay.
Xiaomi attacks it with two-level KV sharing, early prefill exit and token-level sparse retrieval.
Neither approach magically removes the memory requirements of a huge model.
What they demonstrate is much more interesting.
The future of long-context AI will probably depend less on simply adding more memory and more on learning how to avoid wasting the memory we already have.
That is a shift worth watching.
Inside Xiaomi's 1-Trillion Parameter Behemoth
While Xiaomi redesigns AI memory architecture to handle massive context windows, they are also pushing the boundaries of sheer model scale. Read our complete guide to Xiaomi MiMo V2.6 to discover how this massive 1-trillion parameter open model is challenging closed-source giants and redefining local hardware limits.
Read the MiMo V2.6 Guide →Sources checked for this article:
Geeky Gadgets — Xiaomi MiMo V3 and DeepSeek V4.1 Flash Memory Comparison
DeepSeek — Introducing DeepSeek-V4.1-Flash
arXiv — DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
arXiv — HySparse2: Hybrid Sparse Attention With Two-Level KV Sharing
TechNode — Xiaomi MiMo-V3 and HySparse2
Frequently Asked Questions
What is DeepSeek V4.1-Flash's memory limit?
DeepSeek V4.1-Flash supports a context window of up to one million tokens. Its global KV-cache footprint is reported at about 890 bytes per token, but this does not represent the total memory required to load and run the 552B-parameter model.
What is Xiaomi HySparse2?
HySparse2 is a hybrid sparse-attention architecture developed by Xiaomi's LLM-Core team for long-horizon, multi-turn agent workloads. It uses two-level KV sharing and token-level sparse retrieval to reduce prefill computation and KV-cache storage.
How much memory does DeepSeek V4.1-Flash need?
The exact hardware requirement depends on quantization, runtime, model components and offloading strategy. The small KV-cache footprint does not mean the full 552B model can fit into a typical consumer GPU.
Does Xiaomi MiMo-V3 already have a public model?
No. Xiaomi has published the HySparse2 architecture associated with its next-generation MiMo direction, but the research paper does not constitute a release of MiMo-V3 weights. Model availability and final specifications can change.
Why is KV cache important for long-context AI?
KV cache stores attention information from previously processed tokens so the model does not have to recompute the entire history during generation. Smaller or more efficiently shared caches can reduce memory, storage and data-movement requirements.
No comments:
Post a Comment