Latest

Solid AI. Smarter Tech.

How Xiaomi and DeepSeek Are Fixing AI's Memory Flaw

AI MEMORY DeepSeek V4.1 Flash · Xiaomi HySparse2 · KV Cache · 1M Context · AI Agents

I've watched “million-token context” become one of those AI specifications that sounds almost magical until you try to understand what the machine has to remember.

The number of tokens is only half the story.

The real problem is what happens inside the hardware while all that information is being processed and remembered.

DeepSeek and Xiaomi are now attacking that problem from two different architectural directions.

DeepSeek V4.1-Flash dramatically reduces its key-value cache footprint while supporting up to one million tokens. Xiaomi's HySparse2, designed for its next-generation MiMo architecture, tries to reduce both the computation required to process huge contexts and the memory needed to retain them.

The result is a fascinating glimpse at where AI infrastructure is going next.

DeepSeek and Xiaomi AI memory optimization for million-token long-context models

Editorial concept illustrating two different approaches to reducing long-context AI memory and computation.

Important distinction: A model's context window is not the same thing as the amount of memory required to load its weights. DeepSeek V4.1-Flash can support a million-token context, but its 552B-parameter backbone still represents a very large model.
1M
Maximum Context Tokens
890B
Global KV Cache Per Token
4.5×
Reported KV Cache Reduction
5.02×
Prefill FLOPs Reduction

Why AI Memory Is Becoming the Real Bottleneck

When an AI agent searches the web, reads a document, calls a tool and continues reasoning, the amount of information entering its context can grow much faster than the amount of text it actually generates.

That creates a problem called prefill.

Before the model can generate its next token, it has to process the incoming context and build the internal state required for attention.

As context grows, that work becomes increasingly expensive.

The model also needs to retain a KV cache — the working memory used to avoid recomputing every previous token during generation.

For long-running agents, that cache can become a major memory and storage burden.


DeepSeek's Answer: Shrink the KV Cache

DeepSeek V4.1-Flash is a 552-billion-parameter multimodal mixture-of-experts model with a context window of up to one million tokens.

But the striking part is not only the context window.

DeepSeek redesigned the architecture around long-context efficiency.

DeepSeek V4.1-Flash's Memory Strategy

  • Causal Encoder-Decoder: Uses different active parameter counts for prompt processing and generation.
  • CSA2: Enables cross-layer KV cache reuse.
  • FP4 KV caching: Stores cached attention information at very low numerical precision.
  • SWA Bounded Replay: Uses selective recomputation to reduce persistent cache storage.
  • Long context: Supports up to one million tokens.

DeepSeek reports that the global KV cache is only about 890 bytes per token, roughly one quarter of the corresponding footprint in V4-Flash.

The persistent KV-cache footprint can be reduced further through bounded replay.

That 890-byte figure is not the model size. It describes the global KV cache footprint per token. The model's parameters, expert weights, embeddings and other components still require substantial storage and compute resources.

Why 890 Bytes Per Token Is Such a Big Deal

At one million tokens, 890 bytes per token represents roughly 890 megabytes of global KV-cache storage.

That is dramatically smaller than the raw context length might make you expect.

But it would be a mistake to conclude that a million-token DeepSeek model therefore fits into a typical desktop GPU.

It doesn't follow.

The KV cache is only one part of the memory equation.

Model weights can be hundreds of gigabytes before quantization, and the actual runtime also needs memory for active experts, temporary buffers, vision components and other data structures.


Xiaomi's Approach Is Different

Several days later, Xiaomi published HySparse2, an architecture designed for long-horizon, multi-turn agent workloads and identified as the core architecture for its next-generation MiMo-V3 direction.

The research paper focuses on the same fundamental problem: agents accumulate enormous histories from web pages, tool responses, documents, code and execution traces.

Xiaomi's answer is not simply to compress the cache.

It changes how the model processes the history in the first place.


HySparse2 Uses Two Levels of KV Sharing

HySparse2 divides the model backbone into a self-decoder and a cross-decoder.

The first stage combines full attention with sliding-window attention, while the second uses full attention together with sparse attention for global retrieval.

KV Bridging lets cross-decoder full-attention layers derive their cached keys and values from hidden states produced by the self-decoder.

That means prefill can exit after the self-decoder rather than processing the entire backbone.

Overlooked Detail: Xiaomi Is Saving Compute Before It Saves Memory

DeepSeek's V4.1-Flash story is famous for its tiny KV cache. HySparse2 attacks another cost first: the amount of computation required to process long histories before generation can begin.


Token-Level Retrieval Is the Other Big Change

Earlier sparse-attention systems could select groups or blocks of tokens.

HySparse2 moves to token-level selection.

That sounds like a tiny implementation detail.

It isn't.

If a relevant fact occupies only a small part of a large block, block-level selection can waste part of the attention budget on irrelevant information.

Token-level selection lets the model be more precise about which historical information deserves attention.

Xiaomi's paper reports that this improved long-context retrieval while also reducing resource requirements.


The Numbers Xiaomi Reports Are Interesting

Metric at 1M Tokens Hybrid SWA HySparse2
Prefill FLOPs Baseline 5.02× lower
KV Cache About 12.09GB in the cited evaluation About 2.69GB
Selection Hybrid attention strategy Token-level sparse selection
Architecture Hybrid SWA Two-level KV sharing

The paper's evaluation used an 80B-A3B MoE research model.

That matters because these figures are architectural experiment results, not proof that a released MiMo-V3 model will deliver exactly the same performance.


The 80% Number Needs Context

You will see headlines describing HySparse2 as reducing compute costs by 80%.

The underlying paper gives the more precise comparison: at one million tokens, HySparse2 reduces prefill FLOPs by 5.02× relative to Hybrid SWA.

That corresponds to roughly an 80% reduction in the tested setup.

But FLOPs are not the same thing as the final electricity bill, API price or wall-clock latency of every production system.

Those outcomes depend on hardware, kernels, memory bandwidth, batching, network overhead and software implementation.


DeepSeek and Xiaomi Are Solving Different Parts of the Same Problem

DeepSeek V4.1-Flash

  • 552B multimodal MoE backbone.
  • 1M-token context support.
  • 890-byte global KV cache per token.
  • FP4 KV caching.
  • Cross-layer KV reuse and bounded replay.

Xiaomi HySparse2

  • Designed for long-horizon agents.
  • Two-level KV sharing.
  • Token-level sparse retrieval.
  • Prefill can exit early.
  • Research evaluation rather than a released MiMo-V3 model.

The important point is that there isn't a single “AI memory trick.”

Researchers can reduce precision, reuse cached information, select fewer tokens, replay information selectively or redesign the computational path itself.


Nvidia's Jensen Huang Has Been Warning About This

The industry's hardware leaders are increasingly treating AI memory as an architectural problem rather than a simple GPU specification.

Jensen Huang:
“One of the hardest parts is memory.”

Nvidia has been building context-memory infrastructure designed to extend effective GPU memory and move AI working memory closer to the accelerator fabric.

That direction is consistent with what DeepSeek and Xiaomi are doing at the model-architecture level.

The software is learning to remember more efficiently while the hardware ecosystem is being redesigned to move and store that memory more efficiently.


Why This Matters for AI Agents

Long-context chat is only one use case.

Agents are the bigger reason the memory problem is accelerating.

An agent can search the web, inspect a repository, call an API, read the result, perform another action and repeat the cycle dozens or hundreds of times.

That creates a constantly expanding history.

For agents, memory is not just a convenience. It is part of the runtime architecture.

Where Better AI Memory Matters

  • Coding agents: Large repositories and long execution histories.
  • Research agents: Thousands of pages and repeated searches.
  • Enterprise assistants: Large document and workflow histories.
  • Scientific AI: Long experiments and accumulated observations.
  • Autonomous workflows: Multi-step tasks that continue for hours.

Amazon: Hardware for Experimenting With Local AI

NVIDIA GeForce RTX 5090

The RTX 5090's 32GB of VRAM makes it a useful high-end consumer platform for smaller local AI models and experimentation, although memory-intensive models still require more capacity or offloading.

Check RTX 5090 on Amazon

Apple Mac Studio

Apple's unified-memory architecture makes high-memory Mac configurations interesting for local AI workloads where model capacity matters more than conventional discrete-GPU VRAM.

Check Mac Studio on Amazon

High-Speed NVMe SSD

Fast NVMe storage can become important when AI runtimes need to stream model weights or persistent context instead of keeping everything resident in GPU memory.

Browse NVMe SSDs on Amazon

Watch DeepSeek V4.1-Flash Explained

This English-language breakdown from The Stack focuses specifically on DeepSeek V4.1-Flash's KV-cache architecture, 890-byte figure and the memory techniques behind its million-token context.


What AI Developers Should Take Away

The biggest lesson is simple: do not confuse a larger context window with unlimited memory.

Context length tells you how much information the model can theoretically process in one request.

Memory architecture determines how efficiently the system can actually do that.

When evaluating a long-context model, look beyond the headline token count.

Overlooked Checklist

Check KV-cache size, weight memory, precision, prefill compute, memory bandwidth, retrieval accuracy, runtime support and whether the published results came from the actual released model or a research configuration.


The Bottom Line on DeepSeek, Xiaomi and AI Memory

DeepSeek V4.1-Flash and Xiaomi HySparse2 point toward the same future from different directions.

AI systems are becoming more dependent on long histories, tool results and persistent context.

That makes memory one of the defining engineering problems of the agent era.

DeepSeek attacks the problem with aggressive KV-cache compression, cross-layer reuse, FP4 caching and selective replay.

Xiaomi attacks it with two-level KV sharing, early prefill exit and token-level sparse retrieval.

Neither approach magically removes the memory requirements of a huge model.

What they demonstrate is much more interesting.

The future of long-context AI will probably depend less on simply adding more memory and more on learning how to avoid wasting the memory we already have.

That is a shift worth watching.

Inside Xiaomi's 1-Trillion Parameter Behemoth

While Xiaomi redesigns AI memory architecture to handle massive context windows, they are also pushing the boundaries of sheer model scale. Read our complete guide to Xiaomi MiMo V2.6 to discover how this massive 1-trillion parameter open model is challenging closed-source giants and redefining local hardware limits.

Read the MiMo V2.6 Guide →


Frequently Asked Questions

What is DeepSeek V4.1-Flash's memory limit?

DeepSeek V4.1-Flash supports a context window of up to one million tokens. Its global KV-cache footprint is reported at about 890 bytes per token, but this does not represent the total memory required to load and run the 552B-parameter model.

What is Xiaomi HySparse2?

HySparse2 is a hybrid sparse-attention architecture developed by Xiaomi's LLM-Core team for long-horizon, multi-turn agent workloads. It uses two-level KV sharing and token-level sparse retrieval to reduce prefill computation and KV-cache storage.

How much memory does DeepSeek V4.1-Flash need?

The exact hardware requirement depends on quantization, runtime, model components and offloading strategy. The small KV-cache footprint does not mean the full 552B model can fit into a typical consumer GPU.

Does Xiaomi MiMo-V3 already have a public model?

No. Xiaomi has published the HySparse2 architecture associated with its next-generation MiMo direction, but the research paper does not constitute a release of MiMo-V3 weights. Model availability and final specifications can change.

Why is KV cache important for long-context AI?

KV cache stores attention information from previously processed tokens so the model does not have to recompute the entire history during generation. Smaller or more efficiently shared caches can reduce memory, storage and data-movement requirements.

Disclosure: Some of the product links in this article are Amazon affiliate links. If you purchase a product through one of these links, we may earn a small commission at no additional cost to you. Our recommendations are based on the product's relevance and usefulness to our readers.

No comments:

Post a Comment

Explore More