AI infrastructure
KV cache leaves the GPU: Scality puts storage in the inference path
On October 8, 2026, data-infrastructure company Scality launched AI Inference Factory, an open-code software stack for deploying and operating AI inference on enterprise-owned infrastructure. The technical claim worth reading: freeing the KV cache from GPU memory into a multi-petabyte shared storage layer, putting storage directly in the inference serving path with prefill/decode disaggregation. All performance figures are vendor-reported with no independent verification; the sovereignty narrative ultimately has to be verified against KV-cache residency and jurisdiction.
Muse research and drafting; Muse staged review in the same author context. This article was researched, drafted in Chinese, fact-reviewed, and fully translated and semantically reviewed across eight languages by Muse, using staged same-platform review (reviewMode=muse_same_platform_staged, contextRelationship=same_author_context). Same-author context is disclosed as-is; this is not claimed to be a third-party independent audit. The primary source is the full text of Scality's press release (GlobeNewswire, 2026-10-08); performance figures are Scality's stated numbers, each flagged as unverified.
What happened
On October 8, 2026, data-infrastructure company Scality announced in San Francisco that AI Inference Factory is available now. It is an open-code software stack for enterprises, government agencies, and neo-cloud providers to deploy and operate AI inference on infrastructure they own — billed in the release as a supported alternative to cloud AI services, without assembling the whole stack from scratch.
The stack has four parts: validated open-weight models maintained as part of the stack; a disaggregated inference-serving layer with prefill and decode scaling independently; a control plane handling authentication, metering, routing requests to GPUs holding the relevant context, and SLA scheduling; and Scality ADI, a policy-governed autonomous data-management infrastructure that plays the shared-storage role.
It ships as a software license or a fully managed service, and was demonstrated at Scality Day in Paris on October 8. Scality says it validates, ships, and maintains the integrated stack so organizations can keep pace with fast-changing models, serving technologies, and security requirements.
Why now
The release frames the motivation as three bills. First, per-token cloud pricing makes costs unpredictable as workflows get more useful — and impossible to cap. Second, model versions and quantization can change at the provider's discretion, shaking the workflows built on top. Third, where prompts, documents, and proprietary code are processed, under which jurisdiction, and who can cut access — the three sovereignty questions for sensitive and regulated data. Scality's bet: most organizations end up hybrid, with the most critical AI processes running on-premises.
The release quotes IDC analyst Nataliya Yezhkova: where infrastructure can hold and serve model state efficiently, on-premises inference is a credible option for a growing set of enterprise and public-sector workloads. Note: this is an in-release endorsement quote, not an independent IDC report.
The technical substance
The meatiest part is Scality ADI as a shared KV-cache layer. During inference the KV cache grows with context length and quickly overflows GPU HBM; the conventional approach keeps it on the GPU server where it was born. Scality's approach: ADI provides a multi-petabyte shared cache that GPUs read at the same order of magnitude as GPU memory, restoring context from storage instead of recomputing it — improving GPU utilization and inference cost.
Paired with prefill/decode disaggregation: one GPU pool does prefill only, writing KV cache to ADI; another does decode only, reading the cache back to generate tokens. Any decode GPU can pick up any context, prefill no longer interrupts decode, and both pools run full.
CTO Giorgio Regni's words are worth recording: the KV cache on ADI is fast enough to sit in the serving path; restoring a context from ADI is the same order of magnitude as GPU memory and 14 times faster than recomputing it, with GPUs staying busy throughout; storage is no longer the reason to keep the KV cache inside the GPU server. CEO Jérôme Lecat frames the sovereignty narrative: organizations need more control over where inference runs, how models are managed, and what happens to their data.
Reading the numbers
Scality published a set of self-tested figures, all from company testing with no independent verification: Gemma-3 27B loads in 1.9 seconds via RDMA in parallel across the cluster, about 10x local NVMe; 166 ms warm time-to-first-token restoring a 14K-token context from ADI, only 83 ms behind HBM; 14x faster KV-cache retrieval than recomputation at 14K tokens, 72x at 439K; cache over 80x a single GPU's memory with 1,000 concurrent sessions resumable without recomputation; 97% of network line rate between GPUs and storage; context restored before the first token, with no measurable impact on token generation.
For disaggregation itself the release cites two third-party public studies: DistServe (OSDI 2024) serving up to 7.4x more requests under the same latency targets; Mooncake handling 75% more requests on Kimi production traffic. These are public-research figures, not Scality tests — keep them separate from the vendor numbers.
The sweet-spot conditions are explicit: 14K and 439K are the context lengths the release chose; no-measurable-impact depends on restore completing before the first token, with no tail-latency distribution disclosed. Until third-party reproduction exists, capacity planning should discount vendor claims and demand measurements on your own context distribution.
Sovereignty and openness
The compatibility list is broad: OpenCode, Hermes, Goose, LangGraph, and Pydantic AI harnesses and agent frameworks; validated open-weight models including Mistral, Gemma, gpt-oss, Qwen, Kimi, GLM, and DeepSeek; standard servers from Dell, HPE, Lenovo, and Supermicro. Every component ships as open code — customers can inspect how inference state is stored and moved, and submit contributions, but Scality reviews them: the openness is real, and so is the gatekeeper.
Two bills to settle. First, open code with Scality-reviewed contributions is not the same as pure community open source; it is closer to commercial open source with vendor-controlled trunk. Ask about review criteria, security-patch SLAs, and support boundaries after a fork. Second, once the KV cache becomes a persistent, auditable, cross-GPU-portable asset, prompts and documents sit in shared storage as cache — residency policy, encryption, access audit, and deletion semantics all need rebuilding; the release claims policy governance and human-approved lifecycle policies for ADI but stays vague on mechanics, which is exactly the detail regulated buyers must pin down before signing.
Implications and action items
Three implications for enterprise teams considering on-premises or sovereign inference. First, the cost-optimization variables change: KV-cache hit rates, context-restore latency, and prefill/decode ratios are becoming as important as how many GPUs you buy; demand vendor measurements on your own context distribution, not the release's sweet-spot numbers.
Second, the control plane decides the stack: authentication, metering, routing, and SLA scheduling determine whether on-premises inference is an operable service or a pile of bare GPUs; Scality putting it in the four-part stack shows it understands enterprise procurement logic, and buyers should accept against the same checklist.
Third, three things cannot be asserted yet: no independent reproduction of the figures, so discount vendor claims in capacity planning; zero disclosed customers, so production stability and upgrade paths are unknown; and the local stack swaps token bills for fixed hardware-depreciation, power, and ops costs — more predictable is not cheaper, and TCO needs your own workload data. The sovereignty narrative must land at the data plane: open weights are only step one; residency and jurisdiction of KV cache, prompts, and audit logs are the real test.
Sources and further reading
Source records are supplied and reviewed by Muse in the same author context; they have not been independently fact-checked.
- Scality press release (GlobeNewswire, 2026-10-08, read in full)
Scality(经 GlobeNewswire 发布)
Recorded publication date ·
Recorded verification time ·