# MiniMax M3 Prompt Caching on SambaCloud: How It Works

Date: 7 Oct 2026
Released by: SambaNova

## MediaRelease.co Summary

SambaNova has made prompt caching available for MiniMax M3 on SambaCloud, with no code changes required. Requests sharing a stable prefix of at least 4,096 tokens, up to 192,000 tokens, can be served from cache, and cached tokens cost $0.06 per million instead of $0.60. SambaNova reports that cache hits cut time to first token by 35% to 88% across 8,000 to 192,000 tokens of context.

## Original release

Prompt caching, first introduced on SambaCloud with MiniMax M2.7, is now available for MiniMax M3. When requests share a stable prefix of at least 4,096 tokens, SambaCloud can serve that prefix from cache instead of recomputing it, with no code changes required. Cached tokens are billed at 90% below the standard input rate, and across 8k to 192k tokens of context, cache hits cut time to first token (TTFT) by 35% to 88%. TL;DR - Prompt caching is now available for MiniMax M3 on SambaCloud, so when agents resend the same long context, such as a repository, a set of research papers, or a fixed set of tool definitions, and the prefix is served from cache, only the new tokens are processed fresh. - It runs on Automatic Prefix Caching and needs no setup: prefixes of at least 4,096 tokens qualify, up to a maximum cacheable length of 192,000 tokens, and under sustained traffic hit rates typically climb above 90%. - Across 8k to 192k tokens of context, cache hits cut TTFT by 35% at the short end and 88% at the long end, a median speedup of 4.7x, with TTFT dropping from 8.4 seconds to 1.0 seconds at 192k tokens. - Cached tokens on MiniMax M3 are billed 90% below the standard input rate, at $0.06 per million tokens instead of $0.60. - Every response returns a prompt_tokens_details object with cached_tokens and cache_creation_tokens, so you can confirm exactly how many tokens were served from cache. Why Prompt Caching Matters More with MiniMax M3 MiniMax M3 is built for long-horizon agents. Its 1M-token context window, powered by MiniMax Sparse Attention, means whole repositories, multi-day agent logs, and full research papers fit in a single request. Long, repeated contexts like these are where prompt caching pays off. Caching covers prefixes of up to 192,000 tokens, and within that limit, the longer the stable prefix, the larger the share of each request that can be served from cache instead of recomputed. Source: https://sambanova.ai/blog/introducing-prompt-caching-for-minimax-m3-on-sambacloud

## Key details

- Issued by: SambaNova
- Published: 7 Oct 2026
- Publisher country: United States
- Subject country/region: United States
- Topics: AI
- Original source: https://sambanova.ai/blog/introducing-prompt-caching-for-minimax-m3-on-sambacloud
- MediaRelease.co URL: https://mediarelease.co/us/mr01542-minimax-m3-prompt-caching-on-sambacloud-how-it-works-07102026.html

Original source: https://sambanova.ai/blog/introducing-prompt-caching-for-minimax-m3-on-sambacloud
MediaRelease.co canonical URL: https://mediarelease.co/us/mr01542-minimax-m3-prompt-caching-on-sambacloud-how-it-works-07102026.html
