Prompt caching, first introduced on SambaCloud with MiniMax M2.7, is now available for MiniMax M3. When requests share a stable prefix of at least 4,096 tokens, SambaCloud can serve that prefix from cache instead of recomputing it, with no code changes required. Cached tokens are billed at 90% below the standard input rate, and across 8k to 192k tokens of context, cache hits cut time to first token (TTFT) by 35% to 88%.
TL;DR
-
Prompt caching is now available for MiniMax M3 on SambaCloud, so when agents resend the same long context, such as a repository, a set of research papers, or a fixed set of tool definitions, and the prefix is served from cache, only the new tokens are processed fresh.
-
It runs on Automatic Prefix Caching and needs no setup: prefixes of at least 4,096 tokens qualify, up to a maximum cacheable length of 192,000 tokens, and under sustained traffic hit rates typically climb above 90%.
-
Across 8k to 192k tokens of context, cache hits cut TTFT by 35% at the short end and 88% at the long end, a median speedup of 4.7x, with TTFT dropping from 8.4 seconds to 1.0 seconds at 192k tokens.
-
Cached tokens on MiniMax M3 are billed 90% below the standard input rate, at $0.06 per million tokens instead of $0.60.
-
Every response returns a prompt_tokens_details object with cached_tokens and cache_creation_tokens, so you can confirm exactly how many tokens were served from cache.
Why Prompt Caching Matters More with MiniMax M3
MiniMax M3 is built for long-horizon agents. Its 1M-token context window, powered by MiniMax Sparse Attention, means whole repositories, multi-day agent logs, and full research papers fit in a single request. Long, repeated contexts like these are where prompt caching pays off. Caching covers prefixes of up to 192,000 tokens, and within that limit, the longer the stable prefix, the larger the share of each request that can be served from cache instead of recomputed.
Source: https://sambanova.ai/blog/introducing-prompt-caching-for-minimax-m3-on-sambacloud