Kimi K3-256k

monneyboi4 pts0 comments

Model Configuration | Kimi Code Docs

Skip to contentMenuReturn to top

Model Configuration ​<br>This page covers the models Kimi Code provides and how to switch between them in each client.<br>Model Overview ​<br>Kimi Code currently offers two models—Kimi K3 and Kimi K2.7 Code—across four model IDs, selectable by model ID in clients or third-party tools. Model specs:<br>Recommended model launch<br>k3-256k is now available. Within 256k context, it delivers the same results. k3 (1M) consumes about twice as much quota as k3-256k. Ideal for everyday Q&A, code completion, routine feature development, and single-file or small-file edits — video input is not supported.<br>💓 Reminder<br>Switching from K3 (1M) to K3-256k: When switching from k3 (1M) to k3-256k, if the current session's context already exceeds 256k, some coding tools such as Kimi Code CLI and Claude Code will perform a compact on the tool side.<br>Switching recommendations: (1) Because different agent tools handle this differently, manually run compact once before switching to compress the context to within 256k. This preserves the key points of the task, keeps the session intact, and lets you benefit from more durable quota after switching. (2) If the conversation history includes video files, switching directly will fail because K3-256k does not support video input. Please compact first, then switch.<br>Switching from K3-256k to K3 (1M): When switching from k3-256k to k3 (1M), if k3-256k is close to the 256k limit and you don't want compact to lose information, you can switch directly to 1M. The current version switching from 256k to 1M does not affect the cache.

Model IDk3k3-256kkimi-for-codingkimi-for-coding-highspeedModel versionKimi K3Kimi K3Kimi K2.7 CodeK2.7 Code HighSpeedDescriptionKimi's most capable flagship coding model: 2.8T parameters, 1M context windowThe 256K context version of Kimi K3, effectively reducing consumptionGood at code completion and routine development tasksThe high-speed version of K2.7 Code, with the same coding ability and ~5–6× faster outputSpeedRegularRegularRegularHighSpeed (6× speed, 3× quota usage)Context windowUp to 1M (for higher-tier members)256k only256k256kReasoningreasoning_effort:low / high / max(default high)reasoning_effort:low / high / max (default high)Thinking:ONThinking:ONAvailabilityAvailable to Moderato and above; 1M context for Allegretto and aboveAvailable to all Moderato members and aboveAll membersAllegretto plan or aboveMultimodal inputImage, videoImage onlyImage, videoImage, video<br>Need a higher membership plan?<br>Different membership plans unlock different models, context windows, and speeds. Upgrade your plan →

Why did usage go up after the new model launched?After switching models, the context cache built earlier no longer hits on the new model, so that context has to be re-prefilled. Usage therefore looks higher right after switching. Recommended action:<br>Start a new session when using the new model : this gives better results and lower consumption.<br>Why do I still get a 401 with the correct model ID?When the requested capability exceeds your plan's entitlements, the server returns 401. Three common cases:<br>No K3 access : your plan is below Moderato and can't call k3, k3-256k — upgrade to Moderato or above.<br>No 1M access : on a Moderato plan, k3 supports up to 256K context; up to 1M context is available on Allegretto and higher tiers. k3-256k has a fixed 256K context limit.<br>No HighSpeed access : some plans don't include HighSpeed — upgrade to Allegretto or a higher tier to call kimi-for-coding-highspeed.<br>For the full error text and how to handle it, see the Error Reference.<br>Why isn't HighSpeed noticeably faster?Two common reasons:<br>Mistyped model ID : the HighSpeed ID must be kimi-for-coding-highspeed; a wrong value silently falls back to the standard kimi-for-coding — no error, no speedup.<br>Tools and scripts dominate : HighSpeed only speeds up model output . Tool calls (reading/writing files, running commands, etc.) and script execution are unaffected, so when they take up most of a turn the overall speedup feels small.<br>How to reduce the overhead of switching reasoning effort?Switching reasoning effort invalidates the context cache you've built up, so context that would have hit the cache must be re-prefilled. To avoid triggering re-prefill too often:<br>Pick an effort that fits the task and keep it consistent within a session;<br>When you genuinely need a different effort, start a new session rather than switching back and forth in a long session.<br>How to Switch Models ​<br>Usage notes<br>Start a new session when switching model IDs : switching models invalidates the context cache you've built up. We recommend starting a new session to get the best experience and avoid extra token consumption.<br>Fill in the Model ID, not the model version name : when calling a model, use one of the Model IDs from the table above (k3, k3-256k, kimi-for-coding, kimi-for-coding-highspeed). Entering a model version name like Kimi K3 or K2.7 Code will...

256k model context switching kimi code

Related Articles