After building several AI chat projects, I spent time exploring how to retain useful context over longer conversations. Two common approaches are a separate summary process and a memory section included in the model's replies.
A context window is measured in tokens, and the prompt, conversation history, retrieved information, and response all compete for space. There is no fixed conversion from one million tokens to a particular number of Chinese characters. A large window also does not guarantee perfect recall of everything inside it.
Approach 1: a separate rolling summary
A chat model handles the conversation while another model, or a separate call to the same model, updates a summary. The application stores that summary and includes relevant information in later prompts. Updating every few turns can reduce the extra cost compared with summarizing every message.
A cheaper model can handle the summary while a more expensive model handles role-play. The stored material can later grow into a searchable knowledge base, but retrieval and summarization still require careful implementation. Summaries can lose details or preserve an earlier mistake.
For AstrBot, the original article recommends astrbot_plugin_mnemosyne as one project to explore.

Approach 2: include memory in each reply
A simpler setup asks the model to maintain short-term and long-term notes in its output. The following example creates a collapsible section. Replace its long-term material with your own character and story facts:
<details style="background-color: #e6f7ff; border: 2px dashed #69b1ff; border-radius: 10px; padding: 12px 15px; margin-top: 15px; box-shadow: 0 2px 8px rgba(24, 144, 255, 0.15);">
<summary style="cursor: pointer; font-weight: bold; color: #0958d9; outline: none; user-select: none; font-size: 15px;">Conversation memory</summary>
<p style="margin-top: 12px; margin-bottom: 0; font-size: 13px; line-height: 1.6; color: #1677ff; border-top: 1px solid #91caff; padding-top: 10px;">
<strong>[Short-term memory]</strong>(0/5)<br>
(List recent events here. Compress when the list reaches five entries.)<br><br>
<strong>[Long-term memory]</strong><br>
Replace this section with the story, character, and facts that should remain important throughout your chat.
</p>
</details>
Explain when recent events should be compressed and which facts should remain. This is easy to add to a character prompt, but repeatedly generating the notes costs output tokens, and later requests may pay to read the same notes again.
Choose the tradeoff that fits the project
Separate storage and retrieval offer more control but require more work. An in-message summary is convenient but shares the same finite context window and is harder to audit over time. Neither creates unlimited or human-like memory automatically.
I find these techniques useful for preserving personality and story continuity. Check summaries against the conversation and let users correct important facts. That matters more than simply increasing the amount of text labeled “memory.”
Adapted from the original Chinese article, published on April 30, 2026.

Comments NOTHING