You’ve already met our guide on implementing short-term conversational memory using LangChain, which is great for managing context inside a single chat window.

But life, therapy, and enterprise apps sprawl across days, weeks, and years. If our agents are doomed to goldfish-brain amnesia, users end up re-explaining everything from their favorite color to yesterday’s heartbreak.

It's time to graduate from sticky notes to an external brain. This guide will show you what long-term memory in LLMs really is and how to implement it using multiple techniques, like in-memory stores in LangChain, vector databases, Supermemory, etc.

Why Long-Term Memory Matters

Relying on raw token history means stuffing ever-growing chat logs back into each prompt. That quickly blows past context windows, hikes latency and cost, and still buries the signal beneath filler:

  • No prioritization. The model wastes attention on greetings while forgetting allergies or project deadlines
  • Brittle reasoning. Events like “last Friday’s outage” can’t be referenced without replaying the whole week
  • Repetitive answers. The agent re-derives stable traits (“prefers metric units”) every turn

Real-world apps need selective, structured, semantically indexed data points—facts, timelines, preferences, insights gleaned from analysis, just to skim the top.

Selecting, pruning, and handing these nuggets to a language model keeps prompts lean, token counts reasonable, and replies coherent across browser tabs and chat threads.

Therapy Assistant, Short-Term Context Edition

In part one, we built a therapy assistant that remembers within a session via trimming or summarizing. Close the tab or start a new chat session, however, and it’s obvious the bot needs some therapy of its own.

  • Personal details vanish, and users must retell their story
  • No cross-thread recall. Parallel chats never share insights
  • Token bloat. Replaying the whole transcript balloons prompts

In short, we need long-term memory.

Goal

Augment our therapy assistant with long-term, structured memory so sessions feel like genuine continuity. Users should see smooth recall of names, milestones, and coping strategies, plus tighter, more context-aware suggestions because the bot can finally connect the dots between Monday’s panic and Friday’s breakthrough.

We’ll layer in hybrid storage: JSON for crisp facts, a vector database for nuance, and a lightweight retrieval loop that feeds only the relevant slices back to the model.

Persistent Memory: The LangGraph Approach

LangGraph has built-in persistence to support long-term LLM memory using states, threads, and checkpointers. For short-term memory, LangGraph stores the list of messages to the chatbot in the state. Using threads, you can uniquely identify which user session the particular memory belongs to. Refer to our short-term memory guide for a full breakdown of how these work.

However, this memory cannot be shared across threads or across user sessions. For that, LangGraph implements something called stores.

A Store can well, store information as JSON documents across threads and make it available to the graph at any particular point, across different user sessions. Information is organized using namespaces, which basically are folders, but if you want to get technical, they are tuples that are used to uniquely identify a set of memories. Here's an example declaration:

user_id = "1"
namespace_for_memory = (user_id, "memories")

Memory stores are basically like databases managed by LangGraph. Hopefully, this diagram helps you visualize how stores work:

Every time our therapy agent receives a message, we hydrate working memory with just the bits that matter.

Building the Persistent Therapy Bot

Let's start building our therapy chatbot, and then extend it with persistent memory. We'll start by using the code for the chatbot from our previous tutorial.

Setup

Create a directory for the chatbot and open it in your IDE:

mkdir memory-chatbot

Install all the necessary Python libraries:

pip install --upgrade --quiet langchain langchain-openai langgraph

Note: Use pip3 in case just pip doesn’t work.

Basic Conversational Chatbot

Create a file file.py, and start with the following imports:

from langchain.schema import HumanMessage, SystemMessage
from langchain_core.messages import RemoveMessage, trim_messages
from langchain_openai import ChatOpenAI
from langgraph.checkpoint.memory import MemorySaver
from langgraph.graph import START, MessagesState, StateGraph

Set your OpenAI API key as an environment variable by running the following in your terminal:

export OPENAI_API_KEY=”YOUR_API_KEY”

Back to the Python file. Import the API key with:

os.environ.get("OPENAI_API_KEY")

Awesome! LangGraph uses graphs to represent and execute workflows. Each graph contains nodes, which are the individual steps or logic blocks that get executed. Create a function to represent that as follows:

def build_summarized(model: ChatOpenAI, trigger_len: int = 8) -> StateGraph:
    builder = StateGraph(state_schema=MessagesState)

def chat_node(state: MessagesState):
        system = SystemMessage(content="You're a kind therapy assistant.")
        summary_prompt = (
            "Distill the above chat messages into a single summary message. "
            "Include as many specific details as you can."
        )

if len(state["messages"]) >= trigger_len:
            history = state["messages"][:-1]
            last_user = state["messages"][-1]
            summary_msg = model.invoke(history + [HumanMessage(content=summary_prompt)])
            deletions = [RemoveMessage(id=m.id) for m in state["messages"]]
            human_msg = HumanMessage(content=last_user.content)
            response = model.invoke([system, summary_msg, human_msg])
            message_updates = [summary_msg, human_msg, response] + deletions
        else:
            response = model.invoke([system] + state["messages"])
            message_updates = response
        return {"messages": message_updates}

builder.add_node("chat", chat_node)
    builder.add_edge(START, "chat")
    return builder

Let's break this function down. The build_summarized function takes in a model and a trigger_len variable as input. The model is the LLM that'll be used under the hood. We'll touch on the trigger_len variable in a second.

The graph is initialized in the builder variable, and the chat_node function represents the chatbot's node in the graph. It takes the state as an argument, which contains a list of messages stored by the chatbot.

The chat_node function summarizes the conversation history if it passes a certain length, which is what the trigger_len variable expresses. Only the recent message history is stored in memory, and the rest is deleted.

Using an interactive CLI, we'll stitch it together so you can talk to the chatbot:

def interactive_chat(build_graph, thread_id: str):
    model = ChatOpenAI(model="gpt-4o-mini", temperature=0)
    graph = build_graph(model)
    chat_app = graph.compile(checkpointer=MemorySaver())

while True:
        try:
            user_input = input("You: ")
        except (EOFError, KeyboardInterrupt):
            print("\nExiting.")
            break

state_update = {"messages": [HumanMessage(content=user_input)]}
        result = chat_app.invoke(
            state_update, {"configurable": {"thread_id": thread_id}}
        )
        print("Bot:", result["messages"][-1].content)

And:

if __name__ == "__main__":
    interactive_chat(build_summarized, thread_id="demo")

MemorySaver keeps threads alive, thread_id lets you spawn parallel sessions. Right now, our chatbot doesn't have any persistent memory. This is just the basic implementation of a chatbot that you can actually talk to, and that can store the memory for individual sessions.

Memory Types for Human-like Interactions

Before we extend our chatbot with persistent memory, let's understand how effective memory mirrors human cognition:

  • Semantic (Facts): Zara's birthday is September 14th
  • Episodic (Events): Zara went through a breakup in May
  • Procedural (Steps and Habits): Zara uses breathing exercises when anxious

Memory Types Mapped to Therapy

Memory kind Therapy analogue Example
Semantic Stable facts “Name is Zara”
Episodic Lived events “Broke up in May”
Procedural Coping steps “Box breathing on bad days”

Designing the Memory Schema

Field Example Purpose
facts { "name": "Zara", "birthday": "1995-09-14" } Rapport, grounding
traits `[\