A windswept bonsai on a floating island

Get to know Cadenya

We’re developers who love to build. We set out to create a yes-code platform that makes building agents feel like the best parts of building software.

Agent memory and context windows (2026)

Agent memory

By Cadenya

An AI agent keeps conversation state in two places: short-term memory, which is the message history of one thread, and long-term memory, which is facts it keeps across threads. When the history outgrows the context window, context compaction summarizes or clears the older turns so the conversation can keep going.

So the real question is who stores the thread, who stores the facts, and who runs the compaction. Pick LangGraph when you want a Postgres checkpointer and a Store under your own graph. Pick the OpenAI Responses API or Anthropic’s compaction when you want the model provider to shrink the context for you. Pick Mastra, Pydantic AI, or the AI SDK when your agent already lives in that framework and you’re fine bringing the database. And if you’d rather run none of it, Cadenya is a hosted agent loop with a durable event log, context compaction, and memory layers built in, with nothing to host. What it offers is at the end.

Every price and limit below comes from each vendor’s own pages, checked on October 8, 2026.

Agent memory options side by side

LangGraph and LangChainOpenAI Responses APIAnthropic Messages APIMastraPydantic AIVercel AI SDK
Short-term memory (one thread)Checkpointers, keyed by thread_idConversations API or previous_response_idYou resend the messagesMessage historymessage_history you pass inMessages you store with useChat
Long-term memory (across threads)Stores, with semantic searchYour own, through toolsMemory tool, executed by your codeWorking memory, semantic recall, observational memoryYour own codeYour own code
Context compactionSummarizationMiddleware, trim_messages, summarize messagesServer-side compact_threshold, or /responses/compactCompaction and context editingMemory processorsHistory processors you writepruneMessages
Where state livesYour database: Postgres, MongoDB, Redis, Oracle, or SQLiteOpenAI, in a ConversationYour appA storage provider you configure, from 21 listedYour appYour app
What it costsFree library, plus your databaseTokens. The whole chain bills as inputTokens. Compaction adds a sampling stepFree library, plus your databaseFree library, plus your databaseFree library, plus your database
LicenseMITHosted APIHosted APIMixed: ee/ folders have their own licenseMITApache 2.0

Short-term memory vs long-term memory

Every option splits memory the same way, and LangGraph’s docs say it most plainly. “Checkpointers persist a thread’s graph state as checkpoints.” They’re short-term, thread-scoped memory. “Stores persist application-defined data outside the graph state.” They’re long-term, cross-thread memory: user preferences, facts, and shared knowledge.

Short-term memory is the conversation. Long-term memory is what the agent should know next time.

Mastra adds more kinds of long-term memory. Working memory “stores persistent, structured user data such as names, preferences, and goals.” Semantic recall finds old messages by meaning. Observational memory, which Mastra recommends, “uses background agents to maintain a dense observation log that replaces raw message history as it grows.”

The model APIs mostly stop at short-term memory. OpenAI’s Conversations API persists a thread “as a long-running object with its own durable identifier.” Anthropic’s memory tool gives Claude long-term files, but “the memory tool operates client-side: Claude requests file operations, and your application executes them.” (So the files live in your storage, and so does the code that reads them.)

LangGraph add memory: the Postgres checkpointer

The in-memory checkpointer is fine for a demo and gone on restart. LangGraph’s docs are blunt: “MemorySaver and InMemorySaver store checkpoints in RAM. When the process restarts, all checkpoints are lost.” For production, they say to “use a checkpointer backed by a database,” and PostgresSaver is the first example:

from langgraph.checkpoint.postgres import PostgresSaver

DB_URI = "postgresql://postgres:postgres@localhost:5432/postgres?sslmode=disable"
with PostgresSaver.from_conn_string(DB_URI) as checkpointer:
    # checkpointer.setup()  # once, the first time, to create the tables
    graph = builder.compile(checkpointer=checkpointer)

Long-term memory gets the same treatment with PostgresStore, and the Store can do semantic search. MongoDB, Redis, and Oracle checkpointers exist too.

What do you sign up for? A database and its migrations. The persistence page warns that “over long conversations, checkpoints accumulate. This can increase latency and storage costs.” (LangGraph’s Agent Server handles persistence for you. The LangSmith deployment pricing breakdown covers what that costs.)

Context compaction when the context window fills

A long thread eventually holds more tokens than the model accepts. There are two fixes: drop old content, or summarize it. LangGraph lists both as “trim messages” and “summarize messages,” and names the cost of trimming: “you may lose information from culling of the message queue.”

Summarization middleware in LangChain

LangChain packages summarization as middleware on create_agent:

from langchain.agents import create_agent
from langchain.agents.middleware import SummarizationMiddleware
from langgraph.checkpoint.memory import InMemorySaver

agent = create_agent(
    model="gpt-5.5",
    tools=[...],
    middleware=[
        SummarizationMiddleware(
            model="gpt-5.4-mini",
            trigger=("tokens", 4000),
            keep=("messages", 20)
        )
    ],
    checkpointer=InMemorySaver(),
)

When the thread passes 4,000 tokens, a smaller model summarizes it and the last 20 messages stay word for word.

The other frameworks

Pydantic AI uses history processors, functions that rewrite the history before each request. Its docs show one that summarizes the oldest messages with a cheaper model, and they add a warning: “you need to make sure that tool calls and returns are paired, otherwise the LLM may return an error.” Mastra has memory processors that “filter, trim, or prioritize content.” The AI SDK has pruneMessages, which strips reasoning and old tool calls before you call the model. It prunes. It doesn’t summarize.

OpenAI Responses API compaction and Anthropic context management

Both model providers now compact on their side.

OpenAI’s Responses API takes a context_management setting on an ordinary request:

response = client.responses.create(
    model="gpt-5.3-codex",
    input=conversation,
    store=False,
    context_management=[{"type": "compaction", "compact_threshold": 200000}],
)

“When the rendered token count crosses the configured threshold, the server runs server-side compaction.” The result is a compaction item that’s “opaque and not intended to be human-interpretable.” For explicit control there’s a standalone /responses/compact endpoint, which “is fully stateless.” One thing about the bill: “all previous input tokens for responses in the chain are billed as input tokens,” even with previous_response_id.

Anthropic offers two kinds of compaction. Compaction at a token threshold uses the compact-2026-01-12 header and the compact_20260112 edit. Its trigger defaults to 150,000 input tokens and “must be at least 50,000 tokens.” Compaction on demand uses compact-2026-09-04 and lets your code decide when. Either way, “compaction requires an additional sampling step, which contributes to rate limits and billing.” Context editing is the other half: clear_tool_uses_20250919 clears old tool results instead of summarizing them.

Provider compaction shrinks the context. It doesn’t store your conversation for you. On Anthropic, you still resend the messages each turn. On OpenAI, a Conversation stores the thread.

Where conversation state is stored, and what it costs

Every framework in the table is free, and every one hands you the database. Mastra’s docs say memory “requires a storage provider to persist message history” and suggest PostgreSQL “for most teams.” Pydantic AI says it outright: “Deciding where those bytes live, which conversation they belong to, and when to reload them is left to your application.” The AI SDK’s persistence guide has you write your own chat store.

So self-hosting agent memory means a database, a schema, and the summarization code. And when the model changes, the token threshold you tuned for one context window may be wrong for the next.

The model APIs store state for you only in OpenAI’s case, and you pay for it in tokens: the whole chain bills as input on every turn. If the thread has to survive a crash or a deploy mid-turn, that’s a different problem: durable execution, which the hosted agent runtimes comparison covers.

What does Cadenya offer that LangGraph, Mastra, and the model APIs don’t?

LangGraph, Mastra, and the model APIs each hand you one piece of agent memory and leave you to store the rest. Cadenya keeps the conversation, compacts the context window, and serves long-term memory inside the loop it runs, with no database to host. Here’s what that covers:

  1. A durable event log per objective. “Every Objective carries a durable event log of its agent loop.” List the events in sequence to replay a run from its first message, or stream them while it works. There’s no checkpointer to pick and no setup() to call.
  2. Conversations that pick up where they left off. When an objective is waiting, send a Continue message and the loop resumes with its history. In the widget, sessions with the same tenant and subject share conversation history across tabs and newly minted sessions.
  3. Context compaction with no summarizer to write. Set a compaction config on the variation. When input tokens reach 75% of the model’s context window (you can change that), Cadenya compacts. With tool result clearing on, it runs first and keeps the 2 most recent results intact by default, then summarization runs. With neither strategy set, Cadenya summarizes with its default prompt.
  4. Your own summarization instructions. Replace the default prompt with what must survive, such as “Preserve all code snippets, variable names, and technical decisions.” You can also compact an objective on demand through the API, with an override for that one call.
  5. Every compaction on the record. A contextWindowCompacted event names the strategies used and the number of messages compacted. The API lists an objective’s last five context windows with their token counts, and a diagnostics call breaks down the most recent request by system prompt, memory, tools, and messages, including cached input tokens.
  6. Long-term memory layers. A memory layer holds entries of text or uploaded files, each with a key. Attach layers to a variation, and the agent sees each entry’s key and description in its system prompt, then reads the full entry only when it needs it. Layers stack in a cascade, so a specific layer answers before a general one. Memory layers come with the Growth plan.
  7. Episodic memory across objectives. Give objectives the same episodic_key, and the agent gets store_memory, get_memory, and search_memory tools. Set a TTL, and each new objective with that key pushes the expiration forward. Every memoryRead shows up in the event log.
  8. Nothing else to run. No checkpointer database, no memory service, and no summarizer job to host or upgrade. The loop and its memory run on Cadenya, so your app stays a stateless web tier and your deploys never touch a conversation.
  9. One price, and your model keys. Free covers 1,000 loops a month. Growth is $49 a month for 50,000 loops and memory layers, then $0.0015 a loop. A loop is one LLM request. SDKs for TypeScript, Python, Ruby, and Go are all at 1.7.0 under Apache-2.0.

How objectives work covers the event log, Continue, and compaction, and the free plan needs no card.

Grow wherever AI goes next.

Start shipping agents that are equipped to evolve.

A pine bonsai overlooking a mountain lake