---
title: "Agent memory and context windows (2026): context compaction, summarization, and short-term and long-term memory for AI agents — Cadenya"
url: "https://cadenya.com/other-content/agent-memory-and-context-windows"
description: "How AI agents keep conversation state in 2026: LangGraph PostgresSaver checkpointers and Store, SummarizationMiddleware, trim and summarize messages, OpenAI Responses API compaction and Conversations, Anthropic compaction and context management, Mastra memory, Pydantic AI history processors, and a hosted agent loop like Cadenya."
---

An AI agent keeps conversation state in two places: short-term memory, which is the message history of one thread, and long-term memory, which is facts it keeps across threads. When the history outgrows the context window, context compaction summarizes or clears the older turns so the conversation can keep going.

So the real question is who stores the thread, who stores the facts, and who runs the compaction. Pick LangGraph when you want a Postgres checkpointer and a Store under your own graph. Pick the OpenAI Responses API or Anthropic’s compaction when you want the model provider to shrink the context for you. Pick Mastra, Pydantic AI, or the AI SDK when your agent already lives in that framework and you’re fine bringing the database. And if you’d rather run none of it, Cadenya is a hosted agent loop with a durable event log, context compaction, and memory layers built in, with nothing to host. [What it offers](#what-does-cadenya-offer-that-langgraph-mastra-and-the-model-apis-dont) is at the end.

Every price and limit below comes from each vendor’s own pages, checked on October 8, 2026.

## Agent memory options side by side

|  | LangGraph and LangChain | OpenAI Responses API | Anthropic Messages API | Mastra | Pydantic AI | Vercel AI SDK |
| --- | --- | --- | --- | --- | --- | --- |
| Short-term memory (one thread) | Checkpointers, keyed by thread_id | Conversations API or previous_response_id | You resend the messages | Message history | message_history you pass in | Messages you store with useChat |
| Long-term memory (across threads) | Stores, with semantic search | Your own, through tools | Memory tool, executed by your code | Working memory, semantic recall, observational memory | Your own code | Your own code |
| Context compaction | SummarizationMiddleware, trim_messages, summarize messages | Server-side compact_threshold, or /responses/compact | Compaction and context editing | Memory processors | History processors you write | pruneMessages |
| Where state lives | Your database: Postgres, MongoDB, Redis, Oracle, or SQLite | OpenAI, in a Conversation | Your app | A storage provider you configure, from 21 listed | Your app | Your app |
| What it costs | Free library, plus your database | Tokens. The whole chain bills as input | Tokens. Compaction adds a sampling step | Free library, plus your database | Free library, plus your database | Free library, plus your database |
| License | MIT | Hosted API | Hosted API | Mixed: ee/ folders have their own license | MIT | Apache 2.0 |

## Short-term memory vs long-term memory

Every option splits memory the same way, and LangGraph’s docs say it most plainly. “Checkpointers persist a thread’s graph state as checkpoints.” They’re short-term, thread-scoped memory. “Stores persist application-defined data outside the graph state.” They’re long-term, cross-thread memory: user preferences, facts, and shared knowledge.

**Short-term memory is the conversation. Long-term memory is what the agent should know next time.**

Mastra adds more kinds of long-term memory. Working memory “stores persistent, structured user data such as names, preferences, and goals.” Semantic recall finds old messages by meaning. Observational memory, which Mastra recommends, “uses background agents to maintain a dense observation log that replaces raw message history as it grows.”

The model APIs mostly stop at short-term memory. OpenAI’s Conversations API persists a thread “as a long-running object with its own durable identifier.” Anthropic’s memory tool gives Claude long-term files, but “the memory tool operates client-side: Claude requests file operations, and your application executes them.” (So the files live in your storage, and so does the code that reads them.)

## LangGraph add memory: the Postgres checkpointer

The in-memory checkpointer is fine for a demo and gone on restart. LangGraph’s docs are blunt: “`MemorySaver` and `InMemorySaver` store checkpoints in RAM. When the process restarts, all checkpoints are lost.” For production, they say to “use a checkpointer backed by a database,” and `PostgresSaver` is the first example:

```
from langgraph.checkpoint.postgres import PostgresSaver

DB_URI = "postgresql://postgres:postgres@localhost:5432/postgres?sslmode=disable"
with PostgresSaver.from_conn_string(DB_URI) as checkpointer:
    # checkpointer.setup()  # once, the first time, to create the tables
    graph = builder.compile(checkpointer=checkpointer)
```

Long-term memory gets the same treatment with `PostgresStore`, and the Store can do semantic search. MongoDB, Redis, and Oracle checkpointers exist too.

What do you sign up for? A database and its migrations. The persistence page warns that “over long conversations, checkpoints accumulate. This can increase latency and storage costs.” (LangGraph’s Agent Server handles persistence for you. [The LangSmith deployment pricing breakdown](/other-content/langsmith-deployment-alternatives-and-pricing) covers what that costs.)

## Context compaction when the context window fills

A long thread eventually holds more tokens than the model accepts. There are two fixes: drop old content, or summarize it. LangGraph lists both as “trim messages” and “summarize messages,” and names the cost of trimming: “you may lose information from culling of the message queue.”

### Summarization middleware in LangChain

LangChain packages summarization as middleware on `create_agent`:

```
from langchain.agents import create_agent
from langchain.agents.middleware import SummarizationMiddleware
from langgraph.checkpoint.memory import InMemorySaver

agent = create_agent(
    model="gpt-5.5",
    tools=[...],
    middleware=[
        SummarizationMiddleware(
            model="gpt-5.4-mini",
            trigger=("tokens", 4000),
            keep=("messages", 20)
        )
    ],
    checkpointer=InMemorySaver(),
)
```

When the thread passes 4,000 tokens, a smaller model summarizes it and the last 20 messages stay word for word.

### The other frameworks

Pydantic AI uses history processors, functions that rewrite the history before each request. Its docs show one that summarizes the oldest messages with a cheaper model, and they add a warning: “you need to make sure that tool calls and returns are paired, otherwise the LLM may return an error.” Mastra has memory processors that “filter, trim, or prioritize content.” The AI SDK has `pruneMessages`, which strips reasoning and old tool calls before you call the model. It prunes. It doesn’t summarize.

## OpenAI Responses API compaction and Anthropic context management

Both model providers now compact on their side.

OpenAI’s Responses API takes a `context_management` setting on an ordinary request:

```
response = client.responses.create(
    model="gpt-5.3-codex",
    input=conversation,
    store=False,
    context_management=[{"type": "compaction", "compact_threshold": 200000}],
)
```

“When the rendered token count crosses the configured threshold, the server runs server-side compaction.” The result is a compaction item that’s “opaque and not intended to be human-interpretable.” For explicit control there’s a standalone `/responses/compact` endpoint, which “is fully stateless.” One thing about the bill: “all previous input tokens for responses in the chain are billed as input tokens,” even with `previous_response_id`.

Anthropic offers two kinds of compaction. Compaction at a token threshold uses the `compact-2026-01-12` header and the `compact_20260112` edit. Its trigger defaults to 150,000 input tokens and “must be at least 50,000 tokens.” Compaction on demand uses `compact-2026-09-04` and lets your code decide when. Either way, “compaction requires an additional sampling step, which contributes to rate limits and billing.” Context editing is the other half: `clear_tool_uses_20250919` clears old tool results instead of summarizing them.

**Provider compaction shrinks the context. It doesn’t store your conversation for you.** On Anthropic, you still resend the messages each turn. On OpenAI, a Conversation stores the thread.

## Where conversation state is stored, and what it costs

Every framework in the table is free, and every one hands you the database. Mastra’s docs say memory “requires a storage provider to persist message history” and suggest PostgreSQL “for most teams.” Pydantic AI says it outright: “Deciding where those bytes live, which conversation they belong to, and when to reload them is left to your application.” The AI SDK’s persistence guide has you write your own chat store.

So self-hosting agent memory means a database, a schema, and the summarization code. And when the model changes, the token threshold you tuned for one context window may be wrong for the next.

The model APIs store state for you only in OpenAI’s case, and you pay for it in tokens: the whole chain bills as input on every turn. If the thread has to survive a crash or a deploy mid-turn, that’s a different problem: [durable execution](/other-content/restate-vs-temporal-for-ai-agents), which [the hosted agent runtimes comparison](/other-content/hosted-agent-runtimes-2026) covers.

## What does Cadenya offer that LangGraph, Mastra, and the model APIs don’t?

LangGraph, Mastra, and the model APIs each hand you one piece of agent memory and leave you to store the rest. **Cadenya keeps the conversation, compacts the context window, and serves long-term memory inside the loop it runs, with no database to host.** Here’s what that covers:

1.  **A durable event log per objective.** “Every Objective carries a durable event log of its agent loop.” List the events in sequence to replay a run from its first message, or stream them while it works. There’s no checkpointer to pick and no `setup()` to call.
2.  **Conversations that pick up where they left off.** When an objective is waiting, send a Continue message and the loop resumes with its history. In the widget, sessions with the same tenant and subject share conversation history across tabs and newly minted sessions.
3.  **Context compaction with no summarizer to write.** Set a compaction config on the variation. When input tokens reach 75% of the model’s context window (you can change that), Cadenya compacts. With tool result clearing on, it runs first and keeps the 2 most recent results intact by default, then summarization runs. With neither strategy set, Cadenya summarizes with its default prompt.
4.  **Your own summarization instructions.** Replace the default prompt with what must survive, such as “Preserve all code snippets, variable names, and technical decisions.” You can also compact an objective on demand through the API, with an override for that one call.
5.  **Every compaction on the record.** A `contextWindowCompacted` event names the strategies used and the number of messages compacted. The API lists an objective’s last five context windows with their token counts, and a diagnostics call breaks down the most recent request by system prompt, memory, tools, and messages, including cached input tokens.
6.  **Long-term memory layers.** A memory layer holds entries of text or uploaded files, each with a key. Attach layers to a variation, and the agent sees each entry’s key and description in its system prompt, then reads the full entry only when it needs it. Layers stack in a cascade, so a specific layer answers before a general one. Memory layers come with the Growth plan.
7.  **Episodic memory across objectives.** Give objectives the same `episodic_key`, and the agent gets `store_memory`, `get_memory`, and `search_memory` tools. Set a TTL, and each new objective with that key pushes the expiration forward. Every `memoryRead` shows up in the event log.
8.  **Nothing else to run.** No checkpointer database, no memory service, and no summarizer job to host or upgrade. The loop and its memory run on Cadenya, so your app stays a stateless web tier and your deploys never touch a conversation.
9.  **One price, and your model keys.** Free covers 1,000 loops a month. Growth is $49 a month for 50,000 loops and memory layers, then $0.0015 a loop. A loop is one LLM request. SDKs for TypeScript, Python, Ruby, and Go are all at 1.7.0 under Apache-2.0.

[How objectives work](/docs/guides/the-basics/objectives) covers the event log, Continue, and compaction, and the free plan needs no card.