Explore how agents remember information within a run and across runs — ephemeral context, key-value memory stores, and vector memory for semantic retrieval.
In-run context
Within a single agent run, every step has access to the outputs of all previous steps through the expression system. This constitutes the agent's working memory for that run: a structured JSON object that grows as each step completes. The LLM node additionally maintains a conversation history — an ordered list of messages exchanged during the run — which is automatically included in each subsequent LLM call unless you configure the node to use a fresh context. In-run context is ephemeral: it exists only for the duration of the run and is discarded when the run completes, though it remains visible in the run logs for debugging.
Persistent key-value memory
For information that needs to survive across multiple runs — such as a user's preferences, the last-processed record ID, or an accumulated count — use the Memory Store action. The Memory Store is a scoped key-value database tied to your agent: you can read a key at the start of a run, update it during the run, and read the updated value in the next run. Keys are namespaced by agent ID by default, but you can share a namespace between agents in the same workspace to pass state between them. Memory Store entries have a configurable TTL; entries without a TTL persist indefinitely until explicitly deleted.
Vector memory for semantic retrieval
When an agent needs to find relevant information from a large corpus — documentation, past conversations, customer records — rather than retrieve a specific key, use the Vector Memory node. This node embeds a query string using the same embedding model configured for your workspace, searches a vector index for the most similar stored entries, and returns the top-N results as structured text that the LLM can incorporate into its reasoning. You populate the vector index using the Store in Vector Memory action, which accepts any text content and optional metadata for filtering. Vector memory is ideal for retrieval-augmented generation (RAG) patterns.
Managing context window size
Every LLM node has a context window limit measured in tokens. If the accumulated conversation history or retrieved memory content exceeds this limit, the LLM call will fail or produce truncated responses. Cotonity displays a real-time token counter in the LLM node's configuration panel. To stay within limits, consider summarizing earlier parts of the conversation history using the Summarize History action, filtering vector memory results to the top two or three most relevant entries, and avoiding passing large raw documents into the prompt when only a subset is needed. The Run Logs show the token count for every LLM call, which helps you tune these settings over time.