Dual-track memory for conversation and behavior, refined layer by layer; AutoDream digests behavior into insights while idle.
Storing chat logs is not memory
Everyone says agents need memory. But most of what passes for "memory" today just stores chat logs and searches them back later: what users said becomes rows of text waiting to be retrieved — like scrubbing through security footage, not like remembering. That's a database, not memory. The essence of memory isn't storage and retrieval — it's processing: what gets forgotten, refined, and corrected matters as much as what stays.
What is GUMem
GUMem is Qoni's memory layer, and it delivers that processing as a managed service: it remembers not just what users said but what they did; raw information is refined, layer by layer, into judgments ever closer to "who this person is"; what should fade, fades by rule; and recall plans its retrieval around the question and hands over answers with evidence.
The results have public backing: on LoCoMo, the long-conversation memory benchmark, GUMem scored 92.9% (see the website). Behind the number is a plain sentence: across dozens of sessions and several weeks, it still remembers who the user is and what they want.
The six things GUMem does
- Dual-track capture: conversation is one track, behavior is the other (what they clicked, chose, abandoned) — physically separated from capture to processing, each with a full pipeline;
- Layered refinement: conversation → facts → summaries and topics (both part of "condense"), more concentrated at every level; every memory carries a source pointer and a confidence score, traceable back to the sentence or the click where it was born;
- Decay by half-life: identity memories are measured in decades, passing moods vanish in two days; activity drives recall, and stale preferences don't get believed forever;
- Recall that plans: retrieval isn't one vector search but a query agent that plans its path by question type — and must say so when evidence is thin;
- AutoDream: while the agent is idle, it automatically analyzes the behavior logs nobody reads and turns them into usable insights;
- Memory sovereignty and governance: users can see, edit, and delete every memory; enterprise rules attach at explicit hooks, and the data scope is enforced by the backend.
GUMem's two scenarios: shopping guidance and cross-session work
E-commerce guidance. The user says "just browsing," but the behavior track tells another story: repeatedly opening light-colored sneakers, filtering size 42, clicking into carbon-plate runners and backing out within seconds. The preferences the conversation track never held, the behavior track supplies: light colors, size 42, skip the carbon plates. The next recommendation, the agent already knows. That's what dual-track memory is for: the half of the story people never say out loud.
Knowledge work across sessions. In a brand-new session the user asks: "compare this vendor against the one I liked last week." No pasted context, no re-briefing. The query agent recalls across sessions: what last week's vendor was called, that the user cared most about deployment risk, that he prefers concise answers — so the answer lands in one step: "cheaper than last week's, but weaker on audit export and deployment controls." Every supporting memory carries its source pointer — some from conversation, some from the click stream.
GUMem's six kinds of memory, three kinds of storage
Raw input first lands in working memory, the starting point of all processing. It then aggregates into episodic memory: what happened, and in what context. Episodes settle upward along two paths: semantic memory answers "who this person is, what they prefer," and procedural memory answers "how this task gets done." The fifth kind comes from user feedback — satisfaction, frustration, correction — producing sentiment memory: it doesn't describe facts, it describes how the user feels about them. The sixth is associative memory, weaving a web between the others: support, contradiction, extension, causation, similarity, co-occurrence — six kinds of edges that let memories corroborate and correct each other.
The storage layer in one sentence: the vector index handles semantic similarity, the graph store handles relational reasoning, and the transactional database handles the ledger and idempotency; the three are independent, and each can be rebuilt on its own.
GUMem's processing order: recall, generate, write back
The pipeline runs in a deliberate order: recall first, then generate, then write back. Before answering, the agent draws on memory that settled before this turn; only after the answer is generated does this turn's new information enter the pipeline. The order cannot flip — flip it and a subtle disease sets in: the model treats what it just said as "the user's historical fact," one turn's small bias gets written into memory, recalled as fact the next turn, compounding. A memory system's first duty is not to remember more. It is not to pollute itself.
"Layered refinement" isn't mood language. The pipeline runs five stations:
- Collect: messages, clicks, searches, and tool results land as raw events;
- Extract: events become dated facts with a pointer back to their source;
- Condense: repeated facts roll up into summaries, then settle into stable topics — summaries and topics both belong to this one "condense" station;
- Recall: the current task reads the top layer first, the layers beneath hanging along;
- Forget: retention rules, corrections, and deletions all act on the same store.
The fields exist for verifiability: any recalled memory can trace through its anchor back to the original event. A memory system that can't explain where its beliefs came from is a liability inside an enterprise, not an asset.
Memory's five half-life tiers: what should be forgotten, gets forgotten on schedule
GUMem assigns every memory a half-life: identity memories (he's an engineer) are measured in decades, stable relationships in years, preferences and habits in quarters, temporary plans fade in two weeks, passing moods vanish in two days. A memory's activity = confidence × exponential decay parameterized on the half-life: high-activity memories get recalled first, and those that sink to the bottom exit the stage on their own.
The other side of forgetting is reinforcement and overturning, both on the record: a preference confirmed again and again gets its confidence raised and its decay slowed; when the user says "I've stopped drinking americanos, I drink lattes now," the old memory isn't physically deleted — it's marked invalidated and pointed at the new memory that overturned it, and every rewrite leaves a complete evidence chain. This is audit-friendly forgetting: memory can be corrected, but the correction itself is always on the record.
The query agent: recall is not one vector search
Storing well is half the job; recalling precisely is the other half. In most memory systems, "recall" is one vector search: take the question, compare against the store, hand over the closest matches. GUMem's recall is a query agent that plans: it carries a set of memory-access tools (recent messages, session context, record lookup, vector search, graph relations) and decides which memory to consult first by the question's type — "how do I" goes to procedural memory, "what happened" goes to episodic, relationship questions walk the graph, preference questions check the sentiment track first. Evidence is cross-checked across types, then assembled into exactly the context this task needs. No more, no less.
It also obeys an honesty protocol hard-coded into the system: only memories actually retrieved count as evidence — fabrication is forbidden; when evidence is thin, it must say so instead of guessing; when evidence conflicts, it resolves by a fixed priority (the user's explicit correction beats an inferred preference, the current session beats old memory, a specific constraint beats a broad liking), and what can't be resolved gets laid out on the table. And one hard boundary: the data scope of project and user is enforced by the system — the query agent passes business parameters only, and doesn't even have the power to construct a query statement.
AutoDream: digesting behavior logs into insights while the agent is idle
Qoni ships a mechanism that works while the agent is idle, called AutoDream: it automatically analyzes the behavior stream a user accumulates and turns it into genuine insights — "this user always orders late at night," "three straight weeks of searching flights to Tokyo." What makes it different: this isn't a report an engineer pre-wrote. The agent decides what to analyze, writes the analysis code itself, and runs it to a conclusion. The questions nobody thought to ask, it asks first.
Three pieces of engineering behind it, worth more than the metaphor:
- The model never sees raw data. Behavior logs first pass through a deterministic statistics layer that involves no model at all, compressed into a "data digest"; natural-language content keeps only statistical features like length, and the original text never leaves the deterministic statistics layer. The model sees the digest, and proposes analyses based on it;
- The agent's code runs in a sandbox. Generated statistics code cannot import dependencies or touch storage and network — it can only call a controlled set of statistical interfaces;
- The insights carry evidence. Every insight is structured: metric, dimension, current value, baseline, delta, evidence, and confidence.
Memory sovereignty: users can see, edit, and delete
The stronger the memory, the more this question matters: does the user know what the agent has remembered? On the user's side, GUMem's answer is three verbs:
- See: every memory sits in plain sight, with its source and confidence;
- Edit and delete: correct what's wrong, delete what shouldn't be remembered — memory belongs to the user, not the model;
- Watch it form: after writing a batch of messages, the processing can be subscribed to in real time (facts extracted, summaries generated, topics updated); the event stream returns stages and counts only, never the original content, with key-like strings masked.
On the enterprise side, the memory layer holds the system's highest concentration of personal information, so business rules and compliance requirements attach at three hooks — each deliberately given less authority than the last: before writing, data can be rewritten (scrub sensitive fields, enrich business context); before the model, instructions can only be appended (never altering memory itself); after the model, read-only notification (audit pipelines, CRM sync). The real hard boundary sits in the system-enforced data scope: project and user scope is set by the backend, and no hook, no query crosses it.
Recall is also wired into GenAuth's identity boundary: when an agent mounts memory, it declares which user it acts for and which actions it may take (recall only, no writes, for example). What it holds isn't "the key to the memory store" — it's a key that recalls exactly one user. Every recall runs on the same identity, boundary, and audit rails as tool calls and web actions.
Integrating GUMem: three calls for chat, three actions for behavior
The SDK ships as @qoniai/qoni (Node 18+, ESM / CommonJS dual build, full TypeScript types), initialized with the AK / SK from Qoni Console. Step one is a delegation: GUMem's scopes are fine-grained (gumem.memory:read / gumem.memory:write / gumem.message:write), and the common session-write-recall path has a preset bundle, QoniScopeBundles.GUMEM_SESSION_RECALL.
Storing and recalling the conversation track takes three calls:
import { Qoni } from "@qoniai/qoni";
const qoni = new Qoni({ accessKey, secretKey });
const { data } = await qoni.delegateToken({
user: { id: userId },
scopes: ["gumem.memory:read", "gumem.memory:write", "gumem.message:write"],
});
// Chat ingestion: open a session, write messages
await qoni.gumem.createSession({
token: data.token,
userId,
sessionId: "daily-assistant",
title: "Daily assistant memory",
});
await qoni.gumem.addMessages({
token: data.token,
sessionId: "daily-assistant",
messages: [{ role: "user", content: "Keep daily reports short — next actions only." }],
});
// Recall: not keyword search, but task-oriented retrieval
const context = await qoni.gumem.recall({
token: data.token,
sessionId: "daily-assistant",
query: "What confirmed preferences does this user have about reply style?",
details: true,
});
// context.data is ready context that can drop straight into prompt assembly
The behavior track is three symmetric actions: qoni.gumem.actions.record (store), actions.recall, and actions.stream (subscribe) — clicks, choices, and abandonments land here, physically separate from the conversation track. AutoDream has no "integration code" at all: once behavior lands via actions.record, idle-time analysis runs on its own — not one extra line. Delegation narrowing bites here too: an agent holding only gumem.memory:read gets rejected on any write. That's not a documentation convention — it's a credential boundary.
One quick-reference table to gather everything this piece covered:
| Capability | In one line |
|---|---|
| Dual-track memory | Conversation track + behavior track, physically separated from capture to processing |
| Six kinds of memory | Working / episodic / semantic / procedural / sentiment, plus the associative web |
| Refinement pipeline | Five stations — Collect → Extract → Condense → Recall → Forget; every layer's product has fields and anchors |
| Half-life forgetting | Five decay tiers; reinforcement on confirmation; overturning leaves a chain |
| Three stores | Vectors for similarity, graph for relations, transactions for the ledger — each independently rebuildable |
| The query agent | Plans the recall path by question type, cross-checks evidence across types |
| The honesty protocol | Only retrieved memories count; thin evidence must be declared; conflicts resolve by fixed priority |
| AutoDream | Analyzes the behavior stream while idle: the model sees digests, code runs sandboxed, insights carry evidence |
| Memory sovereignty | See, edit, delete; processing observable without leaking content |
| Governance & delegation | Three hooks with shrinking authority; backend-enforced scope; recall rides GenAuth delegation |
| SDK | @qoniai/qoni; three calls for sessions, writes, recall; three symmetric behavior actions |
Not storing every word users said — understanding who they are.
GUMem is currently in private build and will open in stages with Qoni. Join the waitlist on the website.
对话与行为双轨记忆,逐层提炼;AutoDream 趁空闲把行为消化成洞察。
存下聊天记录,不等于有记忆
大家都说 Agent 需要记忆。但今天大多数所谓的「记忆」,只是把聊天记录存起来、要用的时候再搜出来:用户说过的话变成一条条等着被检索的文本,像翻监控录像,不像回忆。那是数据库,不是记忆。记忆的本质不是存取,是加工:遗忘掉的、提炼出的、纠正过的,和留下来的一样重要。
GUMem 是什么
GUMem 是 Qoni 的记忆层,把这套「加工」做成了托管服务:不只记用户说过什么,还记他做过什么;原始信息逐层提炼成越来越接近「这个人是谁」的判断;该忘的按规律忘掉;取用的时候按问题规划检索、带着证据交差。
效果有公开测试背书:在长对话记忆测试 LoCoMo 上,GUMem 拿到了 92.9% 的成绩(官网可查)。这个分数背后是一句朴素的话:隔着几十个会话、隔着几个星期,它依然记得用户是谁、要什么。
GUMem 的六项能力
- 双轨捕获:对话是一条轨道,行为是另一条(点过什么、选过什么、放弃过什么),从采集到加工物理分离,各有完整管线;
- 逐层提炼:对话 → 事实 → 摘要与主题(同属「归纳」),越往上越浓缩;每条记忆带来源指针和可信度,能一路追溯回它诞生的那句话、那次点击;
- 按半衰期遗忘:身份类记忆以十年计,瞬时情绪两天即逝;活性决定召回,过时的偏好不会永远当真;
- 会规划的召回:取用不是一次向量搜索,而是一个按问题类型规划检索路径的查询 Agent,证据不够必须明说;
- AutoDream:趁 Agent 空闲,把没人看的行为日志自动分析成能用的洞察;
- 记忆主权与治理:用户能看、能改、能删每一条记忆;企业规则有明确挂点,数据范围由后端强制。
GUMem 的两个落地场景:导购与跨会话
电商导购。 用户嘴上说「随便看看」,但行为轨里写着另一个故事:反复打开浅色球鞋、筛选 42 码、点开碳板跑鞋又几秒退出。对话轨里没有的偏好,行为轨替他说了:浅色、42 码、避开碳板。下一次推荐,Agent 已经懂了。这就是双轨记忆的意义:补上「人不会说出口」的那一半。
跨会话的知识工作。 用户在一个全新的会话里问:「帮我比一比这家供应商,跟我上周看中的那家怎么样?」没有粘贴上下文,没有重新交代背景。查询 Agent 跨会话召回:上周那家叫什么、用户当时最关心部署风险、他喜欢简洁的回答风格,于是答案一步到位:「比上周的那家便宜,但审计导出和部署管控更弱。」每一条支撑记忆都带来源凭证:有的来自对话,有的来自点击流。
GUMem 的六种记忆,三种存储
原始输入先进入工作记忆,这是所有加工的起点;随后聚合成情节记忆(发生过什么、在什么语境下发生);情节再往上沉淀,分出两条去向:语义记忆回答「这个人是谁、偏好什么」,流程记忆回答「这件事该怎么做」。第五种来自用户的反馈(满意、不满、纠正),是情感记忆:它不描述事实,描述的是用户对事实的态度。第六种是关联记忆,在其他记忆之间织网:支持、矛盾、引申、因果、相似、共用,六种关系边让记忆可以互相印证、互相纠错。
存储层一句话说完分工:向量索引管语义相似,图存储管关系推理,事务数据库管账本与幂等;三者互不依赖,坏了都能独立重建。
GUMem 的加工顺序:先召回,再生成,再写回
管线的运转顺序是一条刻意的纪律:先召回,再生成,再写回。 Agent 回答之前,先取用本轮之前已经沉淀好的记忆;回答生成之后,这一轮的新信息才进入管线加工。顺序不能倒,倒了会出一种很隐蔽的病:模型把自己刚说的话当成「用户的历史事实」,一轮的小偏差写进记忆、下一轮当作事实召回,越聊越偏,而且是复利的。记忆系统的第一要务不是记得多,是不污染自己。
「逐层提炼」不是一句气氛话。流水线是五站:
- Collect:消息、点击、搜索与工具结果,先作为原始事件落库;
- Extract:事件整理成一条带日期的 Fact,并留下回指来源的指针;
- Condense:反复出现的 Fact 归纳成 Summary,再沉淀为稳定的 Topic,摘要与主题同属「归纳」这一站;
- Recall:当前任务先取最上面一层,底下几层一并挂着;
- Forget:保留规则、内容修正与删除请求,都作用在同一个库上。
字段的意义在于可核验:任何一条被召回的记忆,都能顺着锚点回到它的原始事件。一套说不清来历的记忆,在企业里是负债,不是资产。
记忆的五档半衰期:该忘的按规律忘
GUMem 给每条记忆分配一个半衰期:身份类记忆(他是工程师)以十年计,稳定关系以年计,偏好习惯以季度计,临时计划两周就淡,瞬时情绪两天即逝。 一条记忆的活性 = 可信度 × 以半衰期为参数的指数衰减:活性高的优先被召回,活性沉底的自动退出舞台。
遗忘的另一面是加固与推翻,两者都有据可查:同一个偏好被反复印证,可信度上调、衰减放慢;用户说「我不再喝美式了,改喝拿铁」,旧记忆不被物理删除,而是标记为已失效并指向推翻它的那条新记忆,每一次改写都留下完整的证据链。这就是审计友好的遗忘:记忆可以被纠正,但纠正本身永远有据可查。
查询 Agent:取用不是一次向量搜索
存得好只是一半,取得准才是另一半。大多数记忆系统的「取用」是一次向量搜索:拿问题去库里比对,把最像的几条捞出来交差。GUMem 的取用是一个会规划的查询 Agent:它手里有一套记忆工具箱(最近消息、会话上下文、记录检索、向量检索、图关系),按问题的类型决定先查哪种记忆:问「怎么做」先查流程记忆,问「发生过什么」先查情节记忆,问关系走图,问偏好先看情感轨;查到的证据跨类型交叉核对,组装成这项任务刚好需要的上下文,不多也不少。
它还遵守一份写死在系统里的诚实协议:只有真正查到的记忆才算证据,禁止编造;证据不够时必须明说,而不是硬猜;证据冲突时按固定优先级消解(用户的显式更正压过推断出的偏好,本次会话压过旧记忆,具体约束压过宽泛喜好),解不开的就把两边都摆出来。还有一条硬边界:项目与用户的数据范围由系统强制,查询 Agent 只能传业务参数,连构造一条查询语句的权力都没有。
AutoDream:趁空闲把行为日志消化成洞察
Qoni 里有个趁 Agent 空闲时干活的机制,名字叫 AutoDream:自动把用户日积月累的行为流水,分析成真正的洞察,比如「这位用户总在深夜下单」「最近三周都在看去东京的机票」。它特别的地方在于:这不是工程师预先写好的统计报表,是 Agent 自己决定分析什么、自己写分析代码、自己跑出结论。人没想到要问的问题,它先想到了。
三条工程内幕,比「做梦」这个比喻更值得一讲:
- 模型从头到尾看不到原始数据。 行为日志先经过一层不涉及任何模型的确定性统计,压缩成一份「数据概要」;自然语言内容只保留长度等统计特征,原文从不外传;
- Agent 写的代码跑在沙箱里。 生成的统计代码禁止引入依赖、禁止访问存储与网络,只能调用一套受控的统计接口;
- 产出的洞察带证据。 每条洞察都是结构化的:指标、维度、当前值、基线、变化幅度、证据与可信度。
记忆主权:用户能看、能改、能删
记忆越强,另一个问题越重要:Agent 记了什么,用户自己知道吗?在 GUMem 里,用户侧的答案是三个「能」:
- 能看:每一条记忆都摆在明面上,带来源和可信度;
- 能改、能删:记错了就改,不想被记住就删,记忆属于用户,不属于模型;
- 能看着它长出来:写入一批消息后,可以实时订阅加工过程(事实提取好了、摘要生成了、主题更新了),事件流只回阶段和计数、不回内容原文,密钥类字符一律脱敏。
企业侧,记忆层是整个系统里个人信息浓度最高的地方,业务规则和合规要求挂在三个挂点上,权限刻意一级比一级小:写入之前可以改写数据(清洗敏感字段、补全业务信息),进模型之前只能追加指令(不能篡改记忆本身),模型之后只读通知(审计管道、同步 CRM)。真正的硬边界在系统强制的数据范围里:项目与用户的 scope 由后端划定,任何挂点、任何查询都越不过去。
记忆的取用还与 GenAuth 的身份边界打通:Agent 挂载记忆时要声明为哪个用户、允许哪些动作(比如只允许召回、不允许写入),拿到的不是「记忆库的钥匙」,而是「只能召回这一个用户」的钥匙;每一次取用与工具调用、网页操作走同一套身份、边界与审计。
接入 GUMem:对话三个调用,行为三个动作
SDK 包名 @qoniai/qoni(Node 18+,ESM / CommonJS 双构建,TypeScript 类型齐备),凭 Qoni Console 的 AK / SK 初始化。第一步同样是委托:GUMem 的 scope 拆得很细(gumem.memory:read / gumem.memory:write / gumem.message:write),会话、写入、召回这条常用链路有预置组合 QoniScopeBundles.GUMEM_SESSION_RECALL。
对话轨的存入与召回,三个调用:
import { Qoni } from "@qoniai/qoni";
const qoni = new Qoni({ accessKey, secretKey });
const { data } = await qoni.delegateToken({
user: { id: userId },
scopes: ["gumem.memory:read", "gumem.memory:write", "gumem.message:write"],
});
// Chat 存入:开一个会话,写入消息
await qoni.gumem.createSession({
token: data.token,
userId,
sessionId: "daily-assistant",
title: "日常助理记忆",
});
await qoni.gumem.addMessages({
token: data.token,
sessionId: "daily-assistant",
messages: [{ role: "user", content: "日报建议尽量简短,只给下一步行动。" }],
});
// 召回:不是关键词搜索,是面向当前任务的取用
const context = await qoni.gumem.recall({
token: data.token,
sessionId: "daily-assistant",
query: "这个用户对回复风格有什么已确认的偏好?",
details: true,
});
// context.data 就是可直接进提示词组装的就绪上下文
行为轨是对称的三个动作:qoni.gumem.actions.record(存入)、actions.recall(召回)、actions.stream(订阅),点击流、选择、放弃从这里落轨,物理上与对话轨分开。AutoDream 没有「接入代码」这个概念:行为一旦经 actions.record 落轨,空闲期的分析就自动运转,不需要多写一行。委托的收窄在这里直接生效:一个只拿到 gumem.memory:read 的 Agent,调用任何写入都会被拒。这不是文档约定,是凭证边界。
一张速查表,收拢全文讲过的能力:
| 能力 | 一句话说明 |
|---|---|
| 双轨记忆 | 对话轨 + 行为轨,从采集到加工物理分离,各有完整管线 |
| 六种记忆 | 工作 / 情节 / 语义 / 流程 / 情感,加上织网的关联记忆 |
| 提炼管线 | Collect → Extract → Condense → Recall → Forget 五站,每层产物带字段、带锚点 |
| 半衰期遗忘 | 五档半衰期衰减,反复印证加固,推翻留证据链 |
| 三存储分工 | 向量管相似、图管关系、事务库管账本,各自可独立重建 |
| 查询 Agent | 按问题类型规划取用路径,跨类型交叉核对证据 |
| 诚实协议 | 查到的才算证据,不够必须明说,冲突按固定优先级消解 |
| AutoDream | 空闲时自主分析行为流水:模型只看概要,代码跑沙箱,洞察带证据 |
| 记忆主权 | 能看、能改、能删,加工过程可订阅且不泄密 |
| 治理与委托 | 三挂点权限递减,scope 后端强制,取用凭 GenAuth 委托 |
| SDK 接入 | @qoniai/qoni,会话 / 写入 / 召回三调用,行为轨对称三动作 |
不是存下用户说过的每句话,而是理解用户是谁。
GUMem 目前处于 private build,将随 Qoni 分阶段开放,可以在官网加入 waitlist。

