Qoni JournalLaunch

Build a Production-Ready Agent with Qoni

用 Qoni,构建生产级的 Agent

Introducing Qoni · Give your agent a soul.

Qoni 正式亮相 · 给 Agent 一个灵魂。

Introducing Qoni — give your agent a soulQoni 登场 - Give your agent a soul

Introducing Qoni — managed infrastructure for AI agents in production.

A question worth answering properly

Models get smarter every generation, and more teams are building agents than ever. Getting a demo running takes an evening now; the hard part is what comes next. So this piece tries to answer one question properly:

What does it take to build a production-grade agent?

Writing the core business logic is only the start. To send an agent into production and have it work reliably, three more things are missing.

First: an identity of its own

Account systems were designed for people, and agents still have no seat in them. So most agents today borrow a human's identity to get things done: an employee's API key, a shared password. Nobody minds in a demo; in production it's three questions in a row: when something goes wrong, was it the person or the agent? It only needed to do one thing — why does it hold everything this user can do? And what, exactly, allowed this call — a question no log can answer.

And there's a deeper issue: the smarter and more capable the model, the larger the blast radius. The stronger the intelligence, the more it needs boundaries.

Give the agent a clear identity of its own — independent of any human — and the benefits are concrete: every action is attributable, and to whom; every permission is revocable (grants work like visas — scoped and time-boxed); every record is auditable, and reconciled with human identity in one system.

The agent stops being a shadow with borrowed papers, and becomes an employee with a badge.

Second: a productivity tool, and a runtime where it can act

An agent that can only talk isn't productive yet.

Human productivity lives in browsers, operating systems, applications. Agents are no different — an agent needs a browser and a runtime of its own, so it can plan and act autonomously: open pages, fill forms, verify results, retry on failure.

That's when an agent starts producing like a person: not just helping someone think, but doing it for them.

Third: memory, layered and long-term like a human's

A production-grade agent must also understand the people who use it.

Human memory is layered: things that just happened, patterns worked out over time, preferences that are hard to put into words. And people remember what someone did, not just what they said. During sleep, the day's experiences even get sorted — one wakes up with an extra layer of insight.

Agent memory should work the same way: processed in layers rather than logged as a transcript, covering behavior as well as conversation — and ideally able to dream: digesting experience in idle time into deeper understanding.

These three things are Qoni

Today, Qoni makes its official debut.

As managed infrastructure for AI agents in production, Qoni builds identity → action → memory once, integrated into one cloud-hosted runtime. Wiring up a workable identity, running a browser, storing some memory — none of it is hard. What's hard is making those three things secure, reliable, and proven in production. That second part is what Qoni does: a secure, governed, bounded environment that amplifies the model's intelligence safely, makes building agents faster, simpler, and cheaper, and lets agents stand on the web as first-class citizens — with their own identity, actions, and memory. It's a continuation of Tim Berners-Lee's line: this is for everyone.

The trinity — one runtime, the full architecture for production agents.
The trinity — one runtime, the full architecture for production agents.

01 · Identity — GenAuth

Thirty years of identity technology, five generations of protocols — from AD/LDAP and Kerberos to SAML, OAuth 2.0, and OIDC — have all answered the same question: why should this new actor be trusted? The subject was always a human.

Now it's the agent's turn. Researchers break the attribution layer of agent infrastructure into three questions: Agent IDs — who is this agent; identity binding — which person or organization stands behind it; certification — can its behavior and properties be trusted.

GenAuth turns those three questions into a product: every agent gets a passport — an independent Agent ID. Every action is attributable to the delegator. The delegated authorization chain runs on the existing rails of OAuth 2.0 / OIDC — short-lived credentials whose scope only narrows, never widens. MCP tool calls fall under the same governance. All of it, one SDK away.

Four core principles: predictable, governable, revocable, traceable.

Security engineering calls this narrow-only design the principle of least privilege: the closer permissions track what needs doing right now, the more predictable the behavior. Identity technology has traveled thirty years; this time, at last, the subject is the agent.

A passport that can be revoked at any time.
A passport that can be revoked at any time.

02 · Action — Web Agent

For humans and agents alike, almost 90% of productivity happens on the web, in a browser.

Give an agent a browser, and typical browser-use setups turn out slow, expensive, and not accurate enough. Fine in a demo, scary in production. Web Agent's answer is hosted execution: the agent's entire loop — observe, decide, act, retry — runs in the cloud, with no browser to maintain.

Every capability ships as an API:

  • DoAnything — one sentence, a full multi-step web flow: the page is re-observed after every step before the next move; success is never self-reported, only the visible end state counts. Every run is recorded and replayable.
  • WebSearch — not one search, but a round of verification: multiple engines in parallel, sources cross-checked — structured, citation-backed results returned.
  • DeepResearch — turns "looking things up" into "delivering a report." A brief and outline first, direction confirmed before large-scale collection; then cross-checking and synthesis into a document with citations and confidence labels.
  • Track — watches pages; "did it change" is never a feeling. Change is decided by deterministic rules; the moment it does, a webhook lands in the downstream system.

Add reusable login profiles on top: sign in once, and later tasks carry the state.

"Fast and accurate" is an engineering bar Qoni sets for itself: P90 task completion within two minutes, success rates pushed toward 99%, ground out by replaying real tasks round after round until it holds.

Two example scenarios: social-media publishing, comment monitoring, competitor tracking, public-data aggregation; and enterprise legacy systems with no API, brought back within reach through the browser.

When a step needs a human — a login, a QR code, a CAPTCHA — the agent hands over the browser: the operator finishes that step, and the agent resumes from the breakpoint, task unbroken, context intact.

Handed to a human when needed, and back to the agent when done.
Handed to a human when needed, and back to the agent when done.

03 · Memory — GUMem

The industry has reached consensus: an agent without memory meets its users for the first time, every time. But most "memory layers" only do half the job — extract the chat, store it, search it back later.

GUMem does the other half too:

  • Two tracks: conversation and behavior. Not just what users said, but what they did — what they clicked, chose, abandoned. Behavior is often more honest than words.
  • A multi-layer memory pipeline. Conversations → facts → summaries → topics, distilled layer by layer — the higher the layer, the closer to who this person is. Every memory carries its source, confidence, and time decay.
  • Recall that plans. A query agent breaks the question apart, gathers evidence across memory types, and assembles just enough context — and when the evidence runs out, it says so instead of guessing.
  • AutoDream insights. In idle time, GUMem digests accumulated behavior into deep, personal insights — "tends to order late at night," "has been watching Tokyo flights for three weeks." Not a pre-built report: the agent decides what to analyze, writes its own code, and runs it to a conclusion.
  • Formation in plain view. Memories don't grow in a black box: each one's distillation from conversation and behavior can be observed in real time — like watching it take shape.

And all of it belongs to the user: visible, editable, deletable. Memory belongs to the user, not the model. On LoCoMo, a long-conversation memory benchmark, GUMem scores 92.9% — the score is public on Qoni's site.

Memory, distilled layer by layer — the higher, the closer to who this person is.
Memory, distilled layer by layer — the higher, the closer to who this person is.

From one agent to a hundred

When every agent has an identity, a permission boundary, and a full audit trail, something changes: running one agent inspires confidence, and so does running dozens — even hundreds — each doing its own work. If anything goes wrong, it can be traced and stopped, instantly.

From one agent to a hundred — each with an identity, a boundary, and a trail.
From one agent to a hundred — each with an identity, a boundary, and a trail.

The bigger the fleet, the safer and more controllable it gets.

Safer. Stronger. Yours.

When an agent has a soul

In the coming years, the web will be home to agents in the hundreds of millions. They'll book the trips, watch the markets, run the workflows — a real share of the world's work, done on people's behalf.

That world can take two shapes. One is chaos: countless nameless automations charging around on borrowed identities, no one sure who is doing what. The other is order: every agent with a name, a boundary, and a place it came from — brilliant, and worthy of trust.

Qoni is building the second one.

"Give your agent a soul" isn't a romantic line. It's literal: an identity, so it can be recognized. Boundaries on its actions, so it can be trusted. A memory, so it can be entrusted. A soul is everything "trustworthy" means.

Tim Berners-Lee said: this is for everyone. Qoni wants to carry that sentence one step further —

The Agentic Web should be for everyone, too.

The Agentic Web should be for everyone, too.
The Agentic Web should be for everyone, too.

Now

Qoni is in private build.

Join the waitlist on the website — one email when it opens. Nothing else.

Give your agent a soul.

Introducing Qoni — managed infrastructure for AI agents in production. Qoni 正式亮相:面向生产级 AI 智能体的托管式基础设施。

一个值得认真回答的问题

过去一年,模型一代比一代聪明,做 Agent 的团队也越来越多。用一个晚上写出能跑的 demo 已经不难,难的是下一步。所以这篇文章想认真回答一个问题:

如何构造一个生产级的 Agent?

把核心业务逻辑写好,只是开始。要让一个 Agent 真正走进生产环境并可靠地工作,还缺三样东西。

第一样:一个明确的身份

账号体系是为"人"设计的,Agent 还不是其中一员。于是今天绝大多数 Agent 都在借用人的身份干活:拿员工的 API Key、共享的账号密码。Demo 里没人在意,上生产就是三连问:出了事,说不清是人干的还是 Agent 干的;它只需要做一件事,却拿到了一个人的全部权限;日志答不上"这次调用凭什么被允许"。

还有一件根本性的问题:模型越聪明,能力越强,失控的半径越大。智能越强,越需要边界。

给 Agent 一个明确的、独立于人的身份,好处是具体的:每一次行动都能归因,代表谁做的;每一份权限都能收回,授权像签证,有范围、有期限;每一条记录都可审计,和真实人类的身份连成同一体系。

Agent 于是从"借证件的影子",变成"有工牌的正式员工"。

第二样:一个生产力工具和一套能自主行动的运行环境

光会对话的 Agent,还算不上有生产力。

人类的生产力长在浏览器、操作系统、一个个应用上。Agent 也一样,需要一个属于它自己的浏览器和运行环境,让它能自主规划、自主行动:自己打开页面、自己填表、自己核对结果、失败了自己重试。

到这一步,Agent 才开始像人一样具备真正的生产力:不仅「帮人想」,并且「替人做」。

第三样:像人一样分层、长期的记忆

生产级的 Agent,还要更懂使用 TA 的用户。

人的记忆是分层的:有刚发生的事,有总结出的规律,有说不清的偏好。而且人不只记别人说过什么,更记对方做过什么。甚至,人还会在睡梦里整理白天的经历,第二天醒来,多了一层洞察。

Agent 的记忆也应该这样:多级加工而不是流水账,记对话也记行为,最好还能「做梦」:在空闲时自己消化攒下的行为,产生更深的理解。

这三样东西,就是 Qoni

今天,Qoni 正式亮相。

Qoni 作为一个面向生产级 AI 智能体的托管式基础设施平台,把 身份 → 行动 → 记忆 一次做好,整合成一个云端托管运行时。搭一套能用的身份、跑一个浏览器、存一点记忆,都不难;难的是让这三样东西安全、可靠、经得起生产检验。Qoni 做的,正是后者:一个安全、可控、有边界的运行环境,让模型的聪明才智被安全地放大,让 Agent 开发更快、更简单、更低成本,让 Agent 带着自己的身份、行动、记忆,堂堂正正地成为互联网的一等公民。这也是万维网发明人 Tim Berners-Lee 那句 "This is for everyone",在 Agent 时代的续写。

三位一体:一个运行时,生产级 Agent 的完整架构。
三位一体:一个运行时,生产级 Agent 的完整架构。

01 · 身份(GenAuth)

身份技术三十年,五代协议,从 AD/LDAP、Kerberos 到 SAML、OAuth 2.0 再到 OIDC,都在回答同一个命题:这个新出现的主体,凭什么被信任?主语,始终是人。

现在,轮到 Agent 了。学术界把 Agent 基础设施的「归因 Attribution」拆成三问:Agent ID - 这个 Agent 是谁;Identity Binding - TA 背后站着哪个人、哪个组织;Certification - 它的行为与属性是否可信。

GenAuth 把这三问做成了产品:每个 Agent 领一本护照,独立的 Agent ID;每个行为归因到委托者;委托授权链延伸在 OAuth 2.0 / OIDC 轨道上,凭证短时效、只收窄不放大;MCP 工具调用,纳入同一套治理。所有这些,一套 SDK 接入。

四个核心理念:可预期、可管控、可撤销、可溯源。

这套"只收窄、不放大"的设计,安全工程里叫最小权限原则:权限越贴近当下要做的事,行为就越可预期。身份技术走了三十年,这一次,主语终于是 Agent。

一本随时可以吊销的护照。
一本随时可以吊销的护照。

02 · 行动(Web Agent)

人和 Agent 的生产力场景,几乎 90% 都发生在 Web 和浏览器里。

让 Agent 操控浏览器,市面上的 browser-use 方案普遍慢、贵,准确率还不高,demo 能跑,生产不敢上。Web Agent 是云端托管的解法:观察、决策、操作、重试,整个执行循环都跑在云端,浏览器不用自己维护。

能力全部 API 化:

  • DoAnything:一句话下达,跑完整的多步网页流程:每步重新观察页面,再决定下一步;成功只看页面终态,不靠自我汇报;全程录像、可回放。
  • WebSearch:不是一次搜索,而是一轮核对:多引擎并行、交叉核对来源,返回带引用的结构化结果。
  • DeepResearch:把「查资料」升级成「出报告」:先出简报和大纲,确认方向后大规模收集;最后交叉核对成文,带引用与置信度标注。
  • Track:盯页面,「变没变」由确定性规则裁决;一有变化,webhook 直接推送到下游系统。

再配上可复用的登录态档案:登录一次,后续任务接着用。

"又快又准"是 Qoni 给自己定的工程标准:P90 的任务两分钟内完成、成功率朝 99%,用真实任务反复测试、回放,一轮轮磨到稳定。

两个场景举例:社交媒体的内容发布、评论监控、竞品追踪、公开数据聚合;企业内部没有 API、只剩网页的旧系统,也能用浏览器技术接回来。

遇到必须由人来完成的步骤(登录、扫码、验证码),Agent 把浏览器的控制权交回操作者:人做完,它从断点继续,任务不中断、上下文不丢失。

到了需要人的那一步,交给人;做完,再还给它。
到了需要人的那一步,交给人;做完,再还给它。

03 · 记忆(GUMem)

行业已经形成共识:没有记忆的 Agent,每次打交道都是初次见面。但大多数"记忆层"只做了一半:把聊天记录抽取出来、存起来、要用的时候再搜出来。

GUMem 把另一半也做了:

  • 对话和行为的双轨记忆:不只记用户说过什么,还记做过什么:点过什么、选过什么、放弃过什么。行为往往比语言更诚实。
  • 多层记忆加工管线:对话 → 事实 → 摘要 → 主题,逐层提炼,越往上越接近"这个人是谁";每条记忆带来源、可信度和时间衰减。
  • 会规划的取用:查询 Agent 拆解问题、跨记忆类型收集证据,组装成刚好够用的上下文;证据不够时明说,不硬猜。
  • AutoDream 洞察:空闲时自动消化行为流水,产出"总在深夜下单""最近三周都在看去东京的机票"这类个性化洞察:不是预设的报表,是 Agent 自己决定怎么分析、自己写程序、自己跑出结论。
  • 看得见的形成过程:每条记忆如何从对话与行为中被提炼出来,全程可以实时观察。

而这一切都属于用户:看得见、改得了、删得掉。记忆属于用户,不属于模型。在长对话记忆测试 LoCoMo 上,GUMem 拿到了 92.9% 的成绩(官网可查)。

记忆逐层提炼,越往上越接近「这个人是谁」。
记忆逐层提炼,越往上越接近「这个人是谁」。

从一个,到一百个

当每个 Agent 都有身份、有权限边界、每一步都有记录,事情开始起变化:跑一个,可以放心;同时跑几十上百个,各干各的活,也一样放心。出了问题,随时能查、随时能停。

从一个,到一百个:每个都有身份、边界与记录。
从一个,到一百个:每个都有身份、边界与记录。

规模越大,反而越安全、越可控。

Safer. Stronger. Yours.

当 Agent 有了灵魂

未来几年,网络上会生活着数以亿计的 Agent。它们替人订行程、盯市场、跑流程,世界的很大一部分事务,将由它们代为完成。

这样的世界有两种可能。一种是混乱的:无数无名的自动化程序,借着人的身份横冲直撞,没人知道谁在做什么。另一种是有序的:每个 Agent 都有名有姓、有边界、有来处,聪明,而且值得信任。

Qoni 选择建设第二种。

给 Agent 一个灵魂,不是一句浪漫的话。它是具体的:给它身份,让它被认识;给它行动的边界,让它被信任;给它记忆,让它被托付。灵魂,就是"值得信任"的全部含义。

Tim Berners-Lee 说,this is for everyone。Qoni 想做的,是把这句话再往前带一程:

Agentic Web,也应该属于每一个人。

Agentic Web,也应该属于每一个人。
Agentic Web,也应该属于每一个人。

现在

Qoni 正处在内测阶段。

在官网加入候补名单。开放的那一刻,会有一封信,只有那一封,没有别的。

给 Agent 一个灵魂。