Qoni JournalWeb Agent

Put Agents to Work on the Open Web with Qoni Web Agent

用 Web Agent,让 Agent 上网干活

Four APIs — DoAnything, Track, DeepResearch, WebSearch — hosted in the cloud, replayable end to end

四个 API:DoAnything、Track、DeepResearch、WebSearch;云端托管,全程可回放

Let your agent sense the real world让你的 Agent 感知真实世界

Four APIs — DoAnything, Track, DeepResearch, WebSearch — hosted in the cloud, replayable end to end.

What is Web Agent

About 90% of productive work happens in a browser today. Booking, expensing, reporting, publishing, price-checking — all inside one browser tab. Models are already smart enough — they write, calculate, reason — but the world is beyond the screen: pages change, prices move, things happen. For an agent to take on real work, what it needs isn't smarter conversation. It's the interface humans have used for thirty years: a browser.

Web Agent is Qoni's action layer: it hosts the entire execution loop in the cloud — open the page, observe, decide, act, retry on failure. On integration, it's one API: the agent states what it wants, and the rest happens in the cloud.

The Web Agent action plane: from trigger to output, the whole execution loop is hosted in the cloud
Triggers in, results out: scheduling, browsers, the execution loop, and the capability APIs, each in position in the cloud.

There's no shortage of browser-use style automation options, open source and paid alike. But a "library that can drive a browser" is only the starting point: headless browser clusters to maintain, proxies and login state to babysit, concurrency scheduling, crash recovery, and observation capture to assemble by hand — and scripts that shatter every time a page gets redesigned. The library answers "can it click"; the engineering has to answer "usable every day."

Web Agent's four core capabilities

"One API" means DoAnything, the general-purpose entry point; alongside it sit three specialists:

DoAnything — one sentence, and the agent runs a whole flow.

Say "open the back office, go store by store through 20 shops, filter this week's orders, and export them into one consolidated table."

The agent works through that chain in a hosted browser, re-observing the page after every action before deciding the next one. The failure it guards against is "looks successful, actually clicked elsewhere": success is never self-reported — it has to be backed by an end state visible on the page. Every run is recorded, every step replayable. When a step genuinely needs a human — a login, a CAPTCHA — it goes through the takeover mechanism below, not through brute force.

Track — watches pages; "did it change" is not decided by the model.

Say "watch round-trip flights to Tokyo and alert when the price drops more than 10%."

Once the goal and cadence are set, the verdict goes to deterministic rules: the same page is "a price cut" today and "a promotion launch" tomorrow, and if wording decides whether to alert, the watch becomes a random number generator. So the split is: the model only describes what it observed; "did it change, should an alert fire" is decided by rules. When the rules detect a change, a webhook lands at the configured endpoint.

DeepResearch — turns "looking things up" into "delivering a report."

Say "research the agent-identity-auth space and produce a report ready to drop into due-diligence material."

The pipeline is explicit: brief → plan → outline → gather → cross-check → synthesize. The outline stage can stop and wait for sign-off, and the pause sits deliberately before large-scale gathering: when the direction is wrong, diligence multiplies the waste. Otherwise the deliverable is just a page of links; what it hands over is a conclusion document with a citation list and confidence annotations.

WebSearch — one search, one round of verification.

Say "check today's pricing on these three competitors and produce a sourced comparison."

Multiple engines queried in parallel, URLs normalized and deduplicated, sources cross-checked — returning structured, citation-backed results, ready for downstream pipelines. What it really avoids is treating "most similar" as true: a search engine's first answer is not a fact.

Add reusable login profiles on top: sign in once, and later tasks carry the state. One profile is one complete browser identity (deliberately not split per site — single sign-on would trap the user otherwise), and its current validity is queryable.

Four APIs: DoAnything as the general entry, three specialists, plus a reusable login profile
Every capability ships as an API: general flows go through DoAnything; search, research, and standing watch each have a dedicated entry.

The agent does one action at a time, then looks again

The most insidious class of browser-automation accident: the page changes between two actions — a dialog covers the button, a list re-sorts — and an agent operating on old snapshots "looks successful while clicking somewhere else."

The execution loop is therefore broken into strict micro-steps: one action per round — observe → decide → act, then back to observe. The agent never chains operations on a stale page snapshot, and re-verifies the target element before executing.

The micro-step loop: observe, decide, act, then observe again
Never act on a stale snapshot: before and after every action, look at the world again.

The highest law of success judgment: if there is no evidence it succeeded, it failed. Green lights from intermediate layers don't count — a status flag flipped, a script ran to completion, the model said "done." Success must trace to the terminal state the user actually sees: the order confirmation page rendered, the report file landed in the workspace. Regression testing follows the same law: real user tasks are replayed over and over, and the assertions target the final answer, not "the flow completed."

Three observation tiers keep the agent fast and steady

Not every step dumps the whole DOM into the model. Observation has three tiers, light to heavy:

  • Accessibility tree (the default): for most steps, knowing which interactive elements exist is enough;
  • Targeted queries: only the fragment this step needs;
  • Full page snapshot: the last resort.

Raw page data lands in artifact storage; what enters the model context is only what this one step needs. The cleaner the context, the steadier the decisions; the lighter the observation, the faster the loop.

Paths the agent has walked harden into skills

An operation path repeatedly validated on a site hardens into a deterministic action sequence: on a hit, it executes directly, skipping the model call — faster, steadier, cheaper for high-frequency tasks.

Skills are not static scripts: when a redesign makes selectors drift, regeneration and revalidation trigger automatically, and the new deterministic sequence replaces the old — scripts still shatter, but they grow back on their own. One boundary: each site's skills are validated independently — never generalized to another site just because the two look alike. Two checkout flows can be identical except for the position of one button, and that button happens to be "confirm payment."

Dangerous actions pass gates; a human can take over anytime

Submitting forms, paying, deleting — these must pass typed gates: actions are explicitly classified, and high-risk categories trigger additional confirmation logic, not a "please be careful" line in the prompt. Recognizing login walls, error pages, and risk-control pages is a required structured output at every step: required means the model cannot silently skip it. On CAPTCHAs and identity checks, Qoni's position: no silent bypasses, ever — the correct move is to escalate to a human. CAPTCHAs exist precisely to tell humans from programs, and a production system should not sell bypassing them as a feature.

Some steps must be done by a person: logging in, scanning a QR code, solving a CAPTCHA. Web Agent's answer is Take Control: it hands the browser to the operator — a person takes over, finishes that one step, and after handback it resumes from the breakpoint with the context intact. One carefully designed detail: initiating a takeover does not pause the task; it merely issues a control link valid for a few minutes. Only when someone actually opens the link and takes the wheel does the task pause; if nobody comes, the link expires and the agent keeps working the problem itself. The link can be forwarded straight to the end user — SMS or QR code both work — and only one person can hold control at a time. On handback, the agent forcibly re-observes the page before continuing: it doesn't guess what happened during the takeover, it looks.

Web Agent uses the Browser Harness for production stability and reliability

Wiring a model to a browser takes an afternoon; running for hours without falling over is the actual engineering. Long-horizon tasks rarely die because the agent "couldn't click the button" — they die of context management, failure recovery, and state persistence. And an agent's execution path is decided by the model at runtime; nobody can enumerate it in advance. A system like that can only be constrained at runtime: every step observable and recorded, high-risk actions passing through gates.

So the shape of Web Agent isn't "a wrapper around a model." It's four layers, each with one job:

  • Orchestration: prompt and context assembly, task memory, the planning loop, behavioral guardrails — so the model always knows where it is, what's done, and what remains;
  • Browser: sandboxed cloud browsers, one dedicated instance per session; login state, fingerprints, and caches never bleed between sessions. A dedicated instance isn't a luxury — it's the baseline of isolation: two tasks sharing one browser is two people sharing one logged-in computer;
  • Scheduling: large-scale concurrent orchestration — queues, quotas, crash recovery — so a task can resume after an instance failure;
  • Output: structured artifacts with citations, confidence, and a replayable execution trail — objects downstream programs can consume and humans can audit step by step.
The four-layer harness: orchestration, browser, scheduling, output
The model is smart, but "hours without falling over" comes from the four layers around it.

A long task breaks into three objects. The Session holds the environment: one dedicated browser, one login state, one workspace. The Run holds one task: the complete life of one instruction. The event stream records everything: reasoning, actions, input requests, and artifacts pushed with sequence numbers, resumable after a disconnect — live progress reads it, after-the-fact replay reads it too.

All seven run states are explicit: pending, running, awaiting input, paused, done, failed, canceled. Awaiting input is first-class, not an anomaly: when a login, QR scan, or CAPTCHA needs a human, the task suspends in place and resumes from the breakpoint. The workspace passes files across runs, and login profiles carry identity across sessions: tasks are disposable; the environment accumulates.

Web Agent's target: P90 under two minutes, success rate toward 99%

The bar Qoni sets for Web Agent is a target, not an achievement already reached but one being ground toward: P90 task completion latency within two minutes, and the success rate pushed toward 99%. The method isn't mysterious: take the real tasks users have handed to Qoni over the past few years and test against them relentlessly; locate and fix every underperforming scenario, then replay to verify; distill improvements into general capability, not one-off patches.

Two workloads already running in production

Social media automation and data collection. An ops teammate tells the agent, "watch these three competitors' pricing and new-release pages, tell me when anything changes." Track takes over: patrol on schedule, rule-adjudicated verdicts, alerts pushed into the team channel. The same capability set runs multi-platform publishing, comment-section monitoring, and public data aggregation.

The systems inside enterprises that "can't be integrated." Real enterprises are littered with legacy systems that have no API but do have a web page. A finance teammate exports a report from a legacy web UI every week; Web Agent takes over the flow — log in, navigate, export, drop the file into the workspace for the downstream task. Login expired? The task moves into awaiting input, the control link goes to the teammate, and after re-login the task resumes from the breakpoint. Unattended operation doesn't come from "nothing ever goes wrong." It comes from having a viable path after something does.

Integrating Web Agent: scope permissions first, then tell "run finished" from "goal achieved"

The SDK ships as @qoniai/qoni (Node 18+, ESM / CommonJS dual build, full TypeScript types), initialized with the AK / SK issued in Qoni Console. Step one isn't sending a task — it's getting a delegation: issue a scoped delegation credential for the current user, and the agent acts only inside its boundary.

import { Qoni } from "@qoniai/qoni";

const qoni = new Qoni({ accessKey, secretKey });

// Authorization first: the products sugar expands each product into its read + manage scope pair
const { data } = await qoni.delegateToken({
  user: { id: userId },
  products: ["doAnything", "track"],
});

// One instruction, one run
const run = await qoni.doAnything.run({
  token: data.token,
  prompt: "Open the back office, export last week's reconciliation report, drop it into the workspace",
});

const result = await run.wait();
result.status;            // succeeded / failed / canceled: the delivery status of this run
result.isTaskSuccessful;  // whether the task's goal was actually achieved: a separate field

The last two lines deserve a pause: delivery status and goal achievement are two fields — a clean finish can still miss the goal. A run can end normally (succeeded) without achieving its goal; the protocol answers the two questions separately — "no evidence means failure," as seen from inside the SDK.

When a human is needed, the interaction isn't an exception — it's a typed event. All five interaction types (site_login / clarification / confirmation / take_control / wait) arrive on the event stream, wrapped in a handle with methods:

for await (const event of run.events()) {
  if (event.type !== "interaction") continue;
  const i = run.interactionHandle(event.data);
  if (i.type === "clarification" && i.can("answer")) await i.answer("the blue button");
  else if (i.can("confirm")) await i.confirm();
}

i.can() carries a deliberate design: whether each action is available is declared explicitly by the backend, and calling an undeclared method throws — the SDK won't paint a button that does nothing. Standing watch follows the same shape: qoni.track.create() returns a monitor handle; a new monitor first enters intent alignment (pending_clarification), and once active, runNow() triggers an immediate patrol, refine() adjusts the schedule and notification channel, and pause() / resume() / delete() manage its lifecycle. After a disconnect, attach() by id returns to the same monitor or run, resuming the event stream from the breakpoint.

One quick-reference table to gather everything this piece covered:

Capability In one line
DoAnything, general web operations Multi-step flows from one sentence: log in, fill forms, compare — executed via the micro-step loop
Track Patrol on schedule; deterministic rules decide "changed or not"; webhooks to the configured endpoint
DeepResearch Six-stage explicit pipeline; outputs a conclusion document with citations and confidence
WebSearch Parallel multi-engine retrieval, normalized-URL dedup, cross-checking
Login profiles Sign in once, reuse across sessions; one profile is one full browser identity, validity is queryable
Workspaces Files carried across runs: one task's output, the next task's input
Take Control Links valid for minutes; pause only on takeover; forced re-observation on handback
Event stream Structured record of every step, sequence-numbered, resumable, replayable
Skills Validated paths harden into deterministic sequences; never generalized across sites
Risk gates Submit / pay / delete pass typed gates; CAPTCHAs escalate to humans
SDK @qoniai/qoni; delegation first, five typed interactions, reattach across disconnects

Next, two things — no dates promised, both under way: prove on public industry benchmarks and publish the real-task test set accumulated over the years; and let the agent run beyond the cloud — someday, on local browsers and computers.

Not another agent that demos well — "fast and right" delivered as an engineering metric.

Web Agent is currently in private build and will open in stages with Qoni. Join the waitlist on the website.

四个 API:DoAnything、Track、DeepResearch、WebSearch;云端托管,全程可回放。

Web Agent 是什么

当前约 90% 的生产力场景都发生在浏览器里。订票、报销、查数、发布、比价,都在一个浏览器标签页里完成。模型已经会写、会算、会推理,但世界在屏幕之外,页面在变、价格在变、事情在发生。让 Agent 接住真实工作,它需要的不是更聪明的对话,而是人类用了三十年的那个界面:一个浏览器。

Web Agent 是 Qoni 的行动层,把整个执行循环(打开页面、观察、决策、操作、失败重试)托管在云端。接入时,它就是一个 API:Agent 说清楚要什么,剩下的事发生在云上。

Web Agent 行动平面:从触发到产出,整个执行循环托管在云端
触发进来,结果出去:调度、浏览器、执行循环和能力 API,都在云端各就各位。

市面上并不缺 browser-use 类的浏览器自动化方案。但拿到一个「能操作浏览器的库」只是起点:无头浏览器集群要养,代理和登录态要维护,并发调度、崩溃恢复、观测采集都得自己搭;页面一改版,脚本碎一地。库解决的是「能不能点」,工程要解决的是「天天可用」。

Web Agent 的四个核心能力

「一个 API」指的是 DoAnything 这个通用入口;在它旁边,还有三个专职 API:

DoAnything:一句话,跑完整的多步流程。

「打开后台,逐个进入 20 个店铺,把本周订单筛选出来,导出并汇总成一张表」

Agent 在云端浏览器里自己完成这一串动作:每做一个动作,重新看一眼页面,再决定下一步。成功不看模型自己汇报,只看页面上真实出现的终态。它防住的是「看起来成功了,其实点了别处」。全程录像,每一步都能回放。遇到登录、验证码这类必须由人来完成的步骤,走后面的人工接管,而不是硬闯。

Track:盯页面,变没变不由模型说了算。

「盯住东京往返的机票价格,降价超过 10% 就推送提醒」

同一个页面,模型今天说「价格下调」,明天说「促销上线」。靠措辞决定要不要提醒,盯守就成了随机数发生器。所以判定交给确定性规则:模型只负责描述观测,「变没变、该不该提醒」由规则裁决;判定为变化,webhook 直接推送到配置的端点。

DeepResearch:把「查资料」变成「交报告」。

「研究一下 Agent 身份认证这个方向,出一份能直接放进尽调材料的报告」

流程是显式的:简报 → 计划 → 大纲 → 收集 → 交叉核对 → 综合。大纲这一步可以停下来等待确认,暂停点刻意设在大规模收集之前:方向错了,收集得越勤奋,浪费得越彻底。最终交出的,是带引用清单与置信度标注的结论文档,而不是一页链接。

WebSearch:搜一次,核对一轮。

「查一下这三家竞品今天的定价,出一份带来源的对照」

多引擎并行检索、规范化 URL 去重、来源交叉核对之后,返回带引用的结构化结果,下游流程直接消费。要避免的是「把最像的当真的」:搜索引擎的第一条答案不等于事实。

再配上登录态档案:登录一次,后续任务接着用。一份档案是一份完整的浏览器身份(刻意不按站点拆分),档案是否有效,有状态可查。

四个 API:DoAnything 通用入口,三个专职能力,加一份可复用的登录态档案
能力全部 API 化:通用流程走 DoAnything,检索、研究、盯守各有专职入口。

Agent 每步只做一个动作,做完再看一眼

浏览器自动化里最阴险的一类事故:页面在两个动作之间变了(弹窗盖住按钮、列表重新排序),基于旧快照连续操作的 Agent 会「看起来成功了、其实点了别处」。

所以执行循环拆成严格的微步:每轮只执行一个动作,观察 → 决策 → 行动,然后回到观察。动作执行前,还会再验证一次目标元素是否仍然存在。

微步循环:观察、决策、行动,然后回到观察
永远不基于过期的页面快照行动:每个动作前后,都要重新看一眼世界。

判定成功的铁律是:没有证据证明它成功,它就是失败。 中间某一层的绿灯(任务状态翻转了、脚本跑完了、模型说「我做完了」)都不算数,成功必须追到用户真正看到的终态:订单确认页出现了、报表文件落在了工作目录里。回归测试同样如此:拿真实用户的任务反复回放,断言的是最终答案,不是「流程走完了」。

三档观测,让 Agent 又快又稳

不是每一步都把整页 DOM 塞给模型。观察分三档,从轻到重:

  • 无障碍树(默认档):多数步骤知道「有哪些可交互元素」就够了;
  • 定向查询:只取这一步需要的局部;
  • 全量页面快照:兜底档。

原始页面数据一律落进档案存储,进模型上下文的只有这一步刚好需要的观测。上下文越干净,决策越稳;观测越轻,循环越快。

Agent 跑过的路,沉淀成技能

在某个站点反复验证过的操作路径,会固化成确定性的动作序列:命中即执行,跳过模型调用。对高频重复的任务,这意味着更快、更稳、更便宜。

技能不是静态脚本:页面改版导致选择器漂移时,自动触发重新生成、重新验证,新的确定性序列替换旧的;脚本仍然会碎,但碎了会自己长回来。边界是:每个站点的技能独立验证,绝不因为两个网站长得像就把技能推广过去。两个电商站的下单流程可以像到只差一个按钮的位置,而那个按钮恰好是「确认付款」。

Agent 的危险动作要过关卡,人随时能接手

提交表单、付款、删除这类动作必须经过类型化的关卡(Gate):动作被显式分类,高风险类别触发额外的确认逻辑,而不是靠提示词里的一句「请谨慎操作」。登录墙、错误页、风控页的识别是每一步必填的结构化输出:必填意味着模型无法沉默地跳过。遇到验证码和身份核验,Qoni 的立场是:不做也不该做静默绕过,正确的做法是升级给人。验证码存在的意义就是区分人和程序,一个生产系统不应该以绕过它为能力卖点。

总有些步骤必须人来做:登录、扫码、验证码。Web Agent 的做法是 Take Control:把浏览器的控制权交给操作者,人接手做完,交还后它从断点接着往下干,上下文不丢失。发起接管时任务并不暂停,只是签发一个几分钟内有效的接管链接;真的有人打开链接接手,任务才暂停;没人来,链接过期作废,Agent 继续自己想办法。链接可以直接转发给终端用户(短信、二维码都行),同一时刻只允许一个人控制;人做完交还,Agent 强制重新观察一次页面再继续:人接管期间发生了什么,它不猜,重新看。

Web Agent 通过 Browser Harness 保证生产级稳定、可靠

把模型接上浏览器,一个下午就能跑通;跑几个小时不翻车,才是工程。长程任务挂掉,很少死在「不会点按钮」上,而是死于上下文管理、失败恢复、状态保持。而且 Agent 的执行路径是运行时才由模型决定的,没有人能在任务开始前穷举它会做什么;约束一个无法预先穷举的系统,靠的是运行时:让每一步可观察、可记录,让高风险动作过关卡。

所以 Web Agent 的形状不是「模型外面包一层」,而是四层各司其职:

  • 编排层:提示词与上下文的组织、任务记忆、规划循环、行为护栏,让模型每一步都知道「在哪、做到哪、还差什么」;
  • 浏览器层:沙箱化的云端浏览器,每个会话独占实例,登录态、指纹、缓存互不串扰;
  • 调度层:大规模并发编排,排队、配额、崩溃回收,实例故障后任务可恢复续跑;
  • 产物层:结构化输出,带引用、带置信度、带可回放的执行轨迹,能被下游程序直接消费、能被人逐步复核。
Harness 四层:编排、浏览器、调度、产物,各司其职
模型很聪明,但「跑几小时不翻车」靠的是模型外面这四层。

长任务在内部被拆成三个对象。**会话(Session)**装环境:一个独占浏览器、一份登录态、一个工作目录。**执行(Run)**装一次任务:一条指令从开始到结束的完整生命周期。事件流记全过程:推理、动作、请求输入、产物按序编号推送,断线续传;实时进度看它,事后回放也看它。

执行的七个状态全部显式:排队中、运行中、等人输入、已暂停、完成、失败、已取消。其中「等人输入」不是异常,是一等状态:任务跑到登录、扫码、验证码,原地挂起,人来接手,断点续跑。工作目录跨执行传递文件,登录态档案跨会话复用身份:任务是一次性的,环境是可积累的。

Web Agent 的目标线:P90 两分钟、成功率朝 99%

Qoni 给 Web Agent 定的是一条目标线,不是已经达成的成绩,而是正在逼近的目标:任务完成时延的 P90 控制在两分钟以内,成功率朝 99% 去磨。 方法不神秘:拿过去几年真实用户交给 Qoni 的任务反复测试,跑不好的场景逐个定位、修好、再回放验证;有效的改进沉淀成通用能力,而不是某个案例的特殊补丁。

已经在生产里跑的两类活

社交媒体自动化与数据采集。 运营同学对 Agent 说「盯住这三家竞品的价格页和新品页,有变化告诉我」,Track 接手:按节奏巡查、规则裁决、提醒推到工作群;同一套能力也在跑多平台内容发布、评论区监控和公开数据聚合。

企业内部那些「接不进来的系统」。 真实的企业里散落着大量没有 API、但有网页的旧系统。财务同事每周要从老系统网页里导出报表,Web Agent 接管这条流程:登录、导航、导出、把文件落到工作目录交给下游任务。登录过期了?任务挂起进「等人输入」,接管链接发给同事,重新登录之后从断点继续。无人值守靠的不是「永远不出事」,而是出事之后有一条走得通的路。

接入 Web Agent:先圈权限,再分「跑完了」和「做成了」

SDK 包名 @qoniai/qoni(Node 18+,ESM / CommonJS 双构建,TypeScript 类型齐备),凭 Qoni Console 签发的 AK / SK 初始化。接入的第一步不是发任务,而是拿委托:为当前用户签发一张范围明确的委托凭证,Agent 只在这张凭证的边界内行动。

import { Qoni } from "@qoniai/qoni";

const qoni = new Qoni({ accessKey, secretKey });

// 权限先行:products 语法糖会为每个产品展开 read + manage 一对 scope
const { data } = await qoni.delegateToken({
  user: { id: userId },
  products: ["doAnything", "track"],
});

// 一条指令,一次执行
const run = await qoni.doAnything.run({
  token: data.token,
  prompt: "打开后台,导出上周的对账单,放到工作目录",
});

const result = await run.wait();
result.status;            // succeeded / failed / canceled:这次执行的交付状态
result.isTaskSuccessful;  // 任务目标是否真的达成:另一个字段,分开回答

最后两行值得停一秒:交付状态和目标达成是两个字段:跑完,不等于办成。 一次执行可以正常结束(succeeded)但没有达成任务目标;协议把这两个问题分开回答,正是「无证据即失败」在 SDK 里的样子。

需要人的时候,交互不是异常,是一种带类型的事件。五种交互(site_login / clarification / confirmation / take_control / wait)都从事件流里来,包成带方法的句柄:

for await (const event of run.events()) {
  if (event.type !== "interaction") continue;
  const i = run.interactionHandle(event.data);
  if (i.type === "clarification" && i.can("answer")) await i.answer("蓝色那个按钮");
  else if (i.can("confirm")) await i.confirm();
}

i.can() 有个讲究:每个动作是否可用由后端在交互里显式声明,调用未声明的方法直接抛错,SDK 不会画一个点了没用的按钮。盯守走同一套形状:qoni.track.create() 返回监控句柄,新监控先进入意图对齐(pending_clarification),转为 active 后可以 runNow() 立即巡查、refine() 调整节奏与通知通道、pause() / resume() / delete() 管理生命周期;断线之后凭 id attach() 回同一个监控或执行,从事件流断点续传。

一张速查表,收拢全文讲过的能力:

能力 一句话说明
DoAnything 通用网页操作 一句话下达多步流程:登录、填表、比价,微步循环逐步执行
Track 持续盯守 按节奏巡查,变没变由确定性规则裁决,webhook 投递到配置的端点
DeepResearch 深度研究 六阶段显式流水线,产出带引用清单与置信度标注的结论文档
WebSearch 搜索 多引擎并行检索、规范化 URL 去重、交叉核对
登录态档案 登录一次跨会话复用,一份档案是一份完整浏览器身份,是否仍有效可查询
工作目录 跨执行传递文件,上一个任务的产出,下一个接着用
Take Control 接管链接分钟级有效,有人接手才暂停,交还后强制重新观察
事件流 每步结构化记录、按序编号推送,断线续传,支持逐步回放
技能沉淀 验证过的路径固化为确定性动作序列,站点之间不迁移
风险关卡 提交 / 付款 / 删除过类型化 Gate,验证码升级给人
SDK 接入 @qoniai/qoni,委托先行,五种交互类型协议化,断线可 attach 续传

下一步两件事,不承诺时间,但都在路上:去公开的行业测试榜上证明自己,并把这些年积累的真实任务整理成对外发布的测试集;让 Agent 不只跑在云端,将来也能跑在本地浏览器和电脑上。

不是又一个会演示的 Agent,而是把「又快又准」当成工程指标来交付。

Web Agent 目前处于 private build,将随 Qoni 分阶段开放,可以在官网加入 waitlist。