Qoni JournalGenAuth

Give Every Agent Its Own Identity with GenAuth

用 GenAuth,给每个 Agent 一个独立身份

Authority delegated by a human: narrowing-only, short-lived, revocable at any time, audited end to end

权限来自人的显式委托:只收窄、短时效、随时可收回、全程可审计

Give every agent a passport of its own给每个 Agent 一本自己的护照

Authority delegated by a human: narrowing-only, short-lived, revocable at any time, audited end to end.

Agents are the new actors in the enterprise, and the identity stack has no seat for them

Identity technology has been at this for three-plus decades. Compress those decades into one sentence, and they keep answering a single question, over and over: this new kind of subject — why should anyone trust it? And every time that question gets a good answer, the world gains a layer of productivity.

In the LAN era, the question was "who is the person at this machine?" Directory services like AD (accessed over protocols like LDAP) registered every member of an organization, and Kerberos used tickets to prove who they were to each service on the intranet — and directories plus tickets carried the enterprise intranet. In the web era, the question became "why should this app act on the user's behalf?" SAML carried identity federation and single sign-on beyond the intranet; OAuth took on authorization and left authentication out of its scope, and OAuth 2.0 made it the common authorization framework of the entire web (an app can act for a person without surrendering their password) — together they carried SaaS and the open platforms. In the mobile-and-cloud era, the question was "one person, across countless apps and devices": OIDC added a standardized authentication layer on top of OAuth 2.0, "Sign in with ..." became the default front door of the internet — and it carried the mobile internet.

Three-plus decades, generation after generation of protocols — and all sharing one assumption buried so deep almost nobody calls it an assumption: the subject is always a human. The directory holds people, the tickets go to people, the grants come from people, the one signing in is a person; even the newest decentralized-identity (DID) experiments still, in mainstream practice, issue credentials to people.

Every good answer to the trust question adds a layer of productivity: three eras answered it for humans; the fourth era's answer is still blank
Three eras wrote their answer to "why should it be trusted" into a generation of protocols, each carrying a new layer of productivity — and when the agent's turn came, the fourth segment was left hanging on a dashed line.

Now, for the first time, that assumption is breaking. The AI agent is the new actor inside companies: it queries data, files requests, calls internal systems — work that used to belong exclusively to employees. It isn't a user and isn't an app — it's a third kind of digital subject. And it doesn't grow into the human shape: an agent decides its next step at runtime, based on the context and on what the last tool returned; before a task starts, nobody can fully write down its execution path. Statically pre-granting such a subject a fixed bundle of long-lived permissions will never be enough.

Academia has already named the gap. Researchers distill the attribution layer of agent infrastructure into three propositions (Infrastructure for AI Agents, 2025, arXiv:2501.10114): Agent IDs — who is this agent; identity binding — which person or organization stands behind it; certification — can its behavior and properties be trusted. The same paper runs through the alternatives already at hand: OAuth tokens and API keys can help with attribution, but they are typically each bound within a single service — they cannot link one agent's activity across services, let alone trace a sub-agent back to whoever created it; IP addresses are unstable and can be shared; even the newest agent-to-agent protocol (A2A) speaks mainly to what an agent can do — its AgentCard centers on capability discovery, while who stands behind it, and why it should be trusted, is left hanging. Meanwhile, the agent-authorization drafts under discussion at the IETF treat attenuation — authority that can only narrow as it travels down a chain — as a core requirement. The direction is clear: agents need an identity layer of their own.

The three propositions against the alternatives at hand: the gap is structural
The paper's verdict after examining the alternatives at hand: the three propositions are still unanswered.

The reality, though, is that most agents today act on borrowed identity: an employee's API key, a shared service account. The moment an agent heads for production, security asks three questions. When something goes wrong, can anyone tell whether a person did it or an agent? It only needed to do one thing — why does it hold a person's entire permission set? And can anyone answer "what allowed this call"? GenAuth is Qoni's identity layer, and its answer fits in one sentence: give every agent a passport of its own. The agent holds no permissions of its own; every capability it exercises comes from an explicit grant made by a specific person, revocable at any time.

A passport unpacks into three things: identity, attribution, delegation

Unpack the passport and it is three things:

  • Agent ID: an independent identity. Every agent gets dedicated credentials with a full lifecycle — create, rotate, disable, delete. Enabling or retiring an agent's identity depends on no employee account: people come and go; the agent's identity persists and retires on its own.
  • Identity binding: attribution to a person. Every agent is bound to the person and organization behind it, and every action traces back to whoever delegated it. When something goes wrong, it's no longer an unsolved case.
  • A delegation chain, built on the standards. No new protocol invented — delegation semantics for agents extended along the existing rails of OAuth 2.0 / OIDC: short-lived credentials whose scope only ever narrows, never widens. The existing identity infrastructure stays.

An agent's capabilities come half from its tools, so MCP tool calls fall under the same identity, policy, and audit governance: which tool, in whose name, within what scope — all on the record. And all of it is one SDK away.

The attribution layer's three propositions now have an engineering answer: Agent IDs and identity binding are answered by the passport itself, and certification gets its rails from verifiable grant boundaries and end-to-end audit. The model's first-principles constraint fits in one sentence: an agent owns an identity, but owns no inherent permissions — no root authority. From it follow four core principles: predictable, governable, revocable, traceable.

The GenAuth identity plane: from humans to resources, every hop has an identity, a boundary, and a record
Authority starts with a person, is issued by GenAuth, and is enforced at the gateway — every pass and every denial lands on the same audit timeline.

GenAuth's permission model: three sets, one intersection

An agent's effective permissions = what the human actually holds ∩ what was explicitly delegated ∩ what the enterprise has approved

Three sets, one intersection: authority can only narrow as it travels down the chain, never widen (authorization-system design calls this property attenuation). However capable the model, by construction it cannot out-rank the human who delegated it.

Every permission is scoped and time-boxed: it expires after an hour by default, can be as short as a minute, and evaporates when the task ends. An agent with no task in hand holds no standing business authority — security engineering calls this state zero standing privileges. The narrow-only design builds the decades-old principle of least privilege into the runtime: the closer a subject's permissions track what it actually needs right now, the more predictable its behavior — and the easier it is to spot a deviation and revoke the grant behind it.

The intersection formula: an agent's effective permissions are the intersection of three sets
The source of all permission is always a person: delegation scope and enterprise boundary intersect, and the agent holds only the overlap.

GenAuth's two tokens and Token Exchange (RFC 8693)

At the protocol layer, that formula becomes a split between authorization and action — two kinds of tokens:

  • The delegate token can exchange, but can't open doors. When a user approves a delegation, the agent receives a delegate token whose audience points permanently at GenAuth's exchange service (fixed to genauth:token-exchange), never at any business resource. To reach a resource, the agent must perform an RFC 8693 Token Exchange (grant type urn:ietf:params:oauth:grant-type:token-exchange) and trade it for a short-lived access token valid for exactly one resource — a standard JWT that mainstream gateways verify with their stock JWT / OIDC plugins. Every exchange is a live adjudication: is the grant still valid, is the scope in bounds, has the credential been disabled — all re-checked at that moment, instead of issuing one long-lived ticket and waving everything through.
  • Scope attenuation has two gates, plus an allowlist. The scopes requested at exchange must fall simultaneously inside what this delegation granted and what the target resource has registered — crossing either line means rejection — and the exchangeable range is governed by a server-side allowlist. Success and rejection are two distinct audit events (token_exchange.success / token_exchange.rejected): even the scopes that didn't make it are on the record.
  • act is minted only at exchange time. RFC 8693 uses the act (actor) claim to say "who is acting on whose behalf." The delegate token deliberately doesn't carry it: at authorization time, no action has happened yet. And act here is an extended structure with a type discriminator — the docs say plainly that a stock RFC 8693 library will silently fail to parse it, so the resource side must explicitly check the delegation type, and fail closed.
  • Time-to-live is a first-class citizen. Delegate tokens live from 60 seconds to 24 hours, one hour by default, with 5 to 15 minutes recommended for single-task delegations; the interactive-authorization window and the approval session each run 10 minutes. Expiry isn't an exception — it's the norm.
  • Audit is a dual-ID chain, minted into the token. Every delegation mints a grant_id (which authorization) and an audit_id (what happened under it), both embedded in the token itself and traveling wherever it goes: even an attacker holding a stolen token cannot book their actions under someone else's grant.
  • The default posture is fail-closed. If the policy service is unreachable, the request is denied. However hard the model hallucinates, it does not cross that line.

Consent inherits the OAuth authorization-code discipline: a one-time callback code, stored only as a hash, consumed inside a database transaction, with a state machine guaranteeing each code is used exactly once.

A typical day in production: an employee asks the internal data assistant, "pull last quarter's sales for the east region." The agent initiates a delegation under its own passport; the employee confirms; the agent exchanges it for an access token scoped to read-only access on the reporting system, valid for fifteen minutes; the gateway checks it and the query runs. Midway, the agent reaches for a customer-privacy field it wasn't granted? That request is rejected at the gateway — and logged. Afterwards, the security team sees one unbroken chain: who authorized, which agent, what it did, when, on what authority — including the overreach that never went through.

Delegate → exchange → access: authority narrows step by step, with an audit event at every step
The OBO (On-Behalf-Of) path: the delegate token can only be exchanged; the access token is short-lived and valid for one resource.

GenAuth's four-layer permission model: tighten a policy, no code change

Tokens answer "is this credential trustworthy?" The other half of the problem: should this particular request go through? GenAuth follows the classic separation in access control: the policy decision point (PDP) is centralized — rules are managed in one place, GenAuth; the policy enforcement point (PEP) is pushed down — traffic is stopped at the gateway edge or inside the business service.

On the decision side sits a four-layer model: roles (RBAC) answer who, policy conditions answer under what circumstances, resource namespaces answer which class of resource, data permissions answer down to which row. Policy conditions evaluate five operator families — boolean, date, IP, numeric, string — so constraints like "business hours only," "office network only," "within quota only" are expressed as configuration, not code. The operational payoff is immediate: tightening a policy means changing no business code and waiting for no release train — adjust it in the console, and newly issued tokens carry the new scope at once (with one exception: if enforcement lives at the gateway, the gateway plugin configuration still needs a separate update).

This split also answers a very practical objection — approval fatigue. Agents move far faster than people; if every step popped a confirmation dialog, "confirm" would decay into a mindless click within days. GenAuth divides the labor: what a human approves is one scoped, time-boxed delegation, not each individual action; every exchange and access inside it is adjudicated by machine against policy in real time, and anything out of bounds is simply rejected. Humans make few, weighty decisions; machines make many, fast ones.

PDP / PEP split: rules managed centrally, enforcement in place
Policy lives in GenAuth; an out-of-scope request is denied at the enforcement point, never touching business logic, with a record.

GenAuth's revocation and audit: effect boundaries and the five questions

"Revocable at any time" is a sentence every identity vendor writes. The engineering question that matters is the next one: revoked — effective when? GenAuth puts both boundaries in plain sight:

  • Paths validated through introspection or token exchange fail immediately after revocation, because every check goes back to the issuer and reads the delegation's live status;
  • Paths that only verify JWT signatures locally have a hard upper bound: the token's own short TTL. The token cannot outlive its next expiry.

This is also the deeper reason for short lifetimes: when a credential leaks, the blast radius is set by the token's lifespan. In the worst case, three lines of defense remain: dedicated credentials can be disabled and rotated instantly; the credentials themselves carry zero business permissions — actual access rides on short-lived, revocable delegations and access tokens; and enforcement-side checks keep running throughout.

The five audit questions — who, authorized whom, to do what, when, on what authority — draw their answers from two sources that vouch for each other. The authorization side is recorded by GenAuth: delegation approvals (with approver and approval time), exchange successes and rejections, issuance, revocation, policy changes. The execution side is recorded by the gateway's or the business system's access logs: which endpoint, when, with what result. The two sides are recorded independently and joined by correlation keys — request_id, jti, grant_id: an anomaly on one side leaves the other intact, and tampering with a single record means breaking into two systems at once. For compliance teams, this chain provides the audit-data foundation for behavioral traceability under regimes like PIPL and GDPR (the compliance assessment itself still belongs to the broader data-governance program).

A common question from leadership — "which people are using AI agents?" — gets two layers of answers here: the agent level shows cost and load (the gateway reports call volume and consumption per consumer credential, a ready-made dashboard); the person level shows ownership and boundaries (every delegation is, by construction, a structured record of "this person, at this time, put this agent to work within this scope"). The two layers share correlation identifiers and can be checked against each other.

Dual-source audit: authorization-side records and execution-side logs cross-verify through correlation keys
Two sources, recorded independently, joined by correlation keys — answering all five audit questions together.

Integrating GenAuth: two paths, two enforcement shapes

Integration follows the two shapes agents actually take:

  • User-level agents: CLI + Skill. The authorization flow lives outside the agent, in a toolkit: a Skill orchestrates the user's intent, while the CLI handles login, agent identity, user authorization, and short-lived credentials. The agent itself ships zero authorization code. Built for agents running on a user's machine or inside an agent workbench.
  • Enterprise agents: OpenAPI. Identity, user authorization, credential issuance and verification all happen through direct server-side calls to GenAuth's API. Built for existing applications, backends, and automation platforms.

Enforcement comes in two shapes, and they compose:

  • The gateway as the enforcement plane. Gateways like APISIX, Kong, or Higress validate tokens with their JWT / OIDC plugins, then call GenAuth via forward-auth for a live permission decision. An out-of-scope request dies at the gateway — it never touches business logic.
  • The business system as the enforcement plane. Services verify, introspect, and decide in-process via the SDK. Built for teams without a central gateway, or decisions that need to sit close to the data.

The existing identity foundation stays untouched: GenAuth federates with the current IdP over OIDC / SAML — employees keep authenticating against the existing system, and the agent identity and delegation layer grows on top. MCP tooling follows the same management/execution split: the management plane lives in GenAuth (instances, templates, OAuth2 configuration, user-to-tool connections, call logs — bulk-import or unbind one by one), while the execution plane stays with the gateway already in place.

One quick-reference table to gather everything this piece has covered:

Capability In one sentence
Independent agent identity Dedicated credentials with a full lifecycle: create / rotate / disable / delete
Dynamic delegation Explicit approval, layered TTLs, revocable anytime, traceable to the approver
Token Exchange / OBO RFC 8693 standard protocol; authority only narrows; both outcomes audited
Fine-grained authorization RBAC + conditional policies + resource namespaces + data permissions, PDP / PEP split
MCP governance Management plane for instances / templates / user connections; execution on the enterprise's gateway
API / SDK enforcement Services integrate directly: verify, introspect, decide live — no gateway required
The audit chain Dual sources (authorization events + execution logs), five questions answered, denials recorded
Enterprise identity foundation OIDC / OAuth 2.0 / SAML / LDAP, SSO, MFA, org structures — the full enterprise kit

The system is already written up in 37 pages of official documentation, in English and Chinese — concepts, architecture, integration guides, enterprise scenarios; the questions security reviews ask most often already have their own chapters. GenAuth is in private build and will roll out with Qoni in phases — the waitlist is open on the website.

Don't lend the agent the keys. Issue it a passport that can be revoked at any time.

权限来自人的显式委托:只收窄、短时效、随时可收回、全程可审计。

Agent 成了企业里的新角色,账号体系里却没有它的位置

身份技术走了三十多年。把这三十多年压缩成一句话,它反复回答的其实是同一个时代命题:这个新出现的主体,凭什么被信任? 而每一次把这个问题回答好,世界都会多出一层生产力。

局域网时代,问题是「这台机器前坐着的人是谁」:AD 这样的目录服务(配上 LDAP 这样的访问协议)把组织里的每个人登记进目录,Kerberos 用票据替他们向内网的各个服务证明身份,目录与票据支撑起了企业内网。Web 时代,问题变成「这个应用凭什么代表用户」:SAML 把身份联邦与单点登录带出内网;OAuth 专注解决「授权」、把「认证」留在范围之外,OAuth 2.0 让它成为整个 Web 的通用授权框架(应用替人做事,不必交出密码),它们撑起了 SaaS 与开放平台。移动与云的时代,问题是「同一个人,怎么横跨无数应用与设备」:OIDC 在 OAuth 2.0 之上补上标准化的认证层,「用某某账号登录」成了整个互联网的默认动线,撑起了移动互联网。

三十多年,几代协议,它们共享一个藏得很深的预设:主语始终是人。 目录、票据、授权、登录,主语全是人;连最新的去中心化身份(DID)试验,主流实践也还是在为人发证。

信任问题每被回答一次,生产力就多出一层:三个时代的答案主语都是人,第四个时代的答案还空着
三个时代把「凭什么被信任」的答案各写进一代协议、各撑起一层生产力;轮到 Agent 时,第四段还悬在虚线上。

现在,这个预设第一次被打破。AI Agent 成了企业里的新角色:替人查数据、发流程、调系统,做着过去只有员工才能做的事。它不是用户,也不是应用,是第三种数字主体;而且它长不成「人形」:Agent 在运行时才根据上下文和上一步工具的返回,决定下一步访问哪个系统、调用什么工具,任务开始之前,往往没有人能完整写出它的执行路径。给这样一个主体静态地预授一包长期权限,注定不够用。

这个空白,学术界已经点名。研究者把 Agent 基础设施的归因(Attribution)层归纳为三个命题(《Infrastructure for AI Agents》,2025,arXiv:2501.10114):Agent ID(这个 Agent 是谁)、Identity Binding(它背后站着哪个人或组织)、Certification(它的行为与属性是否可信)。同一篇论文逐一检验过现成替代品:OAuth 令牌与 API Key 通常各自绑定在单个服务里,串不起同一个 Agent 跨服务的行为,更追不到子 Agent 背后的创建者;IP 不稳定、可共享;A2A 的 AgentCard 侧重能力发现、回答「这个 Agent 能做什么」,「它背后是谁、凭什么被信任」依然悬着。IETF 讨论中的 Agent 授权草案,则把权限沿链条传递只能收窄(attenuation)列为核心属性。方向已经清楚:Agent 需要一个自己的身份层。

三个命题与现成替代品的逐一检验:空白是结构性的
论文逐一检验现成替代品后的结论:三个命题依然悬空。

而现实是,今天绝大多数 Agent 还在「借用」人的身份行动:拿员工的 API Key,用共享的服务账号。一旦要上生产,安全团队必然会问三个问题:出了事,说得清是人干的还是 Agent 干的吗?它只需要做一件事,为什么拿到了一个人的全部权限?「这次操作凭什么被允许」,答得上来吗?GenAuth 是 Qoni 的身份层,它的回答只有一句话:给每个 Agent 发一本自己的护照。Agent 本身没有任何权限,它做每件事的权限都来自某个人的明确授权,而且随时可以收回。

护照拆开是三件事:身份、归因、委托

护照拆开是三件事:

  • Agent ID,独立身份。 每个 Agent 拥有专属凭据,创建、轮换、禁用、删除全生命周期管理;启用与回收不依赖任何员工账号,人来人走,Agent 的身份独立存续、独立注销。
  • Identity Binding,归因到人。 每个 Agent 与背后的人和组织完成绑定,每一个行为都能归因到委托它的那个人;出了事,不再是无主悬案。
  • 委托授权链,长在标准上。 不发明新协议,在 OAuth 2.0 / OIDC 的既有轨道上延伸面向 Agent 的委托语义:凭证短时效,范围只收窄、不放大。现有的身份基础设施不用推倒重来。

Agent 的能力一半来自工具,所以 MCP 工具调用也被纳入同一套身份、策略与审计治理:Agent 用哪个工具、以谁的名义、在什么范围内,全部有据可查。所有这些,一套 SDK 接入。

归因层的三个命题由此有了工程答案:Agent ID 与 Identity Binding 由护照直接补齐;Certification 由可验证的授权边界与全程审计铺轨。整个模型的第一性约束只有一句:Agent 拥有独立身份,但不拥有固有权限(root authority)。 由此展开四个核心理念:可预期、可管控、可撤销、可溯源。

GenAuth 身份平面:从人到资源,每一跳都有身份、有边界、有记录
授权从人出发,由 GenAuth 签发、在网关处执行;每一次通过与拒绝,都落进同一条审计时间线。

GenAuth 的权限模型:三个集合求交集

Agent 实际权限 = 人的真实权限 ∩ 显式委托范围 ∩ 企业批准边界

三个集合求交集,意味着权限沿链条传递只能收窄、永远不能放大(授权系统设计里的 attenuation 属性):无论模型多强,Agent 天生不可能比委托它的那个人权限更大。

每一份权限都带范围和期限:默认一小时过期,最短一分钟,任务做完权限自动消失。没有任务在身的 Agent 不保有任何常驻业务权限,安全工程里叫零常驻权限(zero standing privileges)。这套「只收窄」的设计,就是把沉淀了几十年的最小权限原则做进了运行时:权限越贴近当下要做的事,行为越可预期,偏差也越容易定位、收回。

交集公式:Agent 的实际权限是三个集合的交集
权限的源头永远是人:委托范围、企业边界层层求交,Agent 拿到的只是交集。

GenAuth 的双令牌与 Token Exchange(RFC 8693)

这行公式在协议层的落实,是把「授权」和「行动」拆成两种令牌:

  • 委托令牌「只能换、不能开门」。 用户批准委托后,Agent 拿到的委托令牌,受众(audience)恒定指向 GenAuth 的兑换服务(固定为 genauth:token-exchange),不指向任何业务资源。要访问资源,必须先做一次 RFC 8693 Token Exchange(标准授权类型 urn:ietf:params:oauth:grant-type:token-exchange),换出一张只对单个资源有效的短时效访问令牌:标准 JWT,主流网关的 JWT / OIDC 插件按配置就能验签。每一次兑换都是一次实时裁决:授权是否仍然有效、范围是否越界、密钥是否已被停用,都在这一刻重新检查,而不是发一张长票放行到底。
  • scope 收窄有两道闸门,外加一张白名单。 兑换请求的范围,必须同时落在「这次委托授予的范围」与「目标资源注册的范围」之内,任何一边越界直接拒绝;可兑换范围还由服务端白名单管控。成功与被拒是两类独立的审计事件(token_exchange.success / token_exchange.rejected):连没通过的那部分 scope,也记录在案。
  • act 只在兑换那一刻铸入。 RFC 8693 用 act(actor)声明「谁在代表谁行动」。委托令牌里刻意没有它:授权发生时,行动还没有发生。而且这个 act 是带类型判别字段的扩展结构,文档写得很直白:用现成的 RFC 8693 库解析会「静默取不到值」,所以资源侧必须显式判断 Agent 委托类型,宁可 fail-closed。
  • 时效是一等公民。 委托令牌 60 秒到 24 小时、默认一小时,单任务委托建议 5 到 15 分钟;交互授权窗口与审批会话各 10 分钟。过期不是异常,是常态。
  • 审计双 ID 铸进令牌。 每次委托生成 grant_id(哪一次授权)与 audit_id(这次授权之后发生了什么),由 GenAuth 直接铸进令牌本身,随之流转:攻击者即使偷到令牌,也无法把行为记到别的授权名下。
  • 兜底姿态 fail-closed。 策略服务不可用时,请求直接被拒。模型再怎么幻觉,也越不过这道边界。

交互式授权的确认环节,沿用 OAuth 授权码的纪律:一次性回调 code,服务端只存哈希,消费发生在数据库事务里,状态机保证一张 code 只能用一次。

一个典型的落地场景:员工对内部数据助手说「帮我拉一下上季度华东区的销售数据」。Agent 用自己的护照发起委托,员工确认授权,Agent 换到一张只覆盖报表系统只读权限、十五分钟有效的访问令牌,经网关校验后完成查询。中途它想顺手访问客户隐私字段?那个请求在网关被拒,并且记录在案。事后安全团队看到的是一条完整链路:谁授权、授权给哪个 Agent、做了什么、什么时候、凭什么,连那次未遂的越权也在。

委托 → 兑换 → 访问三步:权限逐级收窄,每一步都有审计事件
OBO(On-Behalf-Of,代表用户)链路:委托令牌只能换,访问令牌只对单个资源短时有效。

GenAuth 的四层权限模型:收紧一条策略,不改代码

令牌回答「凭证是否可信」,另一半问题是:这一次请求,到底该不该放行? GenAuth 遵循访问控制的经典分层:策略决策点(PDP)集中,规则在 GenAuth 一处管理;策略执行点(PEP)下沉,拦截在网关边缘或业务服务内就地执行。

决策端是四层叠加的权限模型:角色(RBAC)回答「是谁」,策略条件回答「什么情况下」,资源命名空间回答「对哪类资源」,数据权限回答「到哪一行数据」。策略条件支持布尔、日期、IP、数值、字符串五类判断:「仅工作时段」「仅办公网段」「仅额度内」这类约束用配置表达,不用写代码。对运维团队的直接好处:收紧一条策略,不用改业务代码,不用等发版排期;控制台调整后,新签发的令牌即时携带新范围(执行点在网关侧时,插件配置的调整除外)。

这个分工也回答了「审批疲劳」这个很现实的反对意见:Agent 的动作比人快得多,逐次点确认撑不了几天。GenAuth 的分工是:人批准的是一张有范围、有期限的委托,而不是逐次批准每一个动作;委托之内的每一次兑换与访问,由机器按策略实时裁决,越界直接拒绝。人做的决定少而重,机器做的裁决多而快。

PDP / PEP 分离:规则集中管理,拦截就地执行
策略在 GenAuth 集中管理;越权请求在执行点被拒,不触达业务逻辑,并产生留痕。

GenAuth 的撤销与审计:生效边界与审计五问

「随时可以吊销」是所有身份产品都会写的一句话,工程上真正要紧的是下一句:吊销之后,多久生效? GenAuth 把两条边界写在明面上:

  • 经 introspect(凭证内省)或令牌兑换校验的链路:撤销后立即失败,因为每一次校验都会回到签发方核对委托的实时状态;
  • 仅做本地验签的 JWT 链路:失效上界是令牌自己的短时效 TTL,令牌活不过下一次过期。

这也是坚持短时效的深层原因:凭证泄露时的爆炸半径,由令牌寿命决定。 最坏情况发生时有三道防线:专属凭据即时禁用与轮换;凭据本身不携带业务权限,实际访问依赖的委托与访问令牌都是短时效且可撤销的;执行点校验持续生效。

审计五问(谁、授权给谁、做了什么、什么时候、凭什么)由两个源头互相作证:授权侧由 GenAuth 记录委托审批(含授权人与审批时间)、兑换的成功与拒绝、签发、撤销、策略变更;执行侧由网关或业务系统的访问日志记录调了什么接口、何时发生、返回什么结果。两侧独立记录,靠 request_id、jti、grant_id 这类关联键相互连接:单侧异常不影响另一侧的独立性,想篡改一条行为记录,需要同时改动两个系统。对合规团队,这条链为个人信息保护法、GDPR 等法规场景下的行为可追溯要求提供了审计数据基础(合规评估本身,仍需结合企业整体的数据治理开展)。

管理层常问的「哪些人在用 AI Agent」,在这里有两层答案:Agent 级看成本与负载,网关按 consumer 维度出调用量与消耗;人级看归属与边界,每一条委托天然就是一条「某人在某时把某个 Agent 用于某个范围」的结构化记录。两层共享同一组关联标识,可相互核验。

双源审计:授权侧记录与执行侧日志靠关联键互相作证
两个源头独立记录、关联键相互核验,共同回答审计五问。

接入 GenAuth:两条路径,两种执行形态

接入按 Agent 的形态分两条路:

  • 用户级 Agent,走 CLI + Skill。 授权流程整个「外挂」在工具包里:Skill 编排用户的意图,CLI 负责登录、Agent 身份、用户授权和短时效凭证;Agent 本体不用内建任何授权能力。适合跑在用户本地和 Agent 工作台里的 Agent。
  • 企业级 Agent,走 OpenAPI。 身份、用户授权、凭证的签发与校验,由服务端直接调用 GenAuth 的 API 完成。适合已有应用、服务端和自动化平台。

执行点按架构分两种形态,可以并存:

  • 网关作为执行平面。 APISIX、Kong、Higress 这类网关用 JWT / OIDC 插件完成令牌校验,再经 forward-auth 向 GenAuth 请求实时权限判断;越权请求在网关就被拦下,不触达业务逻辑。
  • 业务系统作为执行平面。 服务内直接用 SDK 验签、introspect、实时判断。适合没有统一网关、或需要贴着业务做数据级决策的场景。

企业已有的身份底座不用动:GenAuth 以 OIDC / SAML 联邦方式接入现有 IdP,员工身份仍由企业现有系统认证,Agent 的身份与委托层长在它之上。MCP 工具接入同样是管理与执行分离:管理面在 GenAuth(实例、模板、OAuth2 配置、用户与工具的连接关系、调用日志,支持批量导入、逐个解绑),执行面交给企业已有的网关承载工具流量。

一张速查表,收拢全文讲过的能力:

能力 一句话说明
Agent 独立身份 专属凭据,创建 / 轮换 / 禁用 / 删除全生命周期管理
动态权限委托 显式审批、TTL 分层、随时撤销、可追溯到授权人
Token Exchange / OBO RFC 8693 标准协议,权限逐级收窄,交换双向审计
细粒度鉴权 RBAC + 条件策略 + 资源命名空间 + 数据权限,PDP / PEP 分离
MCP 治理 实例 / 模板 / 用户连接的管理面,执行面对接企业已有的网关
API / SDK 集成鉴权 业务系统直接集成,验签、introspect、实时判断,不依赖网关
审计链 授权侧事件 + 执行侧日志双源,五问可答,拒绝也留痕
企业身份底座 OIDC / OAuth 2.0 / SAML / LDAP、SSO、MFA、组织架构等企业级能力

这套身份体系已经写成 37 页中英双语的官方文档:概念、架构、接入指南、企业场景全部就绪,安全团队评估时最常问的问题,文档里都有对应章节。GenAuth 目前处于 private build,将随 Qoni 分阶段开放,可以在官网加入 waitlist。

不是把人的钥匙借给 Agent,而是给它发一本随时可以吊销的护照。