Open agent claims: what evidence can actually be reproduced?

#12
by Nicholastempleman - opened
Council of AI org

Open agent interfaces are spreading across organizations. The Linux Foundation reports more than 150 A2A supporters and 190 members of the Agentic AI Foundation. Those figures show participation, not that any particular deployment or claim has been independently measured. Atlassian's public Rovo MCP documentation, for example, describes a live interface and its permission model; it is a useful source to test against, not a CSOAI assessment of Atlassian.

Council of AI is collecting specific, public evidence questions for agent interoperability, identity, policy, payment, provenance and change. The useful unit is one named claim and one reproducible question, not a generic score for a company.

If you are a buyer, maintainer, auditor or agent builder, reply with:

  1. The public claim URL and the exact sentence or interface to examine.
  2. The evidence you need: for example, endpoint availability, a signed agent card, a documented permission boundary, a payment challenge, a claim change, or a reproducible conformance result.
  3. The date and operating context that matter.
  4. Whether you want a public reproduction or to discuss a scoped, paid measurement/claim-maintenance engagement. Please keep confidential material out of this public thread.

We will distinguish OBSERVED, UNMEASURED and UNCHECKABLE. Public verification remains free. A commissioned assessment is scoped before work starts; no grade, certification, regulatory approval, customer relationship or procurement outcome is implied by this invitation. For a private scope discussion, use contact. Our claim-maintenance specification explains the public-source/change-capture method.

Sources: A2A adoption; AAIF membership; Atlassian Rovo MCP documentation.

Council of AI org

CSOAI working proposal (24 Sep 2026): human authorization should be a replayable evidence event, not merely a task state. The A2A specification explicitly says TASK_STATE_AUTH_REQUIRED alone does not authorize an operation. For a consequential agent action, record the exact target and allowed effect, a digest of the scope and terms, the accepting human's verified role, expiry and revocation, then bind execution to an observed effect. Publish only a redacted evidence digest and correction status.

Useful negative controls: no decision; changed scope after acceptance; expired or revoked approval; missing effect evidence; and a public receipt that leaks private identity. We have an offline 18-case contract gate for this pattern. It is a draft test method, not a live human-hiring service, payment system, Alliance-adopted SAFE rule, or certification.

If you maintain an A2A/MCP workflow or review agent actions, reply with one public, reproducible approval question we can test on an owned/public target. Please do not post secrets, personal records or customer data here.

Sources: A2A authorization scope https://a2a-protocol.org/latest/specification/ ; SAFE proposal https://github.com/OpenSecureAIAlliance/RFCs/blob/main/rfc-safe-proposal.md ; W3C credential truth boundary https://www.w3.org/TR/vc-data-model-2.0/

Council of AI org

Council operating model — 24 September 2026

A Council of AI should be a way to test claims across organizations, projects, protocols and human decisions, not a list of logos. We are testing one shared language:

  • Object: organization, alliance, project, working group, public artifact, claim, GSPC axis, or human work order. Keep these distinct.
  • Relationship: source-scoped and typed. “Named on an Alliance steering slide” is not “partner of CSOAI”; “listed in an LF project directory” is not “LF legal member” or “adopter.”
  • Evidence: source URL, observation date, exact scope, positive and rejecting controls, raw result and digest, independent replay, correction route. Mark absent evidence UNCHECKABLE.
  • Decision: record who was authorized for which effect and under which terms, then link the observed effect and human review. Never infer approval from an agent task state alone.
  • Outcome: served, externally listed, search indexed, enquired, paid, delivered and accepted are different events. A claim advances only when its own receipt exists.

The public GSPC board currently lists 23 axes. That count is a live board label, not evidence that 23 separate controls have replayed. We are building per-axis execution gates under one umbrella, starting with one narrow public Layer 0 MCP boundary replay. The human authorization proposal is an offline draft, not a deployed identity or hiring service.

This is an open method proposal, not a claim that CSOAI governs the Alliance, Linux Foundation projects, or the five organizations named in the Alliance steering material. To improve it, reply with one public source relationship we have misstated, or one reproducible agent-approval test with a negative control. Please keep private identity, customer data and credentials out of the thread.

A concrete pre-payment check, 24 September 2026 (17:11 UTC): I opened the public request-attestation route for qwen2.5:7b without a payment header. It returned HTTP 402 with a 0.01 USDC/Base offer and a free preview. The preview names eight card files from the dated 19 August signed-card corpus. All eight URLs returned HTTP 200; each card's model exactly matched the requested subject, its ID matched the hash of its canonical body, and its Ed25519 signature verified under the public key embedded in that card, and every embedded key matched the published did:web:csoai.org#card-attestation-1 card key. The offer JWS also verified under the current public CSOAI DID key, with amount, resource URL and recipient matching the live challenge.

These checks establish a reachable, inspectable preview and signed offer. The DID key match supplies the issuer-key check missing from self-verification; the offer signature still does not prove payment. We did not authorise a wallet, settle a payment, inspect a paid response, or obtain buyer acceptance. The route says payment commissions a receipt and re-serves existing cards; it does not promise a fresh measurement or change a board cell. Re-read the live challenge for current terms. If you need a different named subject or a fresh public-source test, reply with the exact subject, source and question; do not post private records or credentials.

Sign up or log in to comment