Methodology · v0.1.0
What we measure, how we weight it, and what we refuse to claim.
Commerce4 measures the technical and semantic readiness of a commerce catalog for AI shopping agents. Every point in the score maps to a condition that can be observed on your storefront and verified independently.
Category weights
Weights reflect where agent-driven revenue is most often lost, in our observation of Shopify catalogs. Weights are versioned; changing them produces a new methodology version so historical scores stay comparable.
| Category | What it measures | Weight |
|---|---|---|
| Catalog & Product Semantics | Whether product titles, attributes and taxonomy describe the product the way a buyer asks for it. | 25 |
| Structured Data & Machine Readability | Validity and completeness of schema.org markup, feeds and machine-readable surfaces. | 15 |
| AI Discovery / Retrieval Readiness | Whether AI crawlers can reach, render and retrieve catalog and content surfaces. | 15 |
| Merchant Policies & Trust Data | Whether shipping, returns, warranty and support terms are machine-answerable. | 10 |
| Inventory, Price & Variant Clarity | Whether price, stock and variant availability are accurate, current and variant-level. | 10 |
| Cart / Checkout Agent Readiness | Whether an agent can build a cart and hand off a completable checkout without human-only steps. | 15 |
| Measurement & AI Traffic Attribution | Whether agent-originated sessions and revenue can be identified and reported. | 10 |
| Total | 100 | |
Readiness tiers
Critical
0–39Agents cannot reliably identify, price or purchase your products. Fundamental catalog or access issues are present.
Emerging
40–59Products are partially retrievable but qualifier-based queries and checkout handoff fail frequently.
Competitive
60–79Most intents resolve. Remaining losses concentrate in edge configurations, bundles and attribution.
Agent-Ready
80–100Catalog, policies and checkout are machine-legible end to end, with monitoring in place to catch regressions.
Category status
- Strong≥ 80% of the category's available points earned.
- Needs Work50–79% of the category's available points earned.
- Critical< 50% of the category's available points earned.
Action ranking
Top actions are ranked deterministically: an impact estimate (0–100) is multiplied by a priority factor (P0 ×1.25, P1 ×1.0, P2 ×0.7) and divided by an effort cost (S ×1.0, M ×1.6, L ×2.6). The result is impact per unit of engineering effort, which is what a commerce team actually optimises for.
What Commerce4 does not claim
- We do not guarantee ranking, recommendation or placement by ChatGPT, Gemini, Copilot or any other AI platform.
- We do not guarantee revenue, traffic or conversion outcomes. Revenue figures in reports are illustrative scenarios, not forecasts.
- We do not have privileged access to any AI platform's retrieval, ranking or merchant selection logic.
- Agent test results are simulations of representative buying intents, not observations of live production agent behaviour.
- A high score reduces avoidable technical and semantic failure. It does not make your product the right answer to a buyer's question.
- Scores are comparable only within the same methodology version.
Second surface: Agent Browser Readiness v0.1.0
The seven categories above score the UCP surface: server-to-server protocol readiness. Since Shopify enabled WebMCP across Liquid storefronts, an agent can also operate the store inside the buyer’s own session. That is a different failure mode, so it gets its own 0–100 score with its own weights. We never average the two — a store can be readable by protocol and unusable in session, or the reverse.
| Browser category | Weight |
|---|---|
| Tool exposure & discoveryWhich WebMCP storefront tools the store exposes, and whether an agent can discover and describe them. | 20 |
| Product & variant resolutionWhether in-session search and product tools resolve to the correct product and the correct variant or configuration. | 20 |
| Cart action fidelityWhether add-to-cart preserves quantity, variant and LINE attributes carrying per-item configuration, plus cart-level attributes for order-wide metadata. | 25 |
| Checkout progressionWhether the agent can progress to checkout without blockers, with truthful totals, shipping and taxes. | 20 |
| Policy & support answerabilityWhether shipping, returns, warranty and FAQ questions are answerable by the agent in session. | 15 |
Browser evals are recorded as six independent metrics — discovery success, correct product, correct variant or configuration, browser-action success, cart fidelity and checkout readiness — rather than a single pass/fail, because losing the configuration between builder and checkout is a different defect from never finding the product. Until evals run in a Chromium origin-trial session with the merchant, every transactional metric stays pending and the browser score is provisional.
UCP capability layers, and why they are not scored
UCP is a fixed standardized schema. Commerce4 reports three separate diagnostic checks that never change the score: base UCP protocol readiness, Shopify enriched catalog extension readiness (dev.shopify.catalog, version 2026-04-08, extending catalog search/lookup with gift_card and collections on products, requires.shipping, requires.selling_plan, checkout_url and selling_plans on variants, plus the available request filter defaulting to true), and decision completeness — how much of the buyer's decision the standardized schema can actually carry.
The pinned stable UCP version is 2026-04-08. Unreleased work on the UCP main branch around request constraints and payment-instrument grammar is watchlist only and must not affect compliance scoring. For a store whose configuration lives outside the schema, the honest diagnosis is: discoverable through UCP, but part of the purchase intent may need WebMCP or merchant-specific application state to complete — which is a hypothesis, not a demonstrated capability.
Line attributes vs cart-level attributes
Cart-level attributes are order-wide or session-wide metadata: gift message, order note, agent provenance. Line attributes carry item-specific configuration: engraving, size, material, page count, design reference, printing instructions. Configuration fidelity is verified on line attributes — a cart-level attribute cannot represent each configured item. We also test that the same variant with different line attributes stays as separate cart lines. Cart-attribute update and event checks are reserved for cart-wide metadata and verification.
MCP security boundary
The Commerce4 Customer/Product MCP at /mcp is public and read-only, and serves seed and public-web data only — never lead records. The Lovable Build/Admin MCP is a separate account and project control surface and is never exposed to customers or autonomous agents. OAuth, least privilege and tier-based access are mandatory before private merchant audits or any write tool are exposed.
Private merchant access & credential isolation
Production merchant credentials never flow through builder-level or shared-workspace connections. Authorization is merchant-owned, tenant-isolated and revocable, or the work does not run.
Public web audit
Independently observable public-web information only. No credentials are requested, held or stored, and no private data is read.
Live
Internal builder test
Builder or shared-workspace connections are permitted only for tightly controlled internal development and testing. They are never merchant-owned production authorization.
Internal only · blocked for production
Private merchant audit
Per-merchant OAuth or an App User Connector equivalent: merchant-owned authorization, least-privilege scopes, isolated per-tenant credentials, server-side or gateway token storage, revocable at any time.
Not connected · planned
Private-access checklist
- · Credential owner
- · Workspace exposure
- · OAuth scopes
- · Per-user isolation
- · Token storage
- · Revocation path
Each item is evaluated by a deterministic validator, not by copy. A missing revocation path, unknown token storage, absent per-tenant isolation, a broad unjustified scope, or any builder / shared-workspace credential blocks the workflow from production. Current private-access state: NOT CONNECTED — no merchant has authorized private access and no private connection exists in this project.
Merchant confidentiality & relationship neutrality
Any storefront named in Commerce4’s public surfaces appears only as an independent public-web reference. Independent public-web reference. Inclusion does not state or imply customer status, partnership, endorsement, participation, authorization, or private access.
Named profiles use only independently observable public-web information. No private, API, account or relationship data appears anywhere in these outputs. Public profiles serialize an explicit allowlist of public-web observations with their evidence level. Private API payloads, credentials, connection details, admin, order, customer, analytics, revenue or non-public catalog data never appear in public UI, MCP output, exports or logs.
Where a merchant authorizes private access, the results stay private. Such work may improve Commerce4 and benefit products delivered to other merchants only as generalized, non-attributable, non-reidentifiable rules, evals, safeguards or product behavior. Raw or merchant-specific data is never transferred between clients, and private data is never used for model training, secondary evaluation or any unrelated purpose without prior written, purpose-specific authorization. No private-derived benchmark or example that could reveal its source is published.
Consent gate: logos, testimonials, case studies, named performance deltas, integration mentions or any statement of relationship require prior written, purpose-specific approval from that merchant. Consent for one asset or channel never implies consent for another. Without it, we use anonymized or synthetic examples.
Evidence standard
Every deduction carries a locatable artefact: a URL, a robots.txt line range, a validator output, a rendered-vs-source diff or a failed agent transaction. If we cannot show you the evidence, we do not deduct the point. Recommendations are written as implementable changes with an effort size, not as generic advice.