Methodology · v0.1.0

What we measure, how we weight it, and what we refuse to claim.

Commerce4 measures the technical and semantic readiness of a commerce catalog for AI shopping agents. Every point in the score maps to a condition that can be observed on your storefront and verified independently.

Category weights

Weights reflect where agent-driven revenue is most often lost, in our observation of Shopify catalogs. Weights are versioned; changing them produces a new methodology version so historical scores stay comparable.

CategoryWeight
Catalog & Product Semantics25
Structured Data & Machine Readability15
AI Discovery / Retrieval Readiness15
Merchant Policies & Trust Data10
Inventory, Price & Variant Clarity10
Cart / Checkout Agent Readiness15
Measurement & AI Traffic Attribution10
Total100

Readiness tiers

  • Critical

    0–39

    Agents cannot reliably identify, price or purchase your products. Fundamental catalog or access issues are present.

  • Emerging

    40–59

    Products are partially retrievable but qualifier-based queries and checkout handoff fail frequently.

  • Competitive

    60–79

    Most intents resolve. Remaining losses concentrate in edge configurations, bundles and attribution.

  • Agent-Ready

    80–100

    Catalog, policies and checkout are machine-legible end to end, with monitoring in place to catch regressions.

Category status

  • Strong≥ 80% of the category's available points earned.
  • Needs Work50–79% of the category's available points earned.
  • Critical< 50% of the category's available points earned.

Action ranking

Top actions are ranked deterministically: an impact estimate (0–100) is multiplied by a priority factor (P0 ×1.25, P1 ×1.0, P2 ×0.7) and divided by an effort cost (S ×1.0, M ×1.6, L ×2.6). The result is impact per unit of engineering effort, which is what a commerce team actually optimises for.

What Commerce4 does not claim

  • We do not guarantee ranking, recommendation or placement by ChatGPT, Gemini, Copilot or any other AI platform.
  • We do not guarantee revenue, traffic or conversion outcomes. Revenue figures in reports are illustrative scenarios, not forecasts.
  • We do not have privileged access to any AI platform's retrieval, ranking or merchant selection logic.
  • Agent test results are simulations of representative buying intents, not observations of live production agent behaviour.
  • A high score reduces avoidable technical and semantic failure. It does not make your product the right answer to a buyer's question.
  • Scores are comparable only within the same methodology version.

Second surface: Agent Browser Readiness v0.1.0

The seven categories above score the UCP surface: server-to-server protocol readiness. Since Shopify enabled WebMCP across Liquid storefronts, an agent can also operate the store inside the buyer’s own session. That is a different failure mode, so it gets its own 0–100 score with its own weights. We never average the two — a store can be readable by protocol and unusable in session, or the reverse.

Browser categoryWeight
Tool exposure & discoveryWhich WebMCP storefront tools the store exposes, and whether an agent can discover and describe them.20
Product & variant resolutionWhether in-session search and product tools resolve to the correct product and the correct variant or configuration.20
Cart action fidelityWhether add-to-cart preserves quantity, variant and LINE attributes carrying per-item configuration, plus cart-level attributes for order-wide metadata.25
Checkout progressionWhether the agent can progress to checkout without blockers, with truthful totals, shipping and taxes.20
Policy & support answerabilityWhether shipping, returns, warranty and FAQ questions are answerable by the agent in session.15

Browser evals are recorded as six independent metrics — discovery success, correct product, correct variant or configuration, browser-action success, cart fidelity and checkout readiness — rather than a single pass/fail, because losing the configuration between builder and checkout is a different defect from never finding the product. Until evals run in a Chromium origin-trial session with the merchant, every transactional metric stays pending and the browser score is provisional.

UCP capability layers, and why they are not scored

UCP is a fixed standardized schema. Commerce4 reports three separate diagnostic checks that never change the score: base UCP protocol readiness, Shopify enriched catalog extension readiness (dev.shopify.catalog, version 2026-04-08, extending catalog search/lookup with gift_card and collections on products, requires.shipping, requires.selling_plan, checkout_url and selling_plans on variants, plus the available request filter defaulting to true), and decision completeness — how much of the buyer's decision the standardized schema can actually carry.

The pinned stable UCP version is 2026-04-08. Unreleased work on the UCP main branch around request constraints and payment-instrument grammar is watchlist only and must not affect compliance scoring. For a store whose configuration lives outside the schema, the honest diagnosis is: discoverable through UCP, but part of the purchase intent may need WebMCP or merchant-specific application state to complete — which is a hypothesis, not a demonstrated capability.

Line attributes vs cart-level attributes

Cart-level attributes are order-wide or session-wide metadata: gift message, order note, agent provenance. Line attributes carry item-specific configuration: engraving, size, material, page count, design reference, printing instructions. Configuration fidelity is verified on line attributes — a cart-level attribute cannot represent each configured item. We also test that the same variant with different line attributes stays as separate cart lines. Cart-attribute update and event checks are reserved for cart-wide metadata and verification.

MCP security boundary

The Commerce4 Customer/Product MCP at /mcp is public and read-only, and serves seed and public-web data only — never lead records. The Lovable Build/Admin MCP is a separate account and project control surface and is never exposed to customers or autonomous agents. OAuth, least privilege and tier-based access are mandatory before private merchant audits or any write tool are exposed.

Private merchant access & credential isolation

Production merchant credentials never flow through builder-level or shared-workspace connections. Authorization is merchant-owned, tenant-isolated and revocable, or the work does not run.

Public web audit

Independently observable public-web information only. No credentials are requested, held or stored, and no private data is read.

Live

Internal builder test

Builder or shared-workspace connections are permitted only for tightly controlled internal development and testing. They are never merchant-owned production authorization.

Internal only · blocked for production

Private merchant audit

Per-merchant OAuth or an App User Connector equivalent: merchant-owned authorization, least-privilege scopes, isolated per-tenant credentials, server-side or gateway token storage, revocable at any time.

Not connected · planned

Private-access checklist

  • · Credential owner
  • · Workspace exposure
  • · OAuth scopes
  • · Per-user isolation
  • · Token storage
  • · Revocation path

Each item is evaluated by a deterministic validator, not by copy. A missing revocation path, unknown token storage, absent per-tenant isolation, a broad unjustified scope, or any builder / shared-workspace credential blocks the workflow from production. Current private-access state: NOT CONNECTED — no merchant has authorized private access and no private connection exists in this project.

Merchant confidentiality & relationship neutrality

Any storefront named in Commerce4’s public surfaces appears only as an independent public-web reference. Independent public-web reference. Inclusion does not state or imply customer status, partnership, endorsement, participation, authorization, or private access.

Named profiles use only independently observable public-web information. No private, API, account or relationship data appears anywhere in these outputs. Public profiles serialize an explicit allowlist of public-web observations with their evidence level. Private API payloads, credentials, connection details, admin, order, customer, analytics, revenue or non-public catalog data never appear in public UI, MCP output, exports or logs.

Where a merchant authorizes private access, the results stay private. Such work may improve Commerce4 and benefit products delivered to other merchants only as generalized, non-attributable, non-reidentifiable rules, evals, safeguards or product behavior. Raw or merchant-specific data is never transferred between clients, and private data is never used for model training, secondary evaluation or any unrelated purpose without prior written, purpose-specific authorization. No private-derived benchmark or example that could reveal its source is published.

Consent gate: logos, testimonials, case studies, named performance deltas, integration mentions or any statement of relationship require prior written, purpose-specific approval from that merchant. Consent for one asset or channel never implies consent for another. Without it, we use anonymized or synthetic examples.

Evidence standard

Every deduction carries a locatable artefact: a URL, a robots.txt line range, a validator output, a rendered-vs-source diff or a failed agent transaction. If we cannot show you the evidence, we do not deduct the point. Recommendations are written as implementable changes with an effort size, not as generic advice.