Claims and evaluation
methodology.
A practical, evidence-based standard for turning observations into defensible statements. We state what we know, how we know it, and what it does not prove.
From question
to impact.
A measurement begins with one decision question and an owner. “Faster” is not a baseline; “median time from submitted request to assigned owner during comparable operating weeks” can be one.
Start with the question
Define the decision question, owner, and baseline before the intervention.
Establish a baseline
Capture the unit, time window, data sources, exclusions, and current process.
Apply the right measures
Use time, quality, reliability, business, and human-centered measures.
Gather evidence
Collect verifiable data, logs, and artifacts from the operating environment.
Review and state claims
Evaluate results, account for limitations, and communicate clear, qualified claims.
What we measure.
A consistent set of measures for AI systems and operational change, selecting the ones that fit the workflow and the business context.
Time
Cycle time, response time, handling time, or staff time sampled from a defined process.
Quality
Error, rework, escalation, and completion or acceptance rates with a written definition.
Reliability
Successful runs, failed runs, recovery time, and manual intervention requirements.
Business effect
Revenue, cost, capacity, or retention when records can be linked to the measured process.
Human effect
Adoption, operator confidence, accessibility, and workload from surveys or observation.
Key evaluation
areas.
Analytical rigor with operational realism. Each area follows a documented procedure and uses verifiable evidence. Open one for the full standard.
AI system evaluations
Use representative tasks with sensitive data removed or controlled. Define expected outputs, prohibited actions, tool permissions, escalation behavior, and evaluation dates.
The evaluation also defines the cost of a false positive and a false negative. Model versions, prompts, retrieval sources, tools, and evaluation dates are recorded because results can change when any of them changes.
High-impact actions require human review or a deterministic control appropriate to the risk. A fluent answer is not evidence of correctness. For retrieval systems, citations or source traces should be checked against the underlying record. For agents, the action log matters as much as the final text.
Automation and software evaluations
Test deterministic workflows for happy path, exceptions, failures, recovery, and integrations. Validate conflict behavior, retries, alerts, and reconciliation methods.
Duplicate events, delayed events, invalid inputs, and permission failures are part of the test set, not an afterthought. Integrations must define the source of truth, conflict behavior, retry policy, alert owner, and reconciliation method.
Backups are not described as recoverable until a restoration procedure has been tested in the relevant environment.
ROI and time-saved claims
Use documented costs and measured benefits over a defined period. Include implementation, licenses, maintenance, training, review time, and change management.
Costs also include discovery, model usage, and hosting. Time returned to staff is not automatically cash savings; the claim should state whether the capacity was actually removed, reassigned, or simply made available.
We do not transfer a percentage from another company into a new prospect's forecast. A range may be modeled during discovery, but it remains a scenario until the client's own baseline and post-launch data support it.
Case-study standard
Document the problem, initial state, intervention, data sources, measurement window, results, limitations, and client-review status. Remove identifying details when needed.
Anonymous evidence must remain auditable even if the public version removes identifying details. A composite or demonstration can explain architecture, but it is labeled illustrative and contains no implied client outcome.
Local facts and service-area claims
Use municipal, county, census, planning, and transit sources. Be clear whether coverage is on-site or day-trip. Do not convert service availability into a false office, address, or team claim.
Population and landmarks are included only when they clarify the operating environment; they are not inserted as keyword decoration. A location page also makes no implied review or customer claim.
Search and answer-engine reporting
Separate impressions, clicks, qualified visits, leads, and attributable outcomes. Rankings are sampled and logged. A single manual search is not a reliable trend.
Rankings are observed by query, location, device, and date where data permits. AI citations and answer-engine visibility are sampled and logged, but no vendor can guarantee inclusion because the result is controlled by the search or answer platform.
Evidence, transparency,
and context.
Questions about a claim or a method can be submitted through the corrections process
Traceable sources
Citations, logs, and source traces back claims.
Clear limitations
We document assumptions and boundaries.
Qualified statements
Claims are specific, measured, and not overstated.