Our methodology

Claims and evaluation
methodology.

A practical, evidence-based standard for turning observations into defensible statements. We state what we know, how we know it, and what it does not prove.

How a claim is made

From question
to impact.

A measurement begins with one decision question and an owner. “Faster” is not a baseline; “median time from submitted request to assigned owner during comparable operating weeks” can be one.

01

Start with the question

Define the decision question, owner, and baseline before the intervention.

02

Establish a baseline

Capture the unit, time window, data sources, exclusions, and current process.

03

Apply the right measures

Use time, quality, reliability, business, and human-centered measures.

04

Gather evidence

Collect verifiable data, logs, and artifacts from the operating environment.

05

Review and state claims

Evaluate results, account for limitations, and communicate clear, qualified claims.

Our measurement areas

What we measure.

A consistent set of measures for AI systems and operational change, selecting the ones that fit the workflow and the business context.

01

Time

Cycle time, response time, handling time, or staff time sampled from a defined process.

02

Quality

Error, rework, escalation, and completion or acceptance rates with a written definition.

03

Reliability

Successful runs, failed runs, recovery time, and manual intervention requirements.

04

Business effect

Revenue, cost, capacity, or retention when records can be linked to the measured process.

05

Human effect

Adoption, operator confidence, accessibility, and workload from surveys or observation.

Detailed methodology

Key evaluation
areas.

Analytical rigor with operational realism. Each area follows a documented procedure and uses verifiable evidence. Open one for the full standard.

AI system evaluations

Use representative tasks with sensitive data removed or controlled. Define expected outputs, prohibited actions, tool permissions, escalation behavior, and evaluation dates.

The evaluation also defines the cost of a false positive and a false negative. Model versions, prompts, retrieval sources, tools, and evaluation dates are recorded because results can change when any of them changes.

High-impact actions require human review or a deterministic control appropriate to the risk. A fluent answer is not evidence of correctness. For retrieval systems, citations or source traces should be checked against the underlying record. For agents, the action log matters as much as the final text.

Automation and software evaluations

Test deterministic workflows for happy path, exceptions, failures, recovery, and integrations. Validate conflict behavior, retries, alerts, and reconciliation methods.

Duplicate events, delayed events, invalid inputs, and permission failures are part of the test set, not an afterthought. Integrations must define the source of truth, conflict behavior, retry policy, alert owner, and reconciliation method.

Backups are not described as recoverable until a restoration procedure has been tested in the relevant environment.

ROI and time-saved claims

Use documented costs and measured benefits over a defined period. Include implementation, licenses, maintenance, training, review time, and change management.

Costs also include discovery, model usage, and hosting. Time returned to staff is not automatically cash savings; the claim should state whether the capacity was actually removed, reassigned, or simply made available.

We do not transfer a percentage from another company into a new prospect's forecast. A range may be modeled during discovery, but it remains a scenario until the client's own baseline and post-launch data support it.

Case-study standard

Document the problem, initial state, intervention, data sources, measurement window, results, limitations, and client-review status. Remove identifying details when needed.

Anonymous evidence must remain auditable even if the public version removes identifying details. A composite or demonstration can explain architecture, but it is labeled illustrative and contains no implied client outcome.

Local facts and service-area claims

Use municipal, county, census, planning, and transit sources. Be clear whether coverage is on-site or day-trip. Do not convert service availability into a false office, address, or team claim.

Population and landmarks are included only when they clarify the operating environment; they are not inserted as keyword decoration. A location page also makes no implied review or customer claim.

Search and answer-engine reporting

Separate impressions, clicks, qualified visits, leads, and attributable outcomes. Rankings are sampled and logged. A single manual search is not a reliable trend.

Rankings are observed by query, location, device, and date where data permits. AI citations and answer-engine visibility are sampled and logged, but no vendor can guarantee inclusion because the result is controlled by the search or answer platform.

Our standards

Evidence, transparency,
and context.

Questions about a claim or a method can be submitted through the corrections process

Traceable sources

Citations, logs, and source traces back claims.

Clear limitations

We document assumptions and boundaries.

Qualified statements

Claims are specific, measured, and not overstated.