THE RIGHT SUPPORT, AT THE RIGHT STAGE

A baseline.
A QA partner.

Start with your chat agent.
Add ongoing QA capacity as your product grows.

01 / Chat Agent Audit: Silent Failure Mapping

Chat Agent Audit:
Silent Failure Mapping.

I find where your chat agent fails and what to fix first, with evidence and a follow-up check.

You get a detailed report covering every test run, with evidence, causes and a prioritized fix plan.

Functional testing & response evaluation

One-time · fixed scope

Completed in 5 business days after access.

See the eight-stage process

Clear for founders.
Ready for engineers.

  • Every test run, with evidence, causes and a prioritized fix plan
  • Exact conversations, test conditions and expected versus actual behaviour
  • Severity, business impact, observed frequency and reproduction status
  • Confirmed or suspected causes, with suggested fixes or investigation steps
  • Behaviours to preserve, coverage limits, evidence gaps and supporting records
Findings walkthrough. Evidence package. Documented baseline.

One bounded fix-verification round within 14 calendar days of report delivery.

What your audit includes.

I test the conversations your agent needs to handle, then verify the failures and investigate their causes.

  • Single-turn and multi-turn conversation testing across the six quality areas
  • Negative testing: unanswerable questions, out-of-scope requests and false premises
  • Boundary testing at the edges of your knowledge base and agreed scope
  • Consistency testing: the same question repeated and rephrased
  • Memory testing across extended conversations
  • Manual exploratory testing alongside the scripted suite
  • Suspected failures repeated to establish frequency, then labelled reproduced, intermittent or needing investigation
  • Test-script errors separated from genuine agent failures
  • Causes confirmed where evidence supports it, and clearly labelled as suspected where it does not

Six quality areas.
Your important workflows.

01

Accuracy & grounding

Correct answers supported by approved documents, sources or tool results.

02

Relevance & completeness

Answers that address the question and include necessary information.

03

Conversation memory

Follow-ups, earlier details and user corrections handled correctly.

04

Instructions & business rules

Required steps, appropriate tone and agreed boundaries.

05

Difficult inputs

Typos, unclear requests, conflicting details and frustrated users.

06

Uncertainty & recovery

Useful clarification, admission of missing knowledge and handoff.

Did the action actually happen?

For agents that make bookings or update a CRM, I also check task completion. Verification requires access to the resulting records; missing evidence is reported explicitly.

WHERE ACCESS ALLOWS

Fixed price, never hourly. Scope agreed in writing before testing.

AGREED BEFORE TESTING

Clear limits.
Useful findings.

The audit covers the agreed scope. It does not certify overall reliability.