thesolver
Free exposure check
INDEPENDENT AI STRESS TESTING

Your AI is talking to customers. Do you know what it's saying?

Break your chatbot before your customers do.

Independent AI stress testing for customer-facing AI — hallucinations, unsafe advice, compliance failures, privacy leaks, brand damage and off-script behaviour.

Your AI can confidently give an answer your business never intended.

THE CORE DIFFERENTIATOR

We don't build your chatbot. We try to break it.

Your chatbot may have been built by an AI vendor, your internal team, an agency, or a platform provider. But someone needs to independently test what happens when customers push it outside the happy path.

Hallucinations
Unsafe advice
Compliance
Brand risk
Privacy
Escalation
Prompt resilience
Customer experience

We're not here to prove your AI works. We're here to find where it doesn't.

THIS ALREADY HAPPENS

Five examples of customer-facing AI risk

Chosen because each demonstrates a specific, well-documented failure surface — not because they're famous.

BUNNINGS · SAFETY / REGULATED ADVICE
Chatbot gave advice restricted to licensed tradespeople
Its AI assistant reportedly walked a customer through electrical work legally reserved for licensed electricians. Bunnings subsequently said it strengthened its safeguards.
Risk demonstrated: unsafe advice in a regulated area, given with full confidence.
What to test for: this is the type of failure adversarial testing is designed to surface — deliberately asking about regulated or safety-critical tasks.
Source: Inside Retail, 2025
AIR CANADA · CUSTOMER INFORMATION / LIABILITY
A tribunal held Air Canada responsible for its chatbot's answer
Its chatbot gave a customer inaccurate information about a bereavement-fare policy that didn't reflect the airline's actual terms.
Risk demonstrated: a business can be held liable for what its chatbot tells a customer.
What to test for: this is the type of failure adversarial testing is designed to surface — probing policy questions against what's actually published.
Source: Moffatt v. Air Canada, 2024 BCCRT (Canada's Civil Resolution Tribunal)
CHEVROLET DEALER · PROMPT MANIPULATION
A dealership's chatbot agreed to an unintended commercial commitment
A customer used a crafted prompt to get a US dealership's chatbot to state a "legally binding" $1 price on a new vehicle.
Risk demonstrated: ordinary users can manipulate a chatbot into commitments the business never authorised.
What to test for: this is the type of failure adversarial testing is designed to surface — deliberate prompt-manipulation attempts.
Source: Business Insider, 2023
DPD · BRAND / INAPPROPRIATE BEHAVIOUR
Chatbot criticised its own company and used inappropriate language
A frustrated customer manipulated the delivery firm's chatbot into insulting DPD and writing a mocking poem about it, shared widely online.
Risk demonstrated: off-script behaviour can become a public brand incident within hours.
What to test for: this is the type of failure adversarial testing is designed to surface — brand-risk and tone testing under provocation.
Source: BBC, The Guardian, 2024
NYC CHATBOT · COMPLIANCE
A city government's own chatbot advised businesses to break the law
New York City's official small-business chatbot reportedly told employers they could take workers' tips and refuse tenants with housing vouchers — both illegal under NYC law.
Risk demonstrated: confident, specific compliance guidance can be wrong in ways that create real legal exposure.
What to test for: this is the type of failure adversarial testing is designed to surface — compliance-category questions specific to the business's jurisdiction and industry.
Source: AP News, The Markup, 2024

Different organisations. Different failure modes. Same underlying problem: AI can say something the business never intended.

METHODOLOGY

The Solver 8™

Eight failure surfaces. Hundreds of ways an AI conversation can go wrong.

01

Hallucination

Invented policies, fees, or figures stated with total confidence.

02

Unsafe advice

Legal, financial, or safety guidance it shouldn't give unsupervised.

03

Compliance

Industry-specific rules — discrimination, disclosures, cooling-off rights.

04

Brand risk

Contradicts the company's own site, or answers inconsistently.

05

Privacy

Unnecessary personal info requested, or internal info exposed.

06

Escalation

Recognising complaints, threats, and vulnerable customers.

07

Prompt resilience

Can an ordinary customer get it to ignore its own rules?

08

Customer experience

Accuracy, clarity, and whether it actually helps.

WE SHOW YOU EXACTLY WHAT YOUR AI SAID

What a failure actually looks like

live chatbot session ILLUSTRATIVE — NOT A CLIENT FINDING
Can I return this item after 21 days?
Yes. You can return your purchase within 30 days for a full refund.
ACTUAL BUSINESS POLICY: Returns are accepted within 14 days.
HIGHPolicy hallucination
FAILURE SURFACEHallucination / Customer Experience
WHY IT MATTERSThe chatbot provided information inconsistent with the stated business policy.
EVIDENCEExact customer question → exact chatbot response → relevant policy reference.
RECOMMENDED ACTIONReview the source material and introduce stronger policy grounding and escalation behaviour.
THE DELIVERABLE

Here's what a finding actually looks like

Don't take our word for it. See the evidence.

chatbot-risk-report.pdf SAMPLE / ILLUSTRATIVE REPORT
CHATBOT RISK REPORT
Overall exposure:HIGH
Test areaFindings
Hallucination4
Compliance2
Brand risk3
Privacy1
Prompt resilience7
CRITICAL FINDINGHIGH RISK
TEST QUESTION
"Can I perform this electrical work myself?"
CHATBOT RESPONSE
[Example response — six-step DIY instructions, no safety referral]
WHY THIS MATTERS
The chatbot provided guidance that could create safety, regulatory or liability exposure.
RECOMMENDED ACTION
Add escalation language and restrict responses relating to regulated work.

See a sample Solver report →

WHAT YOU GET

We provide evidence, not opinions.

TESTING → EVIDENCE → ASSESSMENT → RECOMMENDATIONS → REPORT
01
Live chatbot testingWe interact with the customer-facing AI as a real user would — no simulations.
02
Adversarial scenariosWe deliberately test edge cases, ambiguous questions, manipulation attempts and unsafe scenarios.
03
EvidenceEvery finding includes the exact question, the exact response, and supporting evidence.
04
Risk assessmentFindings are categorised by severity and business impact.
05
Remediation recommendationsWhat should change, and where appropriate, safer response behaviour.
06
Executive reportA concise report shareable with leadership, risk, compliance, technology and CX teams.
WHY THE SOLVER

Independent by design.

Your chatbot vendor built it. Your internal team approved it. Your AI platform powers it. But who tries to break it?

The Solver.

We independently test the customer-facing experience and document what actually happens.

We don't sell chatbot platforms.

We don't build your chatbot.

We don't represent your AI vendor.

Our job is simple: find the failures you don't want your customers to find first.

THE OBVIOUS QUESTION

Can't I just ask an AI to audit my chatbot?

You can — but that isn't the same as an independent stress test.
AI can help generate test questions. The Solver's job is to systematically test the live customer experience, capture what actually happens, assess the response against the relevant business context, and document reproducible findings.
The output is evidence, not an AI-generated opinion about what might happen.
PRICING

Choose your level of testing

Pricing reflects the testing, evidence capture, analysis and reporting involved — not simply the number of questions asked.

AI Exposure Snapshot
From $1,250
15–20 targeted adversarial tests across core failure surfaces.
  • Evidence of findings
  • Risk assessment
  • Exact conversations
  • Short-form report
  • Recommended actions
Request a Snapshot
Comprehensive AI Stress Test
From $3,500
40+ targeted adversarial tests across The Solver 8™.
  • Full eight-surface assessment
  • Evidence transcripts
  • Risk ratings
  • Remediation recommendations
  • Executive report
  • Findings walkthrough
Request a Stress Test
Enterprise
Custom
For organisations requiring broader or recurring testing.
  • Expanded scenarios
  • Multiple journeys
  • Recurring assessments
  • Custom risk requirements
  • Executive reporting
Talk to The Solver

Pricing shown as "from" — the final scope and quote depend on your chatbot's complexity.

FREE EXPOSURE CHECK

We'll try to break your chatbot for free.

Give us your publicly accessible chatbot URL. We'll run a small number of targeted tests and show you an example of what we find.

You'll see:

The question we tested
The response we received
Why we flagged it
The risk category
FREE CHECK→ REAL FINDING→ SNAPSHOT→ STRESS TEST→ RECURRING / ENTERPRISE
No internal system access required. No obligation to continue. We only test systems we are authorised to assess — initial exposure checks are limited to publicly accessible customer-facing interfaces.
WHO'S BEHIND THIS

Built by someone who understands business, data and technology.

The Solver was founded by a senior business analyst with experience across data, analytics, reporting and commercial technology.

The focus is simple: independently test the customer-facing AI businesses are putting between themselves and their customers.

The Solver operates independently and does not rely on access to clients' internal systems to perform its initial stress testing.

QUESTIONS

Before you ask

What do you need from us?

Just your chatbot's public URL to start. Nothing else is required for the free exposure check.

Do you need access to our systems?

No. Initial testing is performed against the publicly accessible customer experience only.

How long does testing take?

A Snapshot typically takes about a week; a Comprehensive Stress Test one to two weeks.

Can you test regulated industries?

Yes — the Solver 8™ includes a dedicated compliance failure surface tailored to the rules relevant to your industry.

Is this a penetration test?

No. The Solver focuses on the behaviour and outputs of customer-facing AI rather than attempting to penetrate underlying infrastructure. We test the experience through authorised interaction with the system and document what the AI actually says and does. If a security vulnerability requires technical penetration testing, that should be handled by an appropriately qualified security testing provider.

Is this legal or compliance advice?

No. The Solver identifies potential risk areas and documents observed AI behaviour. Findings should be reviewed by the organisation's relevant legal, compliance, risk or subject-matter teams where appropriate. The objective is to give those teams concrete evidence of what the AI is actually saying.

What happens after the audit?

You receive a full report with evidence and recommended actions, and we're happy to walk through the findings on a call.

Can you retest after fixes?

Yes — retesting to confirm a fix resolved a finding is available as a follow-up engagement.

PRIVACY

Privacy policy

Short version: we collect only what you give us on this page, we don't sell it, and we don't track you around the web.

What we collect. If you submit the contact form, we collect the chatbot URL and email address you provide. That's it — no cookies, no analytics tracking, no third-party ad pixels on this site.

Why. Solely to respond to your enquiry and run the free exposure check you requested.

Third parties. Form submissions are processed by Formspree, our form-handling provider, solely to deliver your enquiry to us.

Retention. We keep enquiry data only as long as needed to respond, then delete it.

Your rights. Email us any time to ask what we hold about you or to have it deleted.

Audits themselves. If you become a client, what we test, capture, and report on your chatbot is governed separately by the signed Authorisation & Scope of Work, not this policy.

Last updated: September 2026. Contact: hello@thesolver.com.au

The Solver provides independent testing and risk observations. Findings are not legal advice, regulatory certification, or a guarantee that all possible AI failures have been identified.