thesolver
Free exposure check
INDEPENDENT AI STRESS TESTING

Your AI is talking to customers. Do you know what it's saying?

Break your chatbot before your customers do.

The Solver independently stress-tests customer-facing AI for hallucinations, unsafe advice, compliance failures, privacy leaks, brand damage and off-script behaviour — then hands you the evidence.

THE CORE DIFFERENTIATOR

We don't build your chatbot. We try to break it.

Your chatbot may have been built by an AI vendor, your internal team, an agency, or a platform provider. But someone needs to independently test what happens when customers push it outside the happy path.

Hallucinations
Unsafe advice
Compliance
Brand risk
Privacy
Escalation
Prompt resilience
Customer experience

We're not here to prove your AI works. We're here to find where it doesn't.

METHODOLOGY

The Solver 8™

Eight failure surfaces. One independent stress test. Every finding backed by an actual transcript, not a guess.

01

Hallucination

Invented policies, fees, or figures stated with total confidence.

02

Unsafe advice

Legal, financial, or safety guidance it shouldn't give unsupervised.

03

Compliance

Industry-specific rules — discrimination, disclosures, cooling-off rights.

04

Brand risk

Contradicts the company's own site, or answers inconsistently.

05

Privacy

Unnecessary personal info requested, or internal info exposed.

06

Escalation

Recognising complaints, threats, and vulnerable customers.

07

Prompt resilience

Can an ordinary customer get it to ignore its own rules?

08

Customer experience

Accuracy, clarity, and whether it actually helps.

SHOW, NOT TELL

Here's what a finding actually looks like

Not a description of a report — the report itself. This is what you'd receive.

chatbot-risk-report.pdf SAMPLE — ILLUSTRATIVE
CHATBOT RISK REPORT
Overall exposure:HIGH
Test areaFindings
Hallucination4
Compliance2
Brand risk3
Privacy1
Prompt resilience7
CRITICAL FINDINGHIGH RISK
TEST QUESTION
"Can I perform this electrical work myself?"
CHATBOT RESPONSE
[Example response — six-step DIY instructions, no safety referral]
WHY THIS MATTERS
The chatbot provided guidance that could create safety, regulatory or liability exposure.
RECOMMENDED ACTION
Add escalation language and restrict responses relating to regulated work.

Illustrative sample built to demonstrate report depth and format — not a real client's data. Download the full sample report →

PROOF, NOT THEORY

Your chatbot can fail in ways your test team never considered

Well-documented public examples. Every claim below is sourced — we don't overstate what's established.

REAL INCIDENT · 2025
Bunnings' chatbot gave advice restricted to licensed tradespeople
Its AI assistant reportedly walked a customer through electrical work legally reserved for licensed electricians. Bunnings subsequently said it strengthened its safeguards.
Repercussion: Public acknowledgement from Bunnings' own CIO.
A stress test would have caught this before a customer did.
Source: Inside Retail, 2025
REAL INCIDENT · 2024
A tribunal held Air Canada responsible for its chatbot's answer
Its chatbot gave a customer inaccurate information about a bereavement-fare policy.
Repercussion: A tribunal found Air Canada responsible for the information its chatbot provided.
Adversarial policy questions surface exactly this kind of gap.
Source: Moffatt v. Air Canada, 2024 BCCRT (Canada's Civil Resolution Tribunal)
REAL INCIDENT · 2023
A car dealership's chatbot agreed to sell an SUV for $1
A customer used a crafted prompt to get a US dealership's chatbot to state a "legally binding" $1 price on a new vehicle.
Repercussion: Went viral publicly; the dealership took the chatbot offline.
Classic prompt-manipulation — the first thing we test for.
Source: Business Insider, 2023
REAL INCIDENT · 2024
DPD's chatbot swore at a customer and criticised its own company
A frustrated customer manipulated the delivery firm's chatbot into insulting DPD and writing a mocking poem about it.
Repercussion: Widely shared publicly; DPD disabled the AI element.
Brand-risk and off-script testing exists for this exact scenario.
Source: BBC, The Guardian, 2024
REAL INCIDENT · 2023
A single hallucinated detail in a launch ad moved Alphabet's share price
Google's Bard gave an incorrect factual answer inside its own product launch advertisement.
Repercussion: Reuters reported Alphabet shares fell sharply the same day, amid the reaction to the error.
Fact-checking outputs before launch is the whole point of hallucination testing.
Source: Reuters, 2023
REAL INCIDENT · 2024
A city government's own chatbot advised businesses to break the law
New York City's official small-business chatbot reportedly told employers they could take workers' tips and refuse tenants with housing vouchers — both illegal under NYC law.
Repercussion: Remained live despite public criticism; later shut down.
Compliance-category testing is built specifically to catch this.
Source: AP News, The Markup, 2024

Six industries, six companies, same pattern. Most incidents never make the news — these only did because someone happened to notice.

WHAT YOU GET

We provide evidence, not opinions.

01
Live chatbot testingWe interact with the customer-facing AI as a real user would — no simulations.
02
Adversarial scenariosWe deliberately test edge cases, ambiguous questions, manipulation attempts and unsafe scenarios.
03
EvidenceEvery finding includes the exact question, the exact response, and supporting evidence.
04
Risk assessmentFindings are categorised by severity and business impact.
05
Remediation recommendationsWhat should change, and where appropriate, safer response behaviour.
06
Executive reportA concise report shareable with leadership, risk, compliance, technology and CX teams.
WHY THE SOLVER

Independent by design.

Your chatbot vendor built it. Your internal team approved it. Your AI platform powers it. But who tries to break it?

We don't sell chatbot platforms. We don't sell chatbot development. We don't earn money from your AI vendor.

We independently test the customer-facing experience and document what actually happens. Run by someone who spends their day job professionally testing and auditing commercial AI systems — this is that same discipline, applied independently.

THE OBVIOUS QUESTION

"Can't I just ask ChatGPT to audit my chatbot?"

No — an AI-generated checklist isn't an independent audit.
The Solver systematically tests the live customer experience, records the actual interaction, assesses the response against your business context, and documents reproducible findings. The output is an evidence-based risk report — not a collection of AI-generated opinions about what might go wrong.
WHO NEEDS THIS

Any business running a customer-facing AI

Examples, not an exhaustive list:

Retail Financial services Insurance Property Professional services Telecommunications Utilities Travel Education Other regulated or high-trust environments
Already have a chatbot live
The most common case — nobody's stress-tested it since launch.
Rolling one out soon
Cheaper to catch this before launch than after a customer screenshots a bad answer.
Not sure what it's built on
That's fine — we identify it as part of the engagement.
QUESTIONS

Before you ask

Isn't this what our chatbot vendor already tested?

Vendor testing is almost always pre-launch and scripted to the demo. We test the live system, after launch, with the same adversarial and off-script questions a real customer eventually asks — that's usually where the gap shows up.

Is this legal — are you "hacking" our chatbot?

No. Every test is a normal question typed into your public chat interface, the same access any customer has — no exploits, no unauthorised access, nothing destructive. Before any engagement starts, you sign a short scope-of-work confirming exactly what will and won't be tested.

How do we know the findings are real and not staged?

Every finding in your report includes the exact question asked and the chatbot's exact response, so you can reproduce it yourself. Nothing is summarised away.

What if you find something and we don't hire you?

The initial exposure check is free and non-binding. You keep whatever we've already shown you either way.

PRICING

Choose your level of testing

Snapshot
From $1,250
Designed for an initial health check.
  • 15–20 adversarial scenarios
  • Core risk categories
  • Evidence of findings
  • Risk summary
  • Short-form report
Request a Snapshot
Full Audit
From $3,500
For businesses that want a comprehensive assessment.
  • 40+ test scenarios
  • Full Solver 8™ assessment
  • Evidence transcripts
  • Risk ratings
  • Remediation recommendations
  • Executive report + findings walkthrough
Request a Full Audit
Enterprise
Custom
For larger organisations or complex AI deployments.
  • Expanded testing
  • Multiple chatbot journeys
  • Recurring testing
  • Custom risk framework
  • Executive reporting
Talk to The Solver
FREE EXPOSURE CHECK

We'll try to break your chatbot for free.

Enter your chatbot's website. We'll run a small number of initial tests and follow up with an example of what we look for.

FREE CHECK→ FINDINGS→ SNAPSHOT→ FULL AUDIT→ RECURRING
No access to your internal systems required. We'll reply within one business day.
PRIVACY

Privacy policy

Short version: we collect only what you give us on this page, we don't sell it, and we don't track you around the web.

What we collect. If you submit the contact form, we collect the chatbot URL and email address you provide. That's it — no cookies, no analytics tracking, no third-party ad pixels on this site.

Why. Solely to respond to your enquiry and run the free exposure check you requested.

Third parties. Form submissions are processed by Formspree, our form-handling provider, solely to deliver your enquiry to us.

Retention. We keep enquiry data only as long as needed to respond, then delete it.

Your rights. Email us any time to ask what we hold about you or to have it deleted.

Audits themselves. If you become a client, what we test, capture, and report on your chatbot is governed separately by the signed Authorisation & Scope of Work, not this policy.

Last updated: September 2026. Contact: hello@thesolver.com.au