Break your chatbot before your customers do.
Independent AI stress testing for customer-facing AI — hallucinations, unsafe advice, compliance failures, privacy leaks, brand damage and off-script behaviour. Your AI changes. We check that its risks haven't.
Your chatbot may have been built by an AI vendor, your internal team, an agency, or a platform provider. But someone needs to independently test what happens when customers push it outside the happy path.
We're not here to prove your AI works. We're here to find where it doesn't.
Chosen because each demonstrates a specific, well-documented failure surface — not because they're famous.
Different organisations. Different failure modes. Same underlying problem: AI can say something the business never intended.
Eight failure surfaces. Hundreds of ways an AI conversation can go wrong.
Invented policies, fees, or figures stated with total confidence.
Legal, financial, or safety guidance it shouldn't give unsupervised.
Industry-specific rules — discrimination, disclosures, cooling-off rights.
Contradicts the company's own site, or answers inconsistently.
Unnecessary personal info requested, or internal info exposed.
Recognising complaints, threats, and vulnerable customers.
Can an ordinary customer get it to ignore its own rules?
Accuracy, clarity, and whether it actually helps.
Your chatbot might run on OpenAI, Google Gemini, Anthropic's Claude, Microsoft Copilot, Meta's Llama, or a custom model your team built in-house. It doesn't matter — The Solver tests what the chatbot actually says to a customer, regardless of what's powering it.
Don't take our word for it. See the evidence.
| Test area | Findings |
|---|---|
| Hallucination | 4 |
| Compliance | 2 |
| Brand risk | 3 |
| Privacy | 1 |
| Prompt resilience | 7 |
Your chatbot vendor built it. Your internal team approved it. Your AI platform powers it. But who tries to break it?
The Solver.
We independently test the customer-facing experience and document what actually happens.
We don't sell chatbot platforms.
We don't build your chatbot.
We don't represent your AI vendor.
Our job is simple: find the failures you don't want your customers to find first.
Chatbots get updated, retrained, and reconnected to new systems constantly — often without anyone re-checking what that changed about what it tells customers. A one-off stress test is a snapshot of one moment. Ongoing AI Assurance keeps checking.
Your original baseline questions, re-run every month, so nothing that passed quietly breaks.
Fresh attempts to break the chatbot each month — not the same tricks on repeat.
The facts your chatbot has to keep getting right — pricing, refunds, delivery, warranty, escalation — checked every cycle.
A person reviews every material finding before it reaches you. This is never a fully automated report.
Start free, establish a baseline, then keep assurance running month to month.
Launched a new chatbot version or major policy change between scheduled checks? A Major Change Retest (from $500) covers it without waiting for next month. Pricing shown as "from" — final scope depends on your chatbot's complexity.
Give us your publicly accessible chatbot URL. We'll run a small number of targeted tests and show you an example of what we find.
You'll see:
Just your chatbot's public URL to start. Nothing else is required for the free exposure check.
No. Initial testing is performed against the publicly accessible customer experience only.
An Initial AI Stress Test typically takes one to two weeks. Once you're on Ongoing AI Assurance, each monthly cycle runs on its own without you needing to chase it.
Yes — the Solver 8™ includes a dedicated compliance failure surface tailored to the rules relevant to your industry.
No. The Solver focuses on the behaviour and outputs of customer-facing AI rather than attempting to penetrate underlying infrastructure. We test the experience through authorised interaction with the system and document what the AI actually says and does. If a security vulnerability requires technical penetration testing, that should be handled by an appropriately qualified security testing provider.
No. The Solver identifies potential risk areas and documents observed AI behaviour. Findings should be reviewed by the organisation's relevant legal, compliance, risk or subject-matter teams where appropriate. The objective is to give those teams concrete evidence of what the AI is actually saying.
You receive a full report with evidence and recommended actions, and we're happy to walk through the findings on a call.
Yes. On Ongoing AI Assurance, one reasonable retest is included each month. Outside that — or if you've just shipped a major chatbot change — a Major Change Retest (from $500) covers it without waiting for the next cycle.
Short version: we collect only what you give us on this page, we don't sell it, and we don't track you around the web.
What we collect. If you submit the contact form, we collect the chatbot URL and email address you provide. That's it — no cookies, no analytics tracking, no third-party ad pixels on this site.
Why. Solely to respond to your enquiry and run the free exposure check you requested.
Third parties. Form submissions are processed by Formspree, our form-handling provider, solely to deliver your enquiry to us.
Retention. We keep enquiry data only as long as needed to respond, then delete it.
Your rights. Email us any time to ask what we hold about you or to have it deleted.
Audits themselves. If you become a client, what we test, capture, and report on your chatbot is governed separately by the signed Authorisation & Scope of Work, not this policy.
Last updated: September 2026. Contact: hello@thesolver.com.au
The Solver provides independent testing and risk observations. Findings are not legal advice, regulatory certification, or a guarantee that all possible AI failures have been identified.