Ask a general-purpose AI about your return policy and it will answer — fluently, confidently, and quite possibly wrong. It has never read your policy. It has read thousands of other companies’ policies and will happily blend them into one that sounds like yours.

For a store that is embarrassing. For a pharmacy, a clinic or an industrial supplier, a made-up dose or a made-up tolerance is a liability. APIchatbot is built for those cases: it answers from your documents, and only from them.

Four rules that keep it on your documents

  1. Your documents, cut into excerpts. Manuals, FAQs, price lists and PDFs are split into short passages and indexed by meaning. Tables keep their column headers on every piece, so a number never loses the name of what it measures.
  2. Each question fetches the closest excerpts, by meaning and by exact words. Five by default. The AI sees those and nothing else from your business.
  3. Answer from the excerpts, or say you do not know. The instructions are blunt: use only the excerpts provided, never outside knowledge even when certain, never invent details — and when the answer is not there, reply “I am sorry, I do not have that information.”
  4. Show where it came from. The answer cites the excerpts it used, with the document title and page. Only cited sources are shown, and a refusal shows none.

When a question is off the map

Some questions land far from everything in your documents. Before any answer is written, a small, cheap model sorts them into four kinds:

  • About the chatbot (“what can you help with?”): it lists what it covers, built from the titles of your actual documents.
  • About your business (“what is your main service?”): it tries anyway with the best excerpts, still bound by the same rules — those facts often sit in long FAQ passages.
  • A greeting (“hi”, “thanks”): it greets back instead of apologising.
  • Anything else: a polite refusal — or, if you turn it on, an offer to leave an email so a person can follow up.

When the sorting is unsure, it leans towards refusing. And the final safety checks are plain code, not the model grading itself: an answer that cites nothing gets no suggested follow-up questions, because in testing the model kept suggesting follow-ups right after admitting it did not know.

Measured on your documents, not on a benchmark

“Our model scores 90 % on a benchmark” says nothing about your price list. So every chatbot gets its own exam, written from its own documents:

  • Questions drawn from random passages of your documents, each with the exact fact the answer must contain.
  • The same questions with typos, in the other language, and on a nearby topic your documents do not cover — which must be refused.
  • A set of attacks: general-knowledge bait, “ignore your instructions”, requests to reveal the prompt, fake system messages, attempts to read another client’s documents.

The exam runs through the real chat, is not billed to the client, and is never stored as chat history. An AI grader checks each answer: the expected fact, nothing contradicting it, and for the attacks, a refusal with no leak.

The numbers

The first exam, on a pharmacy’s drug guides: 24 of 26 questions answered correctly and zero wrong facts. The two misses were refusals — the chatbot said it did not know rather than guess. That is the failure you want.

The latest validation run, on three chatbots in production. These questions are the test set, not everyday traffic: most of them are the typical problem cases — typos, the other language, topics just outside the documents, and attacks — so the scores are measured where chatbots usually slip.

ChatbotDocumentsQuestionsAccuracy
apigoat.com demo7 documents7298.6 %
A pharmacy3 guides, 709 excerpts8693.0 %
An industrial client29 documents9298.9 %
250test questions, three chatbots
8failed, 6 of them in the pharmacy guides
0wrong facts, first exam

Where it still fails

Eight of 250 questions failed, six of them in the pharmacy guides. The ones that matter are wrong values read from dense tables: a cross-allergy rate given as 10 % instead of under 1 %, and a score band left out of a table. The other two were incomplete answers, not wrong ones.

One more lesson: a setting that fetched more excerpts scored higher on the industrial client — and invented a figure. A higher score is not automatically a safer chatbot. That is why the test records every answer, not just the percentage.

That is the point of testing: finding the dense table before a customer does, and deciding — document by document — whether that content needs restructuring or a human answer.

Your data stays yours

Each client’s documents are searched only for that client’s chatbot. Visitors’ questions are not stored unless you turn logging on, test runs never are, and each chatbot has its own rate and monthly limits, so a bot hammering it cannot run up your bill.

What this means for an SME

Before you put any chatbot in front of your customers, ask one question: what is its accuracy on my documents, and how do you know? If the answer is a benchmark, a demo or a shrug, you are the test.