Work · 03 Document search · Knowledge
Seven sets of help pages, searchable inside the assistant they already used.
The staff at a field-services company already used an AI assistant every day. They ran the business on seven pieces of software, and none of those vendors' help pages were easy to search. One had no working search at all. The question was whether the assistant could answer "how do I do this in that tool" from the vendor's own help pages. It had to name the page it used. And the person had to stay inside the assistant.
- Client
- A field-services company. Not named.
- Layer
- 3 of 4. Knowledge.
- Sources
- Seven vendors' help pages.
- Measured
- 2026-09-19, counted straight from the files.
- Published
Service Knowledge bases for AI assistants How an engagement runs
A search system is worth building only where the cheaper answers fail, so I test those first. If one vendor's help pages are short enough, roughly 500 pages, the whole lot can be handed to the assistant directly and no search system is needed. A public question-and-answer bot is a chat box you rent by the month. A company on Microsoft 365 with Copilot licences should first try pointing Copilot at its own files. This build is for what is left over. Several vendors at once. Pages that sit behind a login or cannot be exported. Answers that must name their source. An index the client keeps on their own machine. And their own assistant as the way in.
01 The constraint
What was actually hard.
Not the first vendor. The seventh. Every vendor builds its help site differently. A mistake in reading one of them does not show up as an error. It shows up weeks later as a wrong answer. And the most common failure of all is cutting a page in the middle of a numbered procedure, so the assistant returns half a set of steps. The work is making seven sources behave the same way, so that adding the eighth is a copy rather than a project.
02 The mechanism
How it works.
- The same six steps for every vendor: collect the pages, read them, store them, index them, search them, then offer that search to the assistant. Adding a vendor means copying the folder and rewriting the first two steps only. Nothing after that changes.
- Plain files the client owns. Each set is a file of pages, a file of passages, a summary file and a search index. They can be opened, compared and rebuilt. Nothing is locked to a vendor. The client can rebuild the index from those files alone, on their own machine. That is the honest answer to "what happens when I stop paying you".
- Pages are cut into passages at their own headings, roughly one to two pages each. Tables and numbered steps are never cut in half.
- Each passage gets a short opening line naming the document and section it came from. That line is added before the passage is indexed, so a passage found on its own still says where it belongs.
- Two searches run at once: one matching words, one matching meaning. Their results are merged and re-sorted. The twenty best sections come back in full rather than as a summary, so the assistant can quote and name its source.
- The index is built on the client's own machine with a free, openly published model. Changing that model later means rebuilding the index, not rebuilding the system.
- Each set is offered to the assistant as a single search command, over the same protocol as build 04. The assistant searches. It does not list and read.
- The test is a list of real questions written by the client's own expert, each with the section that should answer it. That list is re-run after every change to the index.
One vendor's help pages are collected, read into a plain file of pages the client owns, cut into passages at their headings with steps never split, and given a line saying where each passage came from. An index by words and by meaning is built locally. Searches are merged and re-sorted, and the twenty best sections go in full to a search command, one per vendor, offered over the protocol into the assistant already open on the staff's desks. A list of real questions written by the client's own expert tests it, and the client can rebuild all of it from the files alone.
03 The figures
What it measured.
Each figure carries the date it was taken. A number without a date is not on this page.
| What was measured | Value | Measured |
|---|---|---|
| Document sets built and indexed, all to one shape | 7 | 2026-09-19 |
| Help pages across the seven | 5,899 | 2026-09-19 |
| Passages across the seven | 15,788 | 2026-09-19 |
| Where the index is built | on the client's own machine | 2026-09-19 |
Why the design has these steps: published results from a vendor's own test, not from this build
| What the vendor published | Value | Read |
|---|---|---|
| Fewer times the right section was missed, after adding the source line to each passage | 35% fewer | 2026-09-19 |
| After also indexing those passages by word match | 49% fewer | 2026-09-19 |
| After also re-sorting the results | 67% fewer | 2026-09-19 |
| Cost of writing those source lines | about $1.02 per million words of documents | 2026-09-19 |
| Size below which no search system is needed at all | roughly 500 pages | 2026-09-19 |
First: if a vendor's help pages come to less than about 500 pages, the whole lot can be handed to the assistant directly and there is nothing to build. I measure that for each source and each client before proposing anything. What justified a search system here was seven separate sources, each needing to be named as the answer's origin, at a combined size too large to hand over. Second: a company on Microsoft 365 with Copilot licences should first try pointing Copilot at its own files. If that answers their test questions, that is my recommendation. One more thing, because buyers ask: there is no reliable published price for work like this. Every number in circulation comes from a vendor's own blog. So I price this against the result and against the cheaper alternatives, never against a "market rate".
04 Doing it again
What it would take to do again.
Six of the eight parts never change between clients. The shape. The way pages are cut. The line that names each source. The two indexes. The search. And the command the assistant calls. Two parts are rewritten for each new vendor: collecting the pages and reading them. The test questions always belong to the client. Any company already using one of the seven products described here can have that document set on day one, because the files already exist.
A fit
Several separate sources, answers that must name where they came from, and an assistant staff already use daily.
Not a fit
One source under about 500 pages, or a company where Copilot already answers the test questions. I sell this alongside a larger build, not on its own.
Up against a constraint like this one?
Thirty minutes, no prep, no pitch. Bring the thing that is slow; if it is not a fit, I say so and point you somewhere useful.