donatelli.tech

Work · 03 Document search · Knowledge

Seven sets of help pages, searchable inside the assistant they already used.

The staff at a field-services company already used an AI assistant every day. They ran the business on seven pieces of software, and none of those vendors' help pages were easy to search. One had no working search at all. The question was whether the assistant could answer "how do I do this in that tool" from the vendor's own help pages. It had to name the page it used. And the person had to stay inside the assistant.

Client
A field-services company. Not named.
Layer
3 of 4. Knowledge.
Sources
Seven vendors' help pages.
Measured
2026-09-19, counted straight from the files.
Published

Service Knowledge bases for AI assistants How an engagement runs

A search system is worth building only where the cheaper answers fail, so I test those first. If one vendor's help pages are short enough, roughly 500 pages, the whole lot can be handed to the assistant directly and no search system is needed. A public question-and-answer bot is a chat box you rent by the month. A company on Microsoft 365 with Copilot licences should first try pointing Copilot at its own files. This build is for what is left over. Several vendors at once. Pages that sit behind a login or cannot be exported. Answers that must name their source. An index the client keeps on their own machine. And their own assistant as the way in.

01 The constraint

What was actually hard.

Not the first vendor. The seventh. Every vendor builds its help site differently. A mistake in reading one of them does not show up as an error. It shows up weeks later as a wrong answer. And the most common failure of all is cutting a page in the middle of a numbered procedure, so the assistant returns half a set of steps. The work is making seven sources behave the same way, so that adding the eighth is a copy rather than a project.

02 The mechanism

How it works.

  1. The same six steps for every vendor: collect the pages, read them, store them, index them, search them, then offer that search to the assistant. Adding a vendor means copying the folder and rewriting the first two steps only. Nothing after that changes.
  2. Plain files the client owns. Each set is a file of pages, a file of passages, a summary file and a search index. They can be opened, compared and rebuilt. Nothing is locked to a vendor. The client can rebuild the index from those files alone, on their own machine. That is the honest answer to "what happens when I stop paying you".
  3. Pages are cut into passages at their own headings, roughly one to two pages each. Tables and numbered steps are never cut in half.
  4. Each passage gets a short opening line naming the document and section it came from. That line is added before the passage is indexed, so a passage found on its own still says where it belongs.
  5. Two searches run at once: one matching words, one matching meaning. Their results are merged and re-sorted. The twenty best sections come back in full rather than as a summary, so the assistant can quote and name its source.
  6. The index is built on the client's own machine with a free, openly published model. Changing that model later means rebuilding the index, not rebuilding the system.
  7. Each set is offered to the assistant as a single search command, over the same protocol as build 04. The assistant searches. It does not list and read.
  8. The test is a list of real questions written by the client's own expert, each with the section that should answer it. That list is re-run after every change to the index.
ONE SHAPE, SEVEN TIMES help pages one vendor collect written per vendor file of pages plain text the client owns cut into passages at headings · steps never split say where it came from one line on every passage index by words and by meaning built locally merge · re-sort the 20 best sections, in full search command one per vendor, over the protocol existing assistant already open on their desk real questions written by the client's own expert tests it the client can rebuild all of this from the files alone

One vendor's help pages are collected, read into a plain file of pages the client owns, cut into passages at their headings with steps never split, and given a line saying where each passage came from. An index by words and by meaning is built locally. Searches are merged and re-sorted, and the twenty best sections go in full to a search command, one per vendor, offered over the protocol into the assistant already open on the staff's desks. A list of real questions written by the client's own expert tests it, and the client can rebuild all of it from the files alone.

The highlighted box is the one step that is not common practice. Each passage gets a short line naming the document and section it came from, and that line goes into both indexes. The published results below credit this step with most of the improvement. It is cheap, because the lines are written once when the index is built.

03 The figures

What it measured.

Each figure carries the date it was taken. A number without a date is not on this page.

What was measuredValueMeasured
Document sets built and indexed, all to one shape72026-09-19
Help pages across the seven5,8992026-09-19
Passages across the seven15,7882026-09-19
Where the index is builton the client's own machine2026-09-19
Counted straight from the files. A set counts as built only when all four of its files are there: the pages, the passages, the summary and the index.

Why the design has these steps: published results from a vendor's own test, not from this build

What the vendor publishedValueRead
Fewer times the right section was missed, after adding the source line to each passage35% fewer2026-09-19
After also indexing those passages by word match49% fewer2026-09-19
After also re-sorting the results67% fewer2026-09-19
Cost of writing those source linesabout $1.02 per million words of documents2026-09-19
Size below which no search system is needed at allroughly 500 pages2026-09-19
Source: anthropic.com/news/contextual-retrieval, read 2026-09-19. These published results are why the design has those three steps. I claim no accuracy number for this build. The only one that would mean anything is the client's own list of test questions.
One honest limitation: two disqualifiers applied against this line's own revenue.

First: if a vendor's help pages come to less than about 500 pages, the whole lot can be handed to the assistant directly and there is nothing to build. I measure that for each source and each client before proposing anything. What justified a search system here was seven separate sources, each needing to be named as the answer's origin, at a combined size too large to hand over. Second: a company on Microsoft 365 with Copilot licences should first try pointing Copilot at its own files. If that answers their test questions, that is my recommendation. One more thing, because buyers ask: there is no reliable published price for work like this. Every number in circulation comes from a vendor's own blog. So I price this against the result and against the cheaper alternatives, never against a "market rate".

04 Doing it again

What it would take to do again.

Six of the eight parts never change between clients. The shape. The way pages are cut. The line that names each source. The two indexes. The search. And the command the assistant calls. Two parts are rewritten for each new vendor: collecting the pages and reading them. The test questions always belong to the client. Any company already using one of the seven products described here can have that document set on day one, because the files already exist.

A fit

Several separate sources, answers that must name where they came from, and an assistant staff already use daily.

Not a fit

One source under about 500 pages, or a company where Copilot already answers the test questions. I sell this alongside a larger build, not on its own.

Up against a constraint like this one?

Thirty minutes, no prep, no pitch. Bring the thing that is slow; if it is not a fit, I say so and point you somewhere useful.