Every problem you told us about has a fix, built and working on our test site. The last step, rebuilding the search library the new way, is happening now.
After that: we test it, you test it, a quiet fortnight, then live.
Think of a library with a fast assistant and a writer. Every Digest is cut into pages, and each page gets a catalogue card saying what it is about.
Your question is matched against the cards, and the ten best pages go to the writer, who answers only from those ten.
If the right pages are not found, the writer cannot give the right answer, whatever settings we change.
Nine Digests contain the phrase “vertical split head”. TD18-032 says it 28 times. The search found none of them.
Three reasons: the pages were too long (about 500 words, several topics each), the cards carried no number, title, year or author, and the search could not match exact words at all.
Not a temperature problem. It was already at its lowest, and the wrong answers were consistent. That told us the search was at fault.
Pages were about 500 words. Now they are about 200. A short page is about one thing, so it matches a specific question far better.
And because pages are shorter, the writer can be given three times as many for the same cost.
More evidence behind every answer.
Every page now carries its Digest number, title, year and author, written into the page itself so the search can read it.
That is what sends “summarise TD99-002” straight to TD99-002, and what makes author searches work. Author names are filled in automatically. Where one person appears under several spellings, we tidy those by hand.
The search engine has two modes: one for filing documents, one for answering questions. The live site used the filing mode for both.
Questions now use the question mode.
One line of code. Every search a little sharper.
The old search matched exact words. The chat matched meaning only. Rail terms and acronyms need exact words; open questions need meaning.
The chat now does both and merges the results.
“Vertical split head”: your three Digests, 0 of 3 found before, 3 of 3 now.
“VSH” and “vertical split head” now mean the same thing to the search. The glossary lives in the admin area.
It starts with four entries and needs many more: BHC, RCF, ECP, UTP, TTC, FRA and the rest.
We need an acronym list from MxV Rail.
“Publications by Duane E. Otter” used to rely on the search guessing a name. It found 2 relevant Digests in 10.
Now the author is written on every page and the search filters on it: 9 in 10 on our test site.
Once every author is filled in, you get the full list, not a sample.
One Digest with a lot to say could fill all ten pages and crowd out the rest.
There is now a limit per Digest, so a question like “which Digests discuss wheel flange wear” draws on several.
Before, “Did ENSCO become MxV Rail?” got an invented history, because ten loosely related pages were handed to the writer as if they were evidence.
Now: “No. The retrieved MxV Rail sources do not provide information about ENSCO becoming MxV Rail.”
Every citation is checked against the library before it is shown, so made-up filenames cannot appear.
“Summarise TD98-015” and “TD18-032” now give correct, cited summaries. TD99-002 showed the last gap: its PDF writes the number as “TD-99-002”, so the search missed it and summarised the nearest Digest instead. The label on every page (Fix 2) closes that.
“Summarise Stephen Wilk's publications” returned 8 Digests, all correct, out of 56. The full 56 come once every author is filled in.
The old search no longer jumps to the chat tab. Fixed on the live site on 25 August, and you confirmed it the same day.
Footnote numbers used to reload the page. They now jump to the reference.
Until now we only learned of a bad answer when you told us. The chat now keeps a log of every question and what it cited.
We propose keeping it for 30 days.
We need a yes from MxV Rail on keeping question text and IP addresses for 30 days.
The new library is built next to the old one, tested, and only then switched in. The old one is kept in case we need to switch back.
Nothing you use today changes until that switch. This is the step happening now.
18 test questions, built from your own. Each has a pass mark. Today on our test site, with the new code but the old library: 10 pass, 3 nearly, 5 fail.
The 5 fails are exactly the ones waiting on the rebuild. We have the “before” scores from 3 August, so improvement is shown in numbers, not impressions.