The refusal, designed
Every tool that reads documents will answer you. This one is built to decline. When the best passage in your records does not clear two floors, it quotes nothing and says one sentence instead. This page explains the gate behind that sentence, then publishes its one measurement so far — which it failed.
Grounded, or silent
A confident wrong answer looks exactly like a confident right one. If you bill from the number, that is the whole problem. So this engine has a second output besides the quote — a refusal — treated as a designed state with a reason attached, not as an error.
The sentence
ABSTAIN_TEXT"Your records don't contain enough to answer that with confidence."
That is the exact string in the code, and the only thing the engine says when it declines. No partial answer, no best guess, no citation to a passage it did not trust. Beside it, the engine records why: the score the best passage reached, the fraction of your question's words it covered, and which floor it fell under. A refusal is as auditable as an answer.
Note what it does not say. It does not say your records lack the answer; it says the engine could not find it with enough confidence to quote it — the document may never have been filed, may be a photograph not yet read, or may have scored too low. The reason line tells you which.
Two conditions, both
score floor · coverage floorBefore a passage can be quoted it has to pass two tests, and it has to pass both. The first is a score floor: the passage's fused retrieval score, described on the answering page, must reach a minimum. The second is a coverage floor: of the meaningful words in your question, the passage must contain at least a set fraction. A passage that scores well but mentions only half of what you asked is refused; one that mentions everything but scores poorly is refused. Neither floor can rescue the other.
Both floors are set by hand. Nobody derived them from data — the engine's own notes say so — which is why a calibrated gate exists alongside them, described below. If nothing in the index matched at all, the engine refuses before either floor is consulted.
Why silence beats a guess
for someone who bills from the numberTake a contractor pricing a bore. The estimate goes on a bid; the bid becomes a contract; the number on the contract becomes the invoice. They ask what rate per metre the last comparable job ran at. If the records hold it, the engine quotes the line, names the job file, and they check. If they do not — the ticket was never scanned — a tool that answers anyway still produces a rate, in the same confident voice, from whatever passage scored highest. It looks the same as the real one. It goes on the bid.
A refusal costs a phone call. A wrong rate costs the margin on the job, or the job. That asymmetry is the whole argument, and why the refusal is a first-class output here — designed, reasoned and recorded — rather than grey text.
The measurement
once, and it failedWhat happened when the gate was tested
Abstention has been measured once, on three of the founder's own corpora, with questions the records could not answer. The gate failed: it produced 45 unsafe answers — answers to questions it should have declined. Fixing the embedder afterwards moved retrieval and left abstention where it was: better finding did not make the engine better at knowing when it had not found.
A calibrated gate was then built, described below, and fitted on one corpus. Its honest guarantee there: about one unanswerable question in seven still gets answered. That is the measured figure, not a target — and it was fitted on the same cases it was scored against, so it is optimistic by construction. No page on this site claims a better number, because none has been measured.
The hand-set floors still serve any corpus without a calibration; where one exists, it replaces them. The record carries the whole measurement; this page changes when the number does.
What is being done
and what would close itThe calibrated gate works like this. Take questions you know your records cannot answer. Run each through the engine and note the signal its best passage produced. Choose a threshold so that, at a stated error rate, a new unanswerable question falls below it — that choice carries a finite-sample guarantee, provided new questions resemble the calibration ones. Answer only when the signal is strictly above the threshold.
Three properties matter. If the calibration set is too small for the requested error rate, the threshold goes to infinity and the engine abstains on everything rather than pretend. If the corpus is re-ingested after the calibration was fitted, the engine refuses to serve that calibration, because its guarantee no longer holds. And the price of the gate — how many answerable questions it wrongly refuses — carries no guarantee, so it is measured separately.
What would close it is not complicated, only unfinished: a calibration set per corpus, held apart from the questions the result is scored on; the one-in-seven re-measured on questions the gate never saw; the same done for every corpus it serves, and re-done whenever a corpus changes. Until that is published, the gate is a mechanism with one measurement against it — the one above.
Why this page exists
The product that hides this measurement and the one that publishes it are not the same product. One asks you to trust a sentence. The other shows you the count, the caveat and the plan. When the gate is good enough, you will read it here with the same grading as the failure: measured, on which corpus, fitted on what.
Read next
the rule, the record, the engineGrounded or silent
Abstention is the product; the citation is the proof. The one rule every other page is written under.
The rule ▸ Built · specified · never runThe record
Every measured shortfall in one place, each with what would close it — the 45 unsafe answers among them.
Read the record ▸ Steps one to fiveHow it answers
Pooling, three legs fused, a deterministic rerank, the verbatim quote — the path up to the gate.
How it answers ▸