What it cannot read
Every file you give it ends up in one of two places. Readable text, with a note saying where it came from. Or a line on a list that says: this one, not yet, and here is why. That list is not a fault in the product. It is the product. This page is the list.
A silent skip is the only unrecoverable error
A wrong "could not read" is a problem you can see and fix in an afternoon. A file quietly dropped is a problem you find out about years later, from someone who is annoyed with you.
What it cannot read today
counts, not estimatesPhotographs of documents
A picture of a ticket taken on a phone is a picture. It is kept, it is counted, and it is marked picture only — and it is not searchable. Not yet. It is not turned into text, and the software does not pretend it was.
Some very old Word and Excel files
Files from before the current formats come in shapes no reader on the machine opens. In the founder's own records there are 787 of these. Each is on the list by name, marked no reader. They are not gone. They are waiting.
Scans that have not been read yet
A scanned PDF is a picture of a page. Turning it into text is a separate, slower pass. On the founder's main archive that pass has run: 7,996 scanned documents, 35,633 pages, all read on the machine, none of it sent anywhere. On a second set of records it has not started — 0 of 883. Those 883 are listed, each marked waiting to be read. Until that pass runs, those pages are not searchable, and the software says so.
Files that are in the cloud but not on the disk
Some folders show a file that is not really there. The name is there. The icon is there. The bytes are on somebody's server and the machine has never fetched them. In the founder's records 962 files are like this. The software does not skip them — it reports each one as not on disk, by name, so you can fetch it and run again.
Files that hang the reader
Out of 56,687 files in the main archive, two froze the program that opens them. Both were set aside, by name, for a person to look at.
Why the list matters more than the number
a true storyThe storage system on the machine had moved some files off the disk to save space and left the names behind. To the software those files looked present. When it tried to open them, they were not there. It marked them as failed. In one region of the archive, 87 per cent of files were set aside this way — not because the files were bad, but because the disk said it held them and did not.
The lesson written down that day: the storage layer will lie to you before the file formats do.
That sounds like bad news. It is the opposite. The software did not skip those files. It did not mark them read. It wrote each one down as unreadable, with the reason, and the reason pointed straight at the cause. Fetch the files, run again, done.
Had it quietly dropped them, 87 per cent of one region would be missing today and nobody would know. That is the whole idea of this page.
What "not yet" means
Not yet means the file is on the list, with a reason, and there is a known next step. Photographs wait for a reading pass that does not exist yet. Old files wait for a reader that has not been written. Queued scans wait for the reading pass to be run. Cloud files wait for you to fetch them.
None of these happens on its own, and none is promised by a date. When one is done, this page changes.
Where these numbers come from
All of them are from the founder's own records, which is the only body of material this has been run over at scale. There are no customers yet. Your own numbers will be different, and the software will tell you what they are rather than asking you to assume they match these.
If you take one thing from this site, take this. It will tell you what it could not read. Then you decide what to do about it.