Field guide
The Legible Archive
Most of what we call a digital archive is really a pile of files with a search box bolted to the front. It works for a year or two. Then the encoding drifts, the URLs move, the one person who understood the folder structure leaves, and a collection that took a decade to assemble becomes unreadable in the space of a single server migration. This site is about the unglamorous craft that prevents that.
The Legible Archive is an independent, non-commercial publication about how public reference collections get built and — much harder — how they get kept. It is written for the people who actually do the work: the librarian handed a room of unscanned volumes and no budget, the volunteer maintaining a subject encyclopaedia that has outlived three content management systems, the small team that assembled a gazetteer and now has to keep every place name pointing somewhere sensible.
Legibility is a design problem, not a storage problem
Storage is cheap and getting cheaper. Legibility is not. A scanned page is only useful if something can read it; a dictionary entry is only useful if its cross-references resolve; a map is only useful if the place it names can still be found. Every one of those is a decision made at build time by somebody who may never see the consequence.
The through-line of everything published here is that reference material has an ergonomics of its own. Data that is shaped for the convenience of the database will be miserable for the reader and impossible for the next maintainer. Data shaped for the reader — explicit structure, stable addresses, honest encoding, documented provenance — survives handover after handover. The Digital Preservation Coalition has spent years making this argument to institutions with budgets; the same logic applies just as hard to a two-person project running on a single server.
What is covered here
The material falls into four groups, and they build on each other.
Getting text off paper
Everything starts with capture. Digitising reference books covers cradles, lighting, page order and the decisions that quietly determine whether a scan is a preservation master or a dead end. Scanning and OCR deals with what happens after: why recognition accuracy on a dictionary behaves nothing like accuracy on a novel, and why the layout matters more than the letterforms.
Giving it structure
A reference work is not prose with headings — it is a database that happens to have been typeset. Structuring an encyclopaedia looks at entries, senses, cross-references and the modelling mistakes that are almost impossible to reverse later. Metadata and catalogues covers the description layer, and character encoding covers the single most common way a non-English archive quietly destroys itself.
Putting it on the map
Reference collections and geography are old companions — gazetteers, atlases, regional histories. Web cartography and routing traces how map publishing moved from proprietary map servers to open tiles, and point-of-interest data is about the messiest layer of all: the named places themselves.
Keeping it alive
Link rot and permanent URLs is the maintenance discipline nobody budgets for. Open licensing is about making a collection legally reusable, which is often what saves it. And knowledge bases and chat guides looks at the recurring dream of a reference work you can simply ask a question — a much older idea than it currently appears.
There is also one piece of local history: Before the app store, on the strange, short-lived discipline of cataloguing software for hundreds of incompatible mobile handsets. It is here because this domain once hosted exactly such a catalogue, and because the compatibility problem it describes never really went away.
The bias of this site
Three positions run through everything, stated plainly so you can discount for them.
- Plain formats beat clever ones. A directory of well-named files with a documented convention has outlived a great many databases. Complexity has to earn its place.
- URLs are part of the collection. An address that changes is a citation that breaks. This is not a technical detail; it is the difference between a reference work and a website.
- Write down what you did. Scanner settings, encoding choices, the reason a field exists. The maintainer who needs that note is usually you, four years later, with no memory of the decision.
Independence
Nothing here is for sale and there is no organisation behind it. There are no affiliate arrangements with any of the tools or services named, and where a specific product is mentioned it is because it is genuinely the reference implementation, not because of any relationship with its makers. Read more on the about page, which also sets out what this site is not.