story · research_note · en
We Had the Data All Along
How we returned to Chance Brothers factory records with AI and turned an old source into a traceable research layer.

How we returned to Chance Brothers factory records with AI and turned an old source into a traceable research layer.
In 2022, while working with Chance Heritage Trust on the Lighthouse Mapping Project, we used many kinds of sources: official lighthouse pages, historical descriptions, catalogues, maps, photographs, and research publications.
But one source stayed with me.
It looked deceptively simple: several dozen photographed typewritten pages listing equipment supplied by Chance Brothers.
A year. A lighthouse name. A country. A class of apparatus. Line after line.
The opening page is headed "LIGHTS SUPPLIED BY C.B. & CO., LIMITED". Below it, equipment is grouped into classes such as Hyperradial Revolving Section, 1st Order Revolving, and 1st Order Fixed Section. Later pages introduce more classes, including lightvessel equipment and fog signals, and the chronology continues from 1896-1909 into 1910-1917.
At first sight, it is just a list.
In reality, it is a fragment of a global industrial network that existed more than a century ago.
And working with it was remarkably difficult.
When access to information is not yet access to knowledge
The problem was never that the document was inaccessible.
We could open it, enlarge the pages, and read every line. But turning one line into a reliable research record required a completely different kind of work.
What does the historical lighthouse name correspond to today? Has the name changed? Could another lighthouse have the same or a similar name? What exactly did Chance Brothers supply: a lens, burner, lantern, rotating apparatus, or another component? Does the date refer to manufacture, supply, or installation? Is the equipment still there? Was it later replaced, reconstructed, or moved?
And, most importantly: can we establish a traceable relationship between this particular historical piece of equipment and this particular lighthouse?
In 2022, much of this work had to be done manually. The real constraint was therefore not access to data. It was the human time required to turn data into verifiable knowledge.
Four years later
In 2026, we returned to this material.
This time, we had two things that were largely absent from our research process in 2022.
The first was modern AI. The second was our own LUX Light Archive research layer, capable of holding not only published facts but sources, claims, candidate entities, relationships, contradictions, uncertainty, and review status.
That changed the question.
Instead of asking, "Can we manually process this document?", we could begin asking, "What can this document tell us if we systematically compare it with the evidence we have accumulated since?"
Turning a page into evidence
Research layer
From historical source to public record
The stages are intentionally distinct. A readable line is not yet an accepted historical claim.
- 1Historical source
- 2Extracted statement
- 3Candidate claim
- 4Lighthouse and equipment identity
- 5Proposed relationship
- 6Cross-source corroboration
- 7Governed review
- 8Public heritage record
The important part is not allowing AI to transform a historical document directly into a collection of "facts."
Each stage has a different meaning and a different authority. A source is not yet a fact. A line read by AI is not yet a verified claim. A matching name is not proof that two records describe the same lighthouse. Two citations do not necessarily represent two independent pieces of evidence. Correctly identifying a historical piece of equipment does not mean that it remains at the same lighthouse today.
AI proved extremely useful precisely because it can accelerate the early stages of this chain. But it does not need the authority to skip the later ones.
A second source, and the complications begin
A good test came from another document used in the current research: South African lighthouses: Chance Brothers and the rest, a compilation credited in the document to Toby Chance.
It provides an almost perfect example of why simple automated extraction can be dangerous.
The first part contains a table of Chance or Stone-Chance equipment installations. But that does not mean that every row represents a Chance lens. Chance may have supplied another component.
More importantly, a later table is explicitly headed "Lighthouses where a Chance lens has never been installed." Some of those sites may nevertheless have used other Chance equipment, such as a burner or diaphone.
If we reduced the entire document to Lighthouse -> Chance Brothers, we would create more data and simultaneously lose much of its historical meaning.
35 rows as a small experiment
Research layer
August 2026 research snapshot
These figures describe the bounded 35-row experiment at the time of review; they are not current collection totals.
- 35records examined
- 31already represented as draft heritage assets
- 2new review-only leads
- 1attribution narrowed to a burner, not an optic
- 1rejected as non-Chance
- 0automatically published
Discovery is not publication.
We took 35 records from the relevant South African table and compared them with the research layer as it stood in August 2026.
Thirty-one already had corresponding draft heritage assets. Then came the exceptions.
Green Point, Cape Town
The record was missing. But another Green Point already existed in the database, in KwaZulu-Natal. A name-based automated match could therefore have merged two different places.
Green Point, Cape Town was added only as a review-only research lead. The source describes a Chance Brothers third-order optic installed in 1864, but independent host and apparatus confirmation remains separate work.
Ifafa
This record was also missing. The source refers to a Stone-Chance rotating beacon, so it was added as a research lead rather than published. The exact equipment identity and the way Stone-Chance should be represented still require review.
The Hill, Port Elizabeth
This case is even more revealing. There is a genuine Chance connection, but the table attributes the vapour burner to Chance Brothers while identifying J. Pintsch as the supplier of the optic.
Remove equipment type from the data model and it becomes easy to produce the historically incorrect statement "Chance Brothers lens at The Hill." That is not what the source says.
Cape St Martin
This produced the opposite result. The row describes AGA and VEGA equipment without stating a Chance or Stone-Chance component.
Instead of adding another object to the collection, the case was retained as rejected_non_chance. Its presence in the table remains useful audit evidence, but its position is not a manufacturer attribution.
And how many did AI publish?
Zero.
That may be one of the most important results of the experiment.
AI helped us read the material faster. It helped reconcile records, expose gaps, distinguish equipment types, create research leads, and identify potentially misleading attributions.
But none of those operations gave the system the authority to turn a research hypothesis automatically into a public historical record.
AI accelerated discovery. It did not receive authority to decide what becomes heritage record.
For us, this is not a limitation of AI. It is a property of the research architecture.
Sometimes "no" is a good research result
Research layer
Two rows that become clearer when equipment type is preserved
The source fragments are paired with the bounded interpretation recorded by the research workflow.
The Hill, Port Elizabeth

- Source supports
- Chance Brothers supplied a vapour burner; J. Pintsch supplied the optic.
- It does not support
- A Chance Brothers lens at The Hill.
- Research result
component_scope_corrected
Cape St Martin

- Source supports
- The row describes AGA and VEGA apparatus.
- It does not support
- A Chance or Stone-Chance attribution based only on the row's position in the table.
- Research result
rejected_non_chance
AI creates an almost invisible pressure towards quantity: more objects found, more relationships, more data, more automatically populated fields.
Cultural heritage does not work particularly well with that definition of success.
If Cape St Martin cannot be supported as a Chance-related object from the inspected row, the correct result is not to find a way to include it anyway. The correct result is to preserve what the source actually says, record why the tempting attribution was rejected, and make that decision traceable.
Likewise, narrowing The Hill from an apparent "Chance lighthouse" to a Chance burner paired with a J. Pintsch optic is not a smaller result. It is a more accurate one.
Sometimes research progresses by adding an object. Sometimes it progresses by separating two identities, correcting a component type, or recording a well-supported refusal.
The public collection and the hidden research layer
This distinction is visible in the way LUX presents Chance Brothers heritage.
The Chance Brothers Heritage collection is a public reader-facing collection. The Chance Brothers Research Atlas can show public-safe, source-attributed research coverage before every item qualifies for an ordinary public asset page.
Behind those surfaces sits a larger review layer: extracted statements, draft objects, possible matches, collisions, rejected attributions, and unresolved equipment relationships.
The research layer is not a waiting room in which every record is assumed to become public. It is where uncertainty can remain structured without being mistaken for a conclusion.
This matters especially for manufacturer history. A lighthouse may contain a Chance lens, a Chance burner, a Stone-Chance beacon, a Chance lantern, or no Chance component at all. The building, the optic, the light source, the rotating mechanism, and the navigational function can each have different makers and different histories.
What AI actually changed
AI did not make the historical material more authoritative. It made more of the material practically examinable.
It reduced the cost of tasks that previously consumed large amounts of human attention:
- transcribing repeated table structures;
- comparing historical and modern names;
- proposing candidate matches;
- locating likely collisions;
- separating equipment vocabulary;
- comparing source statements;
- finding contradictions and missing qualifiers;
- preparing bounded cases for review.
That shift matters because archives often contain far more information than researchers can afford to connect manually.
The bottleneck is no longer only digitisation. Increasingly, it is the ability to ask structured questions of digitised material while preserving provenance and uncertainty.
From digitising archives to questioning them
The first generation of digital heritage work focused on making material available: photograph the object, scan the page, create the catalogue entry, publish the database.
That work remains essential. But availability is not the end of the research process.
A scanned page can still be difficult to compare. A searchable PDF can still contain ambiguous names. A database can still collapse a lighthouse and its lens into one object. A map can still imply that historical equipment remains at its original location.
AI-assisted research becomes useful when it helps us ask better questions across those boundaries without erasing them.
Which rows probably refer to an existing object? Which names collide? Which statements describe supply rather than installation? Which equipment relationships are contradicted elsewhere? Which negative results deserve to be preserved? Which cases offer the greatest return from another hour of human research?
Those are not merely extraction questions. They are questions about evidence architecture.
The new scarcity is trust
When information was difficult to access, scarcity was often physical: the document was in an archive, the catalogue was out of print, or the photograph was not digitised.
Today, many projects face a different scarcity. We have images, OCR, databases, models, and generated summaries. What remains scarce is a trustworthy path from source to claim.
That path needs visible provenance, bounded interpretation, object identity, equipment specificity, uncertainty, review status, and a publication boundary.
Without those structures, AI can make weak knowledge arrive faster. With them, AI can help researchers find the exact point at which a claim becomes interesting, ambiguous, or unsafe.
We had the data all along
When we first worked with the Chance Brothers factory pages in 2022, the information was already in front of us.
We could see every page and read every line. What we could not yet do efficiently was connect the list to a growing body of lighthouse identities, equipment histories, source conflicts, and review decisions.
Four years later, the documents have not changed. Our ability to question them has.
The result is not simply more Chance Brothers records. It is a more precise account of different places, different kinds of equipment, ambiguous identities, attributions that could easily have become mistakes, and questions that were previously too expensive to ask.
In 2022, we helped turn dispersed research into a global map of lighthouse heritage. Today, we are beginning to explore what else those accumulated sources can tell us when humans and machines can read them together.
AI did not put more history into the archive.
We simply became better able to see the history that had been there all along.
Research note
This article describes an experimental AI-assisted workflow used within LUX Light Archive in continuing research into Chance Brothers lighthouse heritage.
AI is used to assist information extraction, entity reconciliation, detection of possible relationships and contradictions, and research prioritisation. AI-generated interpretations, candidate records, and research leads are not treated as verified historical assertions without separate evidence review and publication status.
The 35-record visual is an August 2026 snapshot of a bounded research experiment. It is not a live total, an independent historical source, or a claim that every represented record has since remained in the same review state.
Decisions concerning object identity, evidence quality, historical attribution, and publication remain part of a separate governed research process.
Sources
Research Documents
- Lights supplied by C.B. & Co., Limited, 1896-1909 / 1910-1917local_primary_sourceTwenty-page photographed company list used in the 2022 mapping research; source SHA-256 9c2ff2245ede02f73dc0103d9249451395669d1c63fadd3d7529c9fec6e9c00d. The project owner approved publication of the page-one excerpt, not the raw PDF.
- We Had the Data All Along - proposed article and visual brieflocal_research_inputOwner-supplied English editorial draft developed in conversation with ChatGPT. Used for article framing and prose, not as historical evidence.
Bibliography
- South African lighthouses: Chance Brothers and the restmanufacturer_referenceDirectly inspected table used for the bounded 35-row experiment and the Green Point, Ifafa, The Hill, and Cape St Martin examples. The two table populations retain different meanings.
- The Lighthouse Mapping Project - Chance Heritage Trustpartner_projectPublic account of the 2022 volunteer lighthouse mapping project.
- Chance Brothers Research Atlas - LUX Light ArchiveprimaryCurrent public LUX research surface; its live totals must not be confused with the article's August 2026 snapshot.
- Chance Brothers Heritage collection - LUX Light ArchiveprimaryPublic collection page linking the partner context, related research, and public member records.
Generated Package Files
Evidence
- No evidence recorded in package yet.
Open Questions
- No open questions recorded.