SYS/WORK ยท 03INTELLIGENCE SYSTEMS

Answers with a paper trail.

Collection, enrichment, correlation and reporting for operational BI, research workflows and OSINT-heavy problems, built so every figure can be traced back to where it came from.

An answer you cannot trace back to a source is a guess with a chart on it.

What this covers

Collection from APIs, registries, documents and scrapes. Normalisation and entity resolution. Enrichment and correlation. Storage that keeps the raw record next to the derived one. And the query or reporting layer people actually open.

Plus the unglamorous half that decides whether any of it survives contact with reality: rate limits, schema drift, retention, access control and a licence check on every source.

OUR POSITION

Provenance is the product. Collection is the easy part.

Any team can pull a dataset and draw a chart. What decides whether anyone can act on the result is being able to answer "where did this number come from" six months later, when the source has changed its schema and the analyst has moved on. So the raw record stays, transformations are recorded as steps rather than overwritten in place, and every derived figure keeps a link back. Language models earn their place in triage and summarising; they stay out of the chain of custody, because a model cannot be cross-examined.

When this is the right work

  • The same question gets a different answer depending on who you ask.
  • The analysis runs in a notebook on one laptop, and the laptop is the system.
  • Sources are gathered by hand and go stale between reports.
  • A finding has to survive scrutiny from someone who was not in the room.
  • Personal data is in scope and retention has never been decided.
  • Two systems disagree about which records describe the same entity.

What changes

Sources are inventoried, with their licence terms and rate limits written down. Collection runs on a schedule and fails loudly instead of quietly returning yesterday. Entities are resolved once rather than per report. Every figure in a report opens back to its raw record.

Retention and access become decisions somebody made, rather than defaults nobody chose.

A note on lawfulness

OSINT work touches personal data, terms of service and, in the Netherlands, the AVG. We scope what is lawful to collect and retain before building the collector, not after it is running. If a source cannot be used for the purpose you have in mind, that is a finding, and it is much cheaper now than later.

How we work on it

Start from the question the system has to answer and work backwards to the sources that can answer it. Model the entities before writing a collector. Keep transformation steps separable so a wrong assumption can be replayed rather than re-collected. Then hand over something an analyst can extend without asking an engineer.

Tools we tend to reach for

Chosen per problem, not per fashion. This is what the shelf looks like.

MESSY INFORMATION PROBLEM? OPEN /CONTACT โ†’