Dark data
Dark data describes stored information that remains outside useful analysis or decision-making. It can include old exports, attachments, logs, or records isolated in a department’s system. “Dark” describes how the organization uses and understands the information; it is not a file format or evidence that the data has commercial value.
How it works
An inventory can identify where records are stored, why they were created, who understands them, and whether they can be interpreted reliably. Some may support a new use once documented or connected to other information. Other records may be redundant, obsolete, inaccurate, or inappropriate to retain.
The useful first step is discovery and assessment, not an indiscriminate export. A system description and data dictionary can reveal more about feasibility than an unexplained file count.
Why it matters for licensing
Unused business records may prompt a licensing assessment, but utility and authority still need to be established. Preparation costs, retention obligations, personal information, and third-party restrictions can change whether further work makes sense.
Example
Fictional example: A facilities business finds historical repair logs in an archived system. It documents event meanings and coverage before considering whether the records support a defined evaluation task. Some incomplete exports are excluded.
Limitations and misconceptions
Unused does not mean valuable, unrestricted, or training-ready. Retaining data indefinitely in case it becomes useful can create costs and obligations. A commercial possibility should not be confused with verified demand.
Questions to ask
- What records exist, and who can explain how they were produced?
- Which plausible use would benefit from them?
- What quality, retention, or rights constraints affect reuse?
Sources
- IBM — Dark data · Accessed
- Gebru et al. — Datasheets for Datasets · Accessed
Explore whether your business data could be a fit.
Start with a description of your systems—not a data upload.