Redaction
Redaction is the selective removal or concealment of information in a record. It can target names, contact details, confidential passages, or other excluded content while retaining the rest. Its effectiveness depends on the file format, method, and subsequent checks, not just how the document looks on screen.
How it works
A workflow identifies the content to exclude, applies an appropriate transformation, and verifies the resulting artifact. In digital files, review may need to consider searchable text, comments, revision history, metadata, embedded objects, or linked copies. Structured datasets can require removal across related tables rather than one visible cell.
Automated detection can help locate candidates, but missed context and false matches require testing. The transformed output should be checked in the form a recipient will actually receive.
Why it matters for licensing
Redaction can help prepare a narrower dataset, but it is one control within a broader rights and disclosure assessment. Records can still reveal identities or commercial facts through remaining context. The license and documentation should describe relevant exclusions without exposing the removed information.
Example
Fictional example: A business removes customer names from service reports and then tests the exported files. It discovers that names remain in document metadata and corrects the export process before conducting further disclosure review.
Limitations and misconceptions
A black rectangle drawn over text may leave the underlying text recoverable. Even genuine removal of direct identifiers does not prove anonymity or resolve all confidentiality obligations. Redaction also changes data utility and may remove context needed for a technical task.
Questions to ask
- Has the information been removed from the actual delivered artifact and its metadata?
- Could remaining context or related files reveal the same information?
- How are redaction accuracy and effects on usefulness checked?