Data retention
Data retention is the policy and practice governing how long data remains stored or accessible. Retention periods can reflect operational needs, contractual duties, and legal requirements. For personal data, applicable rules may limit keeping identifiable information beyond the relevant purpose, subject to specific conditions or exceptions.
How it works
Map the copies and artifacts created during processing, identify the purpose of each, and define the applicable retention rule. Deletion, restricted archival storage, or other end-of-term treatment should be operationally achievable. Backups, audit logs, and legal holds may require explicit handling.
A dataset and a trained model are different artifacts. An agreement should address each where relevant rather than assuming that deleting source files automatically reverses all effects of training.
Why it matters for licensing
A license can specify access periods, return or deletion duties, and permitted retention of particular outputs. These terms affect delivery architecture and recipient obligations. They should also be consistent with upstream permissions and applicable privacy or recordkeeping requirements.
Example
Fictional example: An evaluation license requires working copies to be deleted after the project, while an agreed audit record retains limited administrative evidence. The parties separately define backup handling and whether any derived artifact may remain.
Limitations and misconceptions
A calendar date alone does not prove deletion across every system. Conversely, immediate deletion may conflict with a valid legal hold or other obligation. Retention decisions require coordination between legal requirements and the actual storage and processing workflow.
Questions to ask
- Which source, derived, backup, and model artifacts will exist?
- What purpose or obligation supports each retention period?
- How will end-of-term actions and exceptions be verified?