This project provides a playbook for implementing personal data portability. It's intended for services that host personal data of some kind, and are unsure what approach to take to meet regulatory requirements, to provide data access that works for users and their third-party tools, and to keep their effort and maintenance costs low.
We're not compliance lawyers, so this isn't legal compliance advice! It's a resource to make compliance simpler if this basic approach satisfies your compliance requirements.
Our basic approach here is to help you build a HTTP+OAuth+JSON API quickly and efficiently. A REST style is used due to its overwhelming familiarity and existence of mature and scalable tools. The result should also have good data security, ops, resiliency and scaling characteristics.
┌─────────────────────────────┐
│ Custom Format │
├──────────────┬──────────────┤
│ JSON │ JSON Schema │
├──────────────┼──────────────┤
│ HTTP │ OAuth │
└──────────────┴──────────────┘
This playbook is for online services, such as those with a website or app, that store their users' personal data in the cloud. It might be most immediately useful for organisations assessing the technical feasibility of implementing user-initiated direct transfers. In time, it might become directly relevant for organisations brought into scope of existing or new regulatory requirements or data sharing initiatives. Equally, the playbook may be a useful resource for policy makers considering the design or potential implications of introducing such requirements.
Many such services do other things besides hold personal data. As an example, Ravelry (one of my favourite sites) was initially collecting community data about knitting patterns and yarn, and that led to collecting personal data about knitters' projects and ratings of patterns and yarn. As Ravelry's community grew, knitters' post history became another important type of personal data. How should a service like Ravelry expose all this personal data for third-party access and portability?
Other examples:
| Service Type | Personal Data Types |
|---|---|
| Music streaming services | Playlists and listen history |
| Video streaming services | Playlists and watch history |
| Note-taking services | Notes and folders |
| AI chat services | Chat history |
| Map services | Route history, favourites |
These are illustrative examples and are not intended to be exhaustive. Email and calendar servers should look to standards rather than follow this playbook.
Throughout, third party means the app or service a user authorizes to receive their data, which the data portability ecosystem calls a data destination.
There are two major ways to use this playbook
- Read it and see what ideas and pointers may be helpful. Maybe use the playbook to fill in your project plan.
- Point your AI coding agent at this playbook and ask it to follow the playbook until you're ready to deploy
Because this playbook is documenting common and widespread practices for exposing data, there are many examples in open source and published APIs. Rather than explaining 'endpoint', 'JSON object' or 'cursor', this playbook assumes that the reader already knows, can research, or can delegate the details.
- Identify what counts as portable data
- Pick some libraries and frameworks for the project
- Use JSON Schema to define data formats
- Map your service's data storage to your external format
- Choose appropriate API endpoints
- Implement access control keyed to the user
- Handle references to OTHER users in personal data
- Allow efficient async requests for chunks of bulk data with pagination or cursors
- Use API keys for defensibility and rate-limiting
- Use OAuth for authorization of data access, with limited scopes
- Log grants and data access requests
- Publish documentation
- Design GUI elements
- Plan for cloud deployment/ops
- Support API discovery (including data schemas and OAuth scopes)
- Plan for the future with a schema evolution strategy
Some Internet regulations include data portability requirements.
- GDPR Article 20 — the right to receive personal data in a structured, commonly used, machine-readable format, and to have it transmitted to another controller.
- EU Digital Markets Act (DMA) — Article 6(9), which requires gatekeepers to give end users and the third parties they authorize continuous and real-time access to data generated through use of the service.
- UK Data Protection Act 2018 — which carries the UK GDPR's equivalent portability right post-Brexit.
There are two terms in the DMA text in particular that have driven specific design choices in this playbook:
Continuous personal data access doesn't necessarily require a streaming or push protocol. On a practical basis, we interpret it to mean that the user, or a third party they've authorized, can come back and pull newly created data on a regular ongoing basis, rather than just receiving a one-time export. Using OAuth and pagination or cursors is powerful and flexible.
Real-time also doesn't necessarily require a streaming or push protocol. As long as the API can access data as soon as it's saved in your service, the API can give access to real-time data with latency that satisfies most use cases.
When determining the ideal speed and frequency of transfers, we encourage data portability implementers to consider the parameters that will be most useful for the potential downstream uses of the data by users and third parties. Open Data Institute research on this topic defined this approach as "Functional Real-Time".
It's also possible to augment this API with WebSockets or another streaming or push-based mechanism, but this RESTful approach is still where to start.