The most common surprise in this whole area: publicly available personal data is still personal data. The fact that someone posted their name on a public page does not take it outside data protection law, and the intuition that "public means fair game" is simply wrong under the GDPR and its equivalents.
This matters for product scraping specifically, because personal data arrives in product datasets by accident far more often than by design.
Fields that look commercial and are not
| Field | Why it is personal data | Safer option |
|---|---|---|
| Marketplace seller name | Very often an individual or a sole trader | Keep a hashed seller id, or drop it |
| Review text and author | Written by an identifiable person, and sometimes revealing | Keep the rating count and average; drop the text |
| Q&A on a product page | Same — identifiable individuals | Drop entirely |
| Seller contact details | Directly identifying | Never collect |
| "Sold by" on a retailer listing | Corporate when it is a company, personal when it is not | Check which, or drop |
| Price, stock, SKU, title | Not personal data | Collect freely |
What the rules actually ask of you
Simplified considerably, and for the EU and UK regime specifically, there are four obligations that bite on collected data.
You need a lawful basis. For commercial research the usual candidate is legitimate interests, which requires you to balance your interest against the individual's rights and — importantly — to document that you did.
You have to tell people. There is a transparency obligation when you collect data about someone from a source other than them. There is a carve-out where notifying everyone would be disproportionate, but it is not automatic and it depends on you having thought about it.
You must minimise. Collect what the purpose needs and no more. This is also, conveniently, the rule that makes most of the problem disappear.
And you have to be able to respond. Individuals can ask what you hold and ask you to delete it, which in practice means you need to be able to find a named person in your dataset at all.
The simplest answer is usually the right one
Do not collect it.
A price monitoring system does not need seller names to work. It does not need review text. Dropping those fields at the point of extraction — not filtering them out later, but never writing them down — moves the entire project out of the personal data regime and removes a category of risk, a category of obligation and a category of conversation.
Where a seller identity genuinely matters for the analysis, a stable hash preserves the ability to say "this is the same seller as last week" without holding anybody's name. That is sufficient for almost every commercial question people actually ask of this data.
It is also simply less to defend. A dataset that provably contains no personal data is the shortest possible answer to a security review, and security reviews are where these projects most often stall.
Worked example: auditing a product feed, column by column
A marketplace price feed looks entirely commercial until you list the columns and ask one question of each: could this, alone or combined with the rest, identify a living person? Here is the same feed before and after that audit.
| Column | Identifies a person? | Needed for repricing? | Decision |
|---|---|---|---|
| product_title | No | Yes, for matching | Keep |
| price, currency, in_stock | No | Yes | Keep |
| gtin, mpn | No | Yes, for matching | Keep |
| seller_name | Often yes on a marketplace, where a large share of sellers trade under their own name | No | Drop |
| seller_address | Yes, frequently a home address | No | Drop |
| review_author, review_text | Yes, and review text is a direct opinion attached to a named person | No | Drop |
| q_and_a_username | Yes, pseudonymous but linkable | No | Drop |
What usually goes wrong
Nobody sets out to build a personal dataset. It assembles itself from columns that each looked harmless.
- Scraping the whole page because it was easier than selecting fields, then discovering a year later that the warehouse holds seller names and review text nobody ever used.
- Treating pseudonyms as anonymous. A stable username plus a purchase history plus a town is identifying in practice, whatever it looks like in isolation.
- Keeping seller_name "for debugging" and never removing it, so a temporary convenience becomes a permanent category change.
- Assuming public means unregulated. Publication by the person does not remove the obligations on whoever builds a new collection from it.
- Having no deletion path, so the first time someone asks what you hold about them there is no honest answer and no mechanism to act on it.
When you do need personal data, the minimum discipline
Some research genuinely requires it. If that is you, these are not optional.
- Write the lawful basis down before collecting, with the balancing reasoning, not after
- Keep the narrowest field set that answers the question
- Set a retention period and actually enforce it — indefinite retention is very hard to justify
- Be able to locate and delete an individual's records on request
- Never collect special category data — health, politics, religion, sexuality, biometrics — without specific advice
- Get a real review from someone qualified. This is the part of the course where the stakes justify the fee.
Next: conduct rather than law — rate limits, robots.txt, and being the kind of client nobody has to block.
Lesson 4: rate limits, robots.txt and good conduct