The legal and ethical side, without the hand-waving· Lesson 3 of 5

Personal data, and why product scraping quietly becomes it

Public does not mean unregulated. The categories that catch people out, and the simplest way to stay clear of the whole problem.

  • 11 min read
  • No account needed

The most common surprise in this whole area: publicly available personal data is still personal data. The fact that someone posted their name on a public page does not take it outside data protection law, and the intuition that "public means fair game" is simply wrong under the GDPR and its equivalents.

This matters for product scraping specifically, because personal data arrives in product datasets by accident far more often than by design.

Fields that look commercial and are not

FieldWhy it is personal dataSafer option
Marketplace seller nameVery often an individual or a sole traderKeep a hashed seller id, or drop it
Review text and authorWritten by an identifiable person, and sometimes revealingKeep the rating count and average; drop the text
Q&A on a product pageSame — identifiable individualsDrop entirely
Seller contact detailsDirectly identifyingNever collect
"Sold by" on a retailer listingCorporate when it is a company, personal when it is notCheck which, or drop
Price, stock, SKU, titleNot personal dataCollect freely

What the rules actually ask of you

Simplified considerably, and for the EU and UK regime specifically, there are four obligations that bite on collected data.

You need a lawful basis. For commercial research the usual candidate is legitimate interests, which requires you to balance your interest against the individual's rights and — importantly — to document that you did.

You have to tell people. There is a transparency obligation when you collect data about someone from a source other than them. There is a carve-out where notifying everyone would be disproportionate, but it is not automatic and it depends on you having thought about it.

You must minimise. Collect what the purpose needs and no more. This is also, conveniently, the rule that makes most of the problem disappear.

And you have to be able to respond. Individuals can ask what you hold and ask you to delete it, which in practice means you need to be able to find a named person in your dataset at all.

The simplest answer is usually the right one

Do not collect it.

A price monitoring system does not need seller names to work. It does not need review text. Dropping those fields at the point of extraction — not filtering them out later, but never writing them down — moves the entire project out of the personal data regime and removes a category of risk, a category of obligation and a category of conversation.

Where a seller identity genuinely matters for the analysis, a stable hash preserves the ability to say "this is the same seller as last week" without holding anybody's name. That is sufficient for almost every commercial question people actually ask of this data.

It is also simply less to defend. A dataset that provably contains no personal data is the shortest possible answer to a security review, and security reviews are where these projects most often stall.

Worked example: auditing a product feed, column by column

A marketplace price feed looks entirely commercial until you list the columns and ask one question of each: could this, alone or combined with the rest, identify a living person? Here is the same feed before and after that audit.

ColumnIdentifies a person?Needed for repricing?Decision
product_titleNoYes, for matchingKeep
price, currency, in_stockNoYesKeep
gtin, mpnNoYes, for matchingKeep
seller_nameOften yes on a marketplace, where a large share of sellers trade under their own nameNoDrop
seller_addressYes, frequently a home addressNoDrop
review_author, review_textYes, and review text is a direct opinion attached to a named personNoDrop
q_and_a_usernameYes, pseudonymous but linkableNoDrop

What usually goes wrong

Nobody sets out to build a personal dataset. It assembles itself from columns that each looked harmless.

  • Scraping the whole page because it was easier than selecting fields, then discovering a year later that the warehouse holds seller names and review text nobody ever used.
  • Treating pseudonyms as anonymous. A stable username plus a purchase history plus a town is identifying in practice, whatever it looks like in isolation.
  • Keeping seller_name "for debugging" and never removing it, so a temporary convenience becomes a permanent category change.
  • Assuming public means unregulated. Publication by the person does not remove the obligations on whoever builds a new collection from it.
  • Having no deletion path, so the first time someone asks what you hold about them there is no honest answer and no mechanism to act on it.

When you do need personal data, the minimum discipline

Some research genuinely requires it. If that is you, these are not optional.

  • Write the lawful basis down before collecting, with the balancing reasoning, not after
  • Keep the narrowest field set that answers the question
  • Set a retention period and actually enforce it — indefinite retention is very hard to justify
  • Be able to locate and delete an individual's records on request
  • Never collect special category data — health, politics, religion, sexuality, biometrics — without specific advice
  • Get a real review from someone qualified. This is the part of the course where the stakes justify the fee.

Next: conduct rather than law — rate limits, robots.txt, and being the kind of client nobody has to block.

Lesson 4: rate limits, robots.txt and good conduct

Rather have the feed than build it?

Hand over the list of competitors and get the rows back. Pay per request, no subscription.