New v5.5.1: balance-first checkout — add funds once, buy instantly. Read more

Web scraping and the GDPR — what European law and case law say

Public does not mean free to use. An overview of the rules that apply to collecting data from websites in the EU — the GDPR, the database right, text and data mining exceptions and contract terms — and of the decisions that shaped them.

On this page

This article is general information about European law as it stood on the publication date. It is not legal advice; check your specific situation with a qualified lawyer.

Customers regularly ask us whether web scraping is "legal in Europe". There is no single answer, because scraping is not one activity governed by one law. Collecting public product prices, indexing job ads and building a profile database of named individuals raise completely different questions. This article maps the main legal regimes that apply in the EU, with the decisions that illustrate them.

1. Personal data: the GDPR applies to public data too

The most common misconception is that information published openly on the web is outside the GDPR. It is not. The Regulation applies to any processing of personal data — information relating to an identified or identifiable person — regardless of where it was found.

If what you collect includes names, profile handles, photos, email addresses, phone numbers or anything that can be linked to a person, you are a controller and need:

  • A lawful basis (Article 6). For scraping, that is almost always legitimate interests under Article 6(1)(f), which requires a documented balancing test: your interest, the necessity of the processing, and the reasonable expectations of the people concerned.
  • Transparency (Article 14). When data is not collected from the person directly, you must inform them, within a month, of who you are, what you collect and why. Article 14(5)(b) relieves you of this only where informing people is impossible or would involve disproportionate effort — and even then you must take other measures, such as publishing the information.
  • Minimisation and retention limits (Article 5). Collect only the fields you need and delete them when you no longer do.
  • Data protection by design (Article 25) and, for large-scale or sensitive processing, a data protection impact assessment (Article 35).

Note that IP addresses can themselves be personal data. In Breyer (C-582/14, 2016), the Court of Justice held that a dynamic IP address can be personal data for a website operator that has legal means to identify the user through the access provider.

Decisions worth knowing

  • Bisnode (Poland, 2019). The Polish supervisory authority fined Bisnode about €220,000 for collecting data on millions of sole traders from public business registers without informing them under Article 14. Sending letters to everyone was costly, but the authority did not accept cost as "disproportionate effort" for a commercial database. The decision was challenged in the courts.
  • Clearview AI (Italy, Greece and France, 2022). Three supervisory authorities each fined Clearview AI €20 million for scraping facial images from the web to build a biometric search engine, finding no valid legal basis, and ordered deletion of the data relating to people in their countries.
  • Meta (Ireland, 2022). The Irish Data Protection Commission fined Meta €265 million after personal data of around 533 million users was scraped through Facebook features and published. The fine concerned the platform's failure to protect data by design and by default — a reminder that website operators are themselves under pressure to prevent scraping of personal data.

Practical reading: collecting non-personal data (prices, stock levels, product attributes, public statistics) raises no GDPR issue. Collecting personal data at scale for profiling or resale is where enforcement has concentrated.

2. The database right

The EU Database Directive (96/9/EC) gives the maker of a database a sui generis right when there has been a substantial investment in obtaining, verifying or presenting its contents. It prohibits extracting or re-using a substantial part of the database, or repeatedly and systematically extracting insubstantial parts.

Two judgments of the Court of Justice frame how it applies to scraping:

  • Ryanair v PR Aviation (C-30/14, 2015). Where a database is protected neither by copyright nor by the sui generis right, the Directive does not prevent the website owner from restricting its use through contract terms. In other words: when the database right does not apply, the website's terms of use may.
  • CV-Online Latvia v Melons (C-762/19, 2021). A specialised search engine that copied and indexed job ads from another site was assessed under the sui generis right. The Court held that extraction and re-use are only prohibited where they adversely affect the maker's investment — that is, where they risk depriving the maker of the revenue that funds the database.

National courts have applied this logic. In France, the Paris Court of Appeal in 2021 found that a property portal had infringed Leboncoin's database right by systematically extracting its real-estate listings.

Practical reading: scraping a handful of fields to compare prices is very different from replicating a competitor's listings database. The closer your output comes to substituting for the source, the higher the risk.

3. Text and data mining exceptions

The 2019 Copyright in the Digital Single Market Directive (2019/790) introduced two exceptions for text and data mining:

  • Article 3 allows research organisations and cultural heritage institutions to mine works they have lawful access to, for scientific research.
  • Article 4 allows anyone to mine lawfully accessible works for any purpose, unless the rightholder has reserved this right "in an appropriate manner, such as machine-readable means" for content made publicly available online.

The opt-out under Article 4 is why robots.txt, metadata tags and terms of use matter more than they used to. Ignoring a clear machine-readable reservation removes the protection of the exception.

4. Contract terms and technical measures

Most websites have terms of use that restrict automated access. Whether those terms bind a scraper depends on national contract law — notably whether the terms were accepted (a logged-in user who clicked "I agree" is in a very different position from an anonymous visitor).

Circumventing technical access controls — logging in with credentials you are not entitled to, bypassing paywalls, defeating authentication — moves the activity into a different category altogether, including potential criminal liability for unauthorised access to information systems under national laws implementing Directive 2013/40/EU.

In Germany, the Federal Court of Justice held in 2014 (Flugvermittlung im Internet, I ZR 224/12) that screen scraping of a publicly accessible airline site was not in itself an unfair commercial practice, as long as no technical protection measures were circumvented. That reasoning is often cited, but it is a competition-law ruling and does not displace the GDPR or the database right.

5. A short checklist

Before starting a collection project in Europe, write down the answers to these questions:

  1. Does the data include anything relating to identifiable people? If yes: lawful basis, Article 14 information, minimisation, retention, possibly a DPIA.
  2. Is the source a database with substantial investment behind it, and will you extract a substantial part of it or substitute for it?
  3. Has the site reserved text and data mining rights through machine-readable means?
  4. Have you accepted terms of use (for example by creating an account) that prohibit automated access?
  5. Would you need to bypass any access control? If yes, stop.
  6. Is your request rate respectful of the target's infrastructure? Overloading a site can create liability independently of everything above.

Where proxies fit

A proxy changes where requests come from; it does not change what you are allowed to collect. Our acceptable use policy prohibits using the network to break the law, to access systems without authorisation or to collect special categories of personal data, and we act on substantiated abuse reports. If your project involves personal data, involve your data protection officer or counsel before you start — not after the first complaint.

Related articles

All articles

Guides

Buying proxies with crypto in 2026 — what actually matters

Who accepts cryptocurrency for proxies in 2026, what the payment really changes, and how to pick between dedicated mobile modems, rotating residential pools and static ISP IPs when you pay in BTC, XMR or USDT.

5 min read