PiratProxies

Web scraping

Web Scraping Without Getting Blocked

Sites block scrapers that behave unlike visitors: too fast, too regular, all from one address. Most blocks go away when the scraper is simply better behaved.

Slow down first

The cheapest fix is speed. A crawler that asks for fifty pages a second from one site gets noticed whatever addresses it uses. Add a delay, vary it a little, and run fewer requests at once per site.

Look like a browser

  • Send a current browser User-Agent, and the headers that browser sends with it.
  • Keep cookies between requests that belong together.
  • Follow redirects, and accept compressed responses.

Spread the requests

One address making thousands of requests stands out. With a rotating username, each connection leaves from a different address, so no single one is busy. Put the country in the username when the site shows different content by country.

Rotating username

USERNAME

Keep sessions where they belong

Some flows only work from one address: a login, a search followed by its result pages, anything with a token. Use a sticky username for those and a rotating one for the rest.

Choose the right kind of address

If a site answers a datacenter address with a block page and a home address with content, the address type is the problem, not your code. See residential vs datacenter vs mobile.

Be a good guest

  • Read the site's terms and its robots.txt.
  • Collect public pages, not what sits behind someone else's login.
  • Do not crawl at a rate that would slow the site for its real visitors.
  • Personal data has rules of its own wherever you are. Know them before you collect it.

Pick your tool

  • Python RequestsRequests is the standard HTTP library for Python. A proxy is a dictionary passed with the request, and the username and password go inside the proxy address.
  • HTTPXHTTPX is a modern HTTP client for Python with both normal and async interfaces. The proxy is set on the client.
  • aiohttpaiohttp is the common async HTTP client for Python. The proxy is given per request, which makes it easy to send different requests through different sessions.
  • ScrapyScrapy is the main crawling framework for Python. Its built-in proxy middleware reads a proxy value from each request, including a username and password in the address.
  • Scrapy Playwrightscrapy-playwright lets Scrapy load pages in a real browser. The browser does not use Scrapy's proxy setting; it needs the proxy in its own launch options.
  • Scrapy SplashSplash is a lightweight browser service used with Scrapy to render JavaScript. The proxy is passed to Splash with each request.
  • Beautiful SoupBeautiful Soup reads HTML; it does not download it. The proxy belongs to whatever fetches the page, which is usually Requests.
  • MechanicalSoupMechanicalSoup fills in forms and follows links like a simple browser, on top of Requests and Beautiful Soup. The proxy is set on its session.
  • curl_cfficurl_cffi is a Python client that can copy the TLS fingerprint of a real browser. It takes proxies the same way Requests does.
  • cloudscrapercloudscraper is a Python library built on Requests. Its scraper object takes the same proxies argument.
  • SeleniumSelenium drives a real browser. Chrome accepts a proxy address on the command line but ignores a username and password there, so a proxy with a login needs one extra piece.
  • SeleniumBaseSeleniumBase is a Python framework on top of Selenium. Unlike plain Selenium, it accepts a proxy with a username and password directly.
  • undetected-chromedriverundetected-chromedriver is a patched ChromeDriver for Python. Like Chrome itself, it takes a proxy address but not a username and password.
  • PlaywrightPlaywright drives Chromium, Firefox and WebKit from Python, Node.js, Java or .NET. It takes a proxy with a username and password as launch options, with no extra tools.
  • PuppeteerPuppeteer drives Chrome from Node.js. The proxy address goes in a launch argument, and the username and password are given to the page.
  • CrawleeCrawlee is a crawling library for Node.js with HTTP crawlers and browser crawlers behind one interface. A ProxyConfiguration holds your proxy addresses and hands them out.
  • ApifyApify runs scrapers, called Actors, in the cloud. Most Actors have a proxy setting, and it accepts your own proxy addresses as well as Apify's.
  • AxiosAxios is the most used HTTP client for Node.js. For https sites through a proxy with a login, the reliable way is a proxy agent instead of Axios's own proxy option.
  • CheerioCheerio parses HTML in Node.js with a jQuery-like interface. It does not download pages, so the proxy belongs to the client that does.
  • Node.js fetchThe fetch built into Node.js does not read a proxy from its options. It uses the undici library underneath, and undici's ProxyAgent is how you give it one.
  • GotGot is an HTTP client for Node.js. It takes a proxy through an agent, and its documentation points to the hpagent package for that.
  • curlcurl is the quickest way to test a proxy. One option, -x, takes the whole proxy address.
  • WgetWget downloads files and whole sites from the command line. It reads its proxy from settings or environment variables.
  • HTTrackHTTrack copies a website to your disk for offline reading. Its proxy option takes the login together with the address.
  • KatanaKatana is a fast crawler from ProjectDiscovery, used to map the pages and endpoints of a site you are allowed to test.
  • Go net/httpGo's standard HTTP client takes a proxy on its transport. The username and password are read from the proxy address.
  • CollyColly is a scraping framework for Go. A collector takes one proxy with SetProxy, or a list to rotate through.
  • Rust reqwestreqwest is the usual HTTP client for Rust. A proxy is added to the client builder, with the login set by basic_auth.
  • JsoupJsoup fetches and parses HTML in Java. It takes the proxy host and port on the connection; the login goes through Java's Authenticator.
  • OkHttpOkHttp is the standard HTTP client for Java and Android. It takes the proxy and the proxy login as two separate settings on the client.
  • C# HttpClientIn .NET, the proxy is set on the handler behind HttpClient, with the login as a NetworkCredential.
  • PHP cURLPHP's cURL functions take the proxy address and the proxy login as two options.
  • GuzzleGuzzle is the common HTTP client for PHP. The proxy is one request option, with the login inside the address.
  • RubyRuby's standard library takes the proxy as extra arguments when the connection is created: host, port, username, password.
  • PowerShellPowerShell's Invoke-WebRequest and Invoke-RestMethod take a proxy address and a credential object.
  • Screaming FrogScreaming Frog SEO Spider crawls a site the way a search engine would. It can send the crawl through one proxy, with a username and password.
  • ScrapeBoxScrapeBox is a Windows tool for harvesting search results and checking lists of addresses in bulk. It keeps its proxies as a list of lines.
  • OctoparseOctoparse is a point-and-click scraper. It accepts your own proxies for tasks that run on your computer, but only as an IP address and port, without a username and password.
  • ParseHubParseHub is a visual scraper that runs projects in its cloud. Paid plans can use your own proxies, once ParseHub has switched the option on for your account.
  • WebHarvyWebHarvy is a point-and-click scraper for Windows. Its settings take one proxy or a list to rotate through.
  • Web ScraperWeb Scraper is a browser extension that scrapes with a point-and-click sitemap. In the browser it has no proxy setting: it uses the browser's connection.

Questions

Will proxies alone stop blocks?

No. They remove one signal, the address. Speed, headers and cookies are the others.

How many addresses do I need?

With a rotating username you do not count them: each connection gets one from the pool.