Slow down first
The cheapest fix is speed. A crawler that asks for fifty pages a second from one site gets noticed whatever addresses it uses. Add a delay, vary it a little, and run fewer requests at once per site.
Look like a browser
- Send a current browser
User-Agent, and the headers that browser sends with it. - Keep cookies between requests that belong together.
- Follow redirects, and accept compressed responses.
Spread the requests
One address making thousands of requests stands out. With a rotating username, each connection leaves from a different address, so no single one is busy. Put the country in the username when the site shows different content by country.
Rotating username
USERNAMEKeep sessions where they belong
Some flows only work from one address: a login, a search followed by its result pages, anything with a token. Use a sticky username for those and a rotating one for the rest.
Choose the right kind of address
If a site answers a datacenter address with a block page and a home address with content, the address type is the problem, not your code. See residential vs datacenter vs mobile.
Be a good guest
- Read the site's terms and its
robots.txt. - Collect public pages, not what sits behind someone else's login.
- Do not crawl at a rate that would slow the site for its real visitors.
- Personal data has rules of its own wherever you are. Know them before you collect it.
Pick your tool
- Python RequestsRequests is the standard HTTP library for Python. A proxy is a dictionary passed with the request, and the username and password go inside the proxy address.
- HTTPXHTTPX is a modern HTTP client for Python with both normal and async interfaces. The proxy is set on the client.
- aiohttpaiohttp is the common async HTTP client for Python. The proxy is given per request, which makes it easy to send different requests through different sessions.
- ScrapyScrapy is the main crawling framework for Python. Its built-in proxy middleware reads a
proxyvalue from each request, including a username and password in the address. - Scrapy Playwrightscrapy-playwright lets Scrapy load pages in a real browser. The browser does not use Scrapy's proxy setting; it needs the proxy in its own launch options.
- Scrapy SplashSplash is a lightweight browser service used with Scrapy to render JavaScript. The proxy is passed to Splash with each request.
- Beautiful SoupBeautiful Soup reads HTML; it does not download it. The proxy belongs to whatever fetches the page, which is usually Requests.
- MechanicalSoupMechanicalSoup fills in forms and follows links like a simple browser, on top of Requests and Beautiful Soup. The proxy is set on its session.
- curl_cfficurl_cffi is a Python client that can copy the TLS fingerprint of a real browser. It takes proxies the same way Requests does.
- cloudscrapercloudscraper is a Python library built on Requests. Its scraper object takes the same
proxiesargument. - SeleniumSelenium drives a real browser. Chrome accepts a proxy address on the command line but ignores a username and password there, so a proxy with a login needs one extra piece.
- SeleniumBaseSeleniumBase is a Python framework on top of Selenium. Unlike plain Selenium, it accepts a proxy with a username and password directly.
- undetected-chromedriverundetected-chromedriver is a patched ChromeDriver for Python. Like Chrome itself, it takes a proxy address but not a username and password.
- PlaywrightPlaywright drives Chromium, Firefox and WebKit from Python, Node.js, Java or .NET. It takes a proxy with a username and password as launch options, with no extra tools.
- PuppeteerPuppeteer drives Chrome from Node.js. The proxy address goes in a launch argument, and the username and password are given to the page.
- CrawleeCrawlee is a crawling library for Node.js with HTTP crawlers and browser crawlers behind one interface. A
ProxyConfigurationholds your proxy addresses and hands them out. - ApifyApify runs scrapers, called Actors, in the cloud. Most Actors have a proxy setting, and it accepts your own proxy addresses as well as Apify's.
- AxiosAxios is the most used HTTP client for Node.js. For https sites through a proxy with a login, the reliable way is a proxy agent instead of Axios's own
proxyoption. - CheerioCheerio parses HTML in Node.js with a jQuery-like interface. It does not download pages, so the proxy belongs to the client that does.
- Node.js fetchThe
fetchbuilt into Node.js does not read a proxy from its options. It uses the undici library underneath, and undici'sProxyAgentis how you give it one. - GotGot is an HTTP client for Node.js. It takes a proxy through an agent, and its documentation points to the
hpagentpackage for that. - curlcurl is the quickest way to test a proxy. One option,
-x, takes the whole proxy address. - WgetWget downloads files and whole sites from the command line. It reads its proxy from settings or environment variables.
- HTTrackHTTrack copies a website to your disk for offline reading. Its proxy option takes the login together with the address.
- KatanaKatana is a fast crawler from ProjectDiscovery, used to map the pages and endpoints of a site you are allowed to test.
- Go net/httpGo's standard HTTP client takes a proxy on its transport. The username and password are read from the proxy address.
- CollyColly is a scraping framework for Go. A collector takes one proxy with
SetProxy, or a list to rotate through. - Rust reqwestreqwest is the usual HTTP client for Rust. A proxy is added to the client builder, with the login set by
basic_auth. - JsoupJsoup fetches and parses HTML in Java. It takes the proxy host and port on the connection; the login goes through Java's
Authenticator. - OkHttpOkHttp is the standard HTTP client for Java and Android. It takes the proxy and the proxy login as two separate settings on the client.
- C# HttpClientIn .NET, the proxy is set on the handler behind
HttpClient, with the login as aNetworkCredential. - PHP cURLPHP's cURL functions take the proxy address and the proxy login as two options.
- GuzzleGuzzle is the common HTTP client for PHP. The proxy is one request option, with the login inside the address.
- RubyRuby's standard library takes the proxy as extra arguments when the connection is created: host, port, username, password.
- PowerShellPowerShell's
Invoke-WebRequestandInvoke-RestMethodtake a proxy address and a credential object. - Screaming FrogScreaming Frog SEO Spider crawls a site the way a search engine would. It can send the crawl through one proxy, with a username and password.
- ScrapeBoxScrapeBox is a Windows tool for harvesting search results and checking lists of addresses in bulk. It keeps its proxies as a list of lines.
- OctoparseOctoparse is a point-and-click scraper. It accepts your own proxies for tasks that run on your computer, but only as an IP address and port, without a username and password.
- ParseHubParseHub is a visual scraper that runs projects in its cloud. Paid plans can use your own proxies, once ParseHub has switched the option on for your account.
- WebHarvyWebHarvy is a point-and-click scraper for Windows. Its settings take one proxy or a list to rotate through.
- Web ScraperWeb Scraper is a browser extension that scrapes with a point-and-click sitemap. In the browser it has no proxy setting: it uses the browser's connection.
Questions
Will proxies alone stop blocks?
No. They remove one signal, the address. Speed, headers and cookies are the others.
How many addresses do I need?
With a rotating username you do not count them: each connection gets one from the pool.