PiratProxies

Web scraping

Web Scraping Tools and Libraries: Proxy Setup Index

Every scraping tool takes a proxy in its own way. This index links to a working example for each one.

All the guides

  • Python RequestsRequests is the standard HTTP library for Python. A proxy is a dictionary passed with the request, and the username and password go inside the proxy address.
  • HTTPXHTTPX is a modern HTTP client for Python with both normal and async interfaces. The proxy is set on the client.
  • aiohttpaiohttp is the common async HTTP client for Python. The proxy is given per request, which makes it easy to send different requests through different sessions.
  • ScrapyScrapy is the main crawling framework for Python. Its built-in proxy middleware reads a proxy value from each request, including a username and password in the address.
  • Scrapy Playwrightscrapy-playwright lets Scrapy load pages in a real browser. The browser does not use Scrapy's proxy setting; it needs the proxy in its own launch options.
  • Scrapy SplashSplash is a lightweight browser service used with Scrapy to render JavaScript. The proxy is passed to Splash with each request.
  • Beautiful SoupBeautiful Soup reads HTML; it does not download it. The proxy belongs to whatever fetches the page, which is usually Requests.
  • MechanicalSoupMechanicalSoup fills in forms and follows links like a simple browser, on top of Requests and Beautiful Soup. The proxy is set on its session.
  • curl_cfficurl_cffi is a Python client that can copy the TLS fingerprint of a real browser. It takes proxies the same way Requests does.
  • cloudscrapercloudscraper is a Python library built on Requests. Its scraper object takes the same proxies argument.
  • SeleniumSelenium drives a real browser. Chrome accepts a proxy address on the command line but ignores a username and password there, so a proxy with a login needs one extra piece.
  • SeleniumBaseSeleniumBase is a Python framework on top of Selenium. Unlike plain Selenium, it accepts a proxy with a username and password directly.
  • undetected-chromedriverundetected-chromedriver is a patched ChromeDriver for Python. Like Chrome itself, it takes a proxy address but not a username and password.
  • PlaywrightPlaywright drives Chromium, Firefox and WebKit from Python, Node.js, Java or .NET. It takes a proxy with a username and password as launch options, with no extra tools.
  • PuppeteerPuppeteer drives Chrome from Node.js. The proxy address goes in a launch argument, and the username and password are given to the page.
  • CrawleeCrawlee is a crawling library for Node.js with HTTP crawlers and browser crawlers behind one interface. A ProxyConfiguration holds your proxy addresses and hands them out.
  • ApifyApify runs scrapers, called Actors, in the cloud. Most Actors have a proxy setting, and it accepts your own proxy addresses as well as Apify's.
  • AxiosAxios is the most used HTTP client for Node.js. For https sites through a proxy with a login, the reliable way is a proxy agent instead of Axios's own proxy option.
  • CheerioCheerio parses HTML in Node.js with a jQuery-like interface. It does not download pages, so the proxy belongs to the client that does.
  • Node.js fetchThe fetch built into Node.js does not read a proxy from its options. It uses the undici library underneath, and undici's ProxyAgent is how you give it one.
  • GotGot is an HTTP client for Node.js. It takes a proxy through an agent, and its documentation points to the hpagent package for that.
  • curlcurl is the quickest way to test a proxy. One option, -x, takes the whole proxy address.
  • WgetWget downloads files and whole sites from the command line. It reads its proxy from settings or environment variables.
  • HTTrackHTTrack copies a website to your disk for offline reading. Its proxy option takes the login together with the address.
  • KatanaKatana is a fast crawler from ProjectDiscovery, used to map the pages and endpoints of a site you are allowed to test.
  • Go net/httpGo's standard HTTP client takes a proxy on its transport. The username and password are read from the proxy address.
  • CollyColly is a scraping framework for Go. A collector takes one proxy with SetProxy, or a list to rotate through.
  • Rust reqwestreqwest is the usual HTTP client for Rust. A proxy is added to the client builder, with the login set by basic_auth.
  • JsoupJsoup fetches and parses HTML in Java. It takes the proxy host and port on the connection; the login goes through Java's Authenticator.
  • OkHttpOkHttp is the standard HTTP client for Java and Android. It takes the proxy and the proxy login as two separate settings on the client.
  • C# HttpClientIn .NET, the proxy is set on the handler behind HttpClient, with the login as a NetworkCredential.
  • PHP cURLPHP's cURL functions take the proxy address and the proxy login as two options.
  • GuzzleGuzzle is the common HTTP client for PHP. The proxy is one request option, with the login inside the address.
  • RubyRuby's standard library takes the proxy as extra arguments when the connection is created: host, port, username, password.
  • PowerShellPowerShell's Invoke-WebRequest and Invoke-RestMethod take a proxy address and a credential object.
  • Screaming FrogScreaming Frog SEO Spider crawls a site the way a search engine would. It can send the crawl through one proxy, with a username and password.
  • ScrapeBoxScrapeBox is a Windows tool for harvesting search results and checking lists of addresses in bulk. It keeps its proxies as a list of lines.
  • OctoparseOctoparse is a point-and-click scraper. It accepts your own proxies for tasks that run on your computer, but only as an IP address and port, without a username and password.
  • ParseHubParseHub is a visual scraper that runs projects in its cloud. Paid plans can use your own proxies, once ParseHub has switched the option on for your account.
  • WebHarvyWebHarvy is a point-and-click scraper for Windows. Its settings take one proxy or a list to rotate through.
  • Web ScraperWeb Scraper is a browser extension that scrapes with a point-and-click sitemap. In the browser it has no proxy setting: it uses the browser's connection.

Three patterns cover nearly all of them

  • Libraries take the whole proxy as one address, login included.
  • Browsers take the address and the login separately. Chrome on its own takes no login at all.
  • Desktop tools take a line or a form, and a few take only an IP address and a port.

The address most libraries want

http://USERNAME:PASSWORD@resi.piratproxies.com:8080

Before you scale up

Read web scraping without getting blocked, and estimate your traffic with how much proxy traffic do you need.