Web scraping: proxy guides
Every scraping library takes a proxy a little differently, and most of the trouble is in the details: where the password goes, whether SOCKS5 works, how to keep or change the IP. Each guide has an example you can copy and run.
Python Requests
Requests is the standard HTTP library for Python. A proxy is a dictionary passed with the request, and the username and password go inside the proxy address.
HTTPX
HTTPX is a modern HTTP client for Python with both normal and async interfaces. The proxy is set on the client.
aiohttp
aiohttp is the common async HTTP client for Python. The proxy is given per request, which makes it easy to send different requests through different sessions.
Scrapy
Scrapy is the main crawling framework for Python. Its built-in proxy middleware reads a
proxyvalue from each request, including a username and password in the address.Scrapy Playwright
scrapy-playwright lets Scrapy load pages in a real browser. The browser does not use Scrapy's proxy setting; it needs the proxy in its own launch options.
Scrapy Splash
Splash is a lightweight browser service used with Scrapy to render JavaScript. The proxy is passed to Splash with each request.
Beautiful Soup
Beautiful Soup reads HTML; it does not download it. The proxy belongs to whatever fetches the page, which is usually Requests.
MechanicalSoup
MechanicalSoup fills in forms and follows links like a simple browser, on top of Requests and Beautiful Soup. The proxy is set on its session.
curl_cffi
curl_cffi is a Python client that can copy the TLS fingerprint of a real browser. It takes proxies the same way Requests does.
cloudscraper
cloudscraper is a Python library built on Requests. Its scraper object takes the same
proxiesargument.Selenium
Selenium drives a real browser. Chrome accepts a proxy address on the command line but ignores a username and password there, so a proxy with a login needs one extra piece.
SeleniumBase
SeleniumBase is a Python framework on top of Selenium. Unlike plain Selenium, it accepts a proxy with a username and password directly.
undetected-chromedriver
undetected-chromedriver is a patched ChromeDriver for Python. Like Chrome itself, it takes a proxy address but not a username and password.
Playwright
Playwright drives Chromium, Firefox and WebKit from Python, Node.js, Java or .NET. It takes a proxy with a username and password as launch options, with no extra tools.
Puppeteer
Puppeteer drives Chrome from Node.js. The proxy address goes in a launch argument, and the username and password are given to the page.
Crawlee
Crawlee is a crawling library for Node.js with HTTP crawlers and browser crawlers behind one interface. A
ProxyConfigurationholds your proxy addresses and hands them out.Apify
Apify runs scrapers, called Actors, in the cloud. Most Actors have a proxy setting, and it accepts your own proxy addresses as well as Apify's.
Axios
Axios is the most used HTTP client for Node.js. For https sites through a proxy with a login, the reliable way is a proxy agent instead of Axios's own
proxyoption.Cheerio
Cheerio parses HTML in Node.js with a jQuery-like interface. It does not download pages, so the proxy belongs to the client that does.
Node.js fetch
The
fetchbuilt into Node.js does not read a proxy from its options. It uses the undici library underneath, and undici'sProxyAgentis how you give it one.Got
Got is an HTTP client for Node.js. It takes a proxy through an agent, and its documentation points to the
hpagentpackage for that.curl
curl is the quickest way to test a proxy. One option,
-x, takes the whole proxy address.Wget
Wget downloads files and whole sites from the command line. It reads its proxy from settings or environment variables.
HTTrack
HTTrack copies a website to your disk for offline reading. Its proxy option takes the login together with the address.
Katana
Katana is a fast crawler from ProjectDiscovery, used to map the pages and endpoints of a site you are allowed to test.
Go net/http
Go's standard HTTP client takes a proxy on its transport. The username and password are read from the proxy address.
Colly
Colly is a scraping framework for Go. A collector takes one proxy with
SetProxy, or a list to rotate through.Rust reqwest
reqwest is the usual HTTP client for Rust. A proxy is added to the client builder, with the login set by
basic_auth.Jsoup
Jsoup fetches and parses HTML in Java. It takes the proxy host and port on the connection; the login goes through Java's
Authenticator.OkHttp
OkHttp is the standard HTTP client for Java and Android. It takes the proxy and the proxy login as two separate settings on the client.
C# HttpClient
In .NET, the proxy is set on the handler behind
HttpClient, with the login as aNetworkCredential.PHP cURL
PHP's cURL functions take the proxy address and the proxy login as two options.
Guzzle
Guzzle is the common HTTP client for PHP. The proxy is one request option, with the login inside the address.
Ruby
Ruby's standard library takes the proxy as extra arguments when the connection is created: host, port, username, password.
PowerShell
PowerShell's
Invoke-WebRequestandInvoke-RestMethodtake a proxy address and a credential object.Screaming Frog
Screaming Frog SEO Spider crawls a site the way a search engine would. It can send the crawl through one proxy, with a username and password.
ScrapeBox
ScrapeBox is a Windows tool for harvesting search results and checking lists of addresses in bulk. It keeps its proxies as a list of lines.
Octoparse
Octoparse is a point-and-click scraper. It accepts your own proxies for tasks that run on your computer, but only as an IP address and port, without a username and password.
ParseHub
ParseHub is a visual scraper that runs projects in its cloud. Paid plans can use your own proxies, once ParseHub has switched the option on for your account.
WebHarvy
WebHarvy is a point-and-click scraper for Windows. Its settings take one proxy or a list to rotate through.
Web Scraper
Web Scraper is a browser extension that scrapes with a point-and-click sitemap. In the browser it has no proxy setting: it uses the browser's connection.
Web Scraping Without Getting Blocked
The habits that keep a scraper running: sensible speed, real browser headers, rotating IP addresses, sessions where they belong, and respect for the site.
Web Scraping Tools and Libraries: Proxy Setup Index
An index of proxy setup guides for web scraping: Python, Node.js, Go, Java, PHP, Ruby and command-line tools, plus desktop and no-code scrapers.