In This Article

Back to blog

Playwright Web Scraping: The Complete 2026 Guide

Tutorials

Vilius Dumcius

Last updated - ‐ 13 min read

Key Takeaways

  • Playwright is a multi-language scraping framework that runs faster than Selenium because it uses a persistent WebSocket connection instead of HTTP overhead.

  • It natively auto-waits for dynamic elements and allows you to intercept background network traffic.

  • To scrape websites at scale, open isolated browser contexts for concurrent tasks, use exponential backoff to handle rate limits, and get rotating proxies.

Web scraping with Playwright is becoming increasingly common because the framework provides many features that make data extraction much easier. It’s one of the few tools that are available in multiple programming languages, such as Python and JavaScript (Node.js).

Multi-language support is not the only reason to pick Playwright for web scraping. Numerous advanced features make Playwright web scraping a breeze by letting you customize browser sessions for specific websites.

What Is Playwright?

Playwright is an open-source browser automation framework that has been in active development since 2020. Managed by Microsoft, Playwright is designed to provide an easy way to automate browsers based on Chromium, WebKit, and Firefox.

One of its strengths has been multi-language, cross-browser support, since Playwright is available in many popular programming languages. As such, it quickly became popular among developers and has been used for many different use cases.

Two primary applications for Playwright have been web scraping and website testing. Both use the framework’s powerful features, allowing developers to automate many processes with ease.

Playwright web scraping is particularly popular because the framework can automate multiple browsers while offering extensive customization for each. Additionally, Playwright is fairly well-optimized, making it a great choice for web scraping.

Playwright vs. Puppeteer vs. Selenium

The framework you choose depends heavily on your project stack. Here’s how the main options compare:

Feature Playwright Puppeteer Selenium
Language support Multi-language (JS, Python, Java, .NET) JavaScript / TypeScript Multi-language (Java, Python, C#, JS, Ruby)
Browser support Chromium, Firefox, WebKit Chrome, Firefox Chrome, Firefox, Edge, Safari
Auto-waiting Built-in Via Rotator API (experimental) Manual
Speed Very fast (WebSocket) Very fast (CDP/WebDriver BiDi) Slower (WebDriver HTTP overhead)
Community support Rapidly growing Mature Highly mature

Playwright downloads its own browser binaries for Chromium, Firefox, and WebKit, so you don’t need to manage separate drivers. Because it natively auto-waits for dynamic content to load, you write less boilerplate code compared to Selenium.

Selenium has the longest track record and supports the most programming languages, but Playwright has become a go-to choice for many developers working on modern web scraping and testing projects.

Ready to get started?
Register now

Is Playwright Web Scraping Worth It?

Playwright web scraping is highly effective, customizable, and applicable to most scenarios. Many browser automation or HTTP request libraries lack certain features such as asynchronous programming or auto-waiting until specific elements load.

All of these are available in Playwright. Web scraping solution developers can make great use of these features, making Playwright a better option than most other browser automation libraries .

There are a few minor drawbacks to web scraping with Playwright. One of the major ones is the learning curve, as many of the advanced features can be a bit harder to understand.

Additionally, Playwright is a bit heavier because it provides both headful and headless browser drivers for multiple browsers. While disk space is rarely an issue, it can be a bit frustrating in some scenarios.

Finally, while the Playwright Python library has matured quickly since its 2020 release, Selenium, the most established browser automation library, still has a larger archive of tutorials and public code, so niche problems can take longer to research.

Web Scraping With Playwright

While Playwright code can be written in JavaScript/TypeScript, Python, Java, and .NET (C#), we’ll be using JavaScript through Node.js and Python to create a web scraping tool with each one.

You may need to rewrite some of the syntax if you use a different language, but the underlying logic will be identical.

Setting Up Your Environment

Node.js

To get started with Playwright web scraping, we’ll first need the language package and an IDE. You can find and install the Node.js package from the official website. In terms of IDE choices, there are plenty of options, but we’ll be using IntelliJ IDEA .

Now, once both are installed, you’ll need to install Playwright itself. Open up your IDE, start a new project, boot up the Terminal, and type in the following two commands one after the other.

The first adds the Playwright library to your project, and the second downloads the browser binaries for Chromium, Firefox, and WebKit:

npm install playwright
npx playwright install

Python

For Python, you’ll need, again, the language package and an IDE. You can download Python from the official website, and there are also plenty of options for the IDE. PyCharm’s free tier (JetBrains merged the former Community Edition into a single PyCharm product in 2025) covers everything you need for this tutorial.

Once both are installed, open up your IDE, start a new project, boot up the Terminal, and install Playwright along with its browser binaries:

pip install playwright
playwright install

Locating Elements

Playwright provides many ways to find the elements you need, each of which has its own drawbacks and benefits. Over time, you’ll learn to pick the right method just from experience, but here are a few popular options:

  • CSS selectors. One of the most common ways to find elements. You can find them based on CSS classes, IDs, attributes, and many other features.
  • XPath. Another popular way to find elements. XPath lets you move through the HTML DOM structure and pick out the elements you need.
  • Text-based locations. Playwright gives you the ability to find elements through strings. It’s highly useful when you want to extract information from buttons or other short, interactable elements.

Here are a few examples of how you can locate elements using these methods.

Example #1: CSS Selectors

const titles = await page.$$eval('.post-title', elements =>
    elements.map(el => el.textContent.trim())
  );

 console.log('Article Titles:', titles);
        titles = page.query_selector_all('.post-title')
        titles_text = [title.text_content().strip() for title in titles]

        print('Article Titles:', titles_text)

In a web scraping context, Playwright would find all CSS elements that match the “.post-title” selector. Then, all of the elements found according to the selector would be stored in a different variable and trimmed. All data would be output into the console.

Example #2: XPath

  const prices = await page.$$eval(
    '//div[@class="product"]//span[@itemprop="price"]',
    elements => elements.map(el => el.textContent.trim())
  );

  console.log('Product Prices:', prices);
        prices = page.query_selector_all('//div[@class="product"]//span[@itemprop="price"]')
        prices_text = [price.text_content().strip() for price in prices]

        print('Product Prices:', prices_text)

Instead of looking at the selectors, we now look at all the price elements that are stored within product divs. The rest of the HTML code follows the same logic.

Example #3: Text-Based Locations

  const buttonText = await page.locator('button:has-text("Read More")').innerText();

  console.log('Button Text:', buttonText);
        button = page.locator('button:has-text("Read More")')
        button_text = button.inner_text()

        print('Button Text:', button_text)

Text-based extraction is the most intuitive. All you have to do is find the button in your browser, verify that the string exists in the HTML file, and input it into your code. As long as the string matches, Playwright will be able to find it.

Scraping Text

To avoid cluttering the article with too much code, we’ll be using a classic method for text scraping – CSS selectors. But before we can start scraping text, we first need to create a Playwright browser instance within Node.js :

const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch();
  const page = await browser.newPage();
  await page.goto('https://iproyal.com/blog/');

    await browser.close();
})();

Our code creates a Chromium instance through Playwright, then creates a new browser page and goes to the assigned URL, in this case, the IPRoyal blog. Since it’s running a headless browser, all you’ll get is an exit code.

Now, we’ll use CSS selectors to find the page titles and output them into the console:

const { chromium } = require('playwright');

(async () => {
    const browser = await chromium.launch();
    const page = await browser.newPage();
    await page.goto('https://iproyal.com/blog/');

      const titles = await page.$$eval('h2.tp-headline-s', elements =>
    elements.map(el => el.textContent.trim())
  );

  console.log('Article Titles:', titles);
    await browser.close();
})();

We’re using the previous “$$eval” method to extract titles. You have to find the correct CSS selector, however.

To find the CSS selector, you can visit the web page and use the Inspect function. All you need to do is right-click on the element and select “Inspect”. Find the text you need in the code, right-click again, hover over the “Copy” function, and click on “Copy selector”.

Sometimes, the CSS selector will be way too precise or complicated. You can then try to look at the code to find a simpler option. For example, on the IPRoyal website, selectors will be highly specific and complicated.

But all titles are stored in the selector “h2.tp-headline-s”. Filling that in, as we did in our code, will perform all of the functions you need.

Here’s identical code for a Python implementation:

from playwright.sync_api import sync_playwright

def scrape_titles():
    with sync_playwright() as p:
        browser = p.chromium.launch()
        page = browser.new_page()
        page.goto('https://iproyal.com/blog/')

        # Use CSS selector to select all titles inside <h2> tags with the class 'tp-headline-s'
        titles = page.query_selector_all('h2.tp-headline-s')
        titles_text = [title.text_content().strip() for title in titles]

        print('Article Titles:', titles_text)

        browser.close()

scrape_titles()

Scraping Images

The Playwright package will allow you to extract any type of data from a website, images included. Take note, however, that images might take up a lot more space than textual data, so you might need some optimization strategies.

We only need a few tweaks to our previous code to download images from the blog:

const { chromium } = require('playwright');
const fs = require('fs');
const path = require('path');

(async () => {
  const browser = await chromium.launch();
  const page = await browser.newPage();
  await page.goto('https://iproyal.com/blog/');

  // Collect absolute image URLs, skipping inline data: URIs
  const images = await page.$$eval('img', imgs =>
    imgs.map(img => img.currentSrc || img.src).filter(src => src.startsWith('http'))
  );

  for (let i = 0; i < images.length; i++) {
    // page.request downloads the file without navigating away from the page
    const response = await page.request.get(images[i]);
    if (!response.ok()) continue;
    const extension = path.extname(new URL(images[i]).pathname) || '.jpg';
    await fs.promises.writeFile(`image-${i}${extension}`, await response.body());
  }

  await browser.close();
})();

We first include “fs” (File System) in our code to allow us to make use of the *writeFile *function within Node.js.

Everything else is the same until we hit the “img” CSS selector. We need the URLs of images, which are stored in each image’s “src” attribute.

The “imgs” parameter holds every element matched by the “img” CSS selector, and the “map()” function extracts the URL from each one. We also filter out inline data: URIs, which aren’t downloadable files.

After that, we create a “for” loop that matches the length of our “images” variable. Our loop takes each URL, goes to the page, and uses writeFile to create a file from the image. They’re all downloaded to the default directory (usually, the project directory).

For Python, we’re following an identical process, except that open() creates a file object we write each image’s bytes into:

import os
from urllib.parse import urlparse
from playwright.sync_api import sync_playwright

def scrape_images():
    with sync_playwright() as p:
        browser = p.chromium.launch()
        page = browser.new_page()
        page.goto('https://iproyal.com/blog/')

        # Collect absolute image URLs, skipping inline data: URIs
        image_urls = page.locator('img').evaluate_all(
            'imgs => imgs.map(img => img.currentSrc || img.src).filter(src => src.startsWith("http"))'
        )

        for i, image_url in enumerate(image_urls):
            # page.request downloads the file without navigating away from the page
            response = page.request.get(image_url)
            if not response.ok:
                continue
            extension = os.path.splitext(urlparse(image_url).path)[1] or '.jpg'
            with open(f'image-{i}{extension}', 'wb') as file:
                file.write(response.body())

        browser.close()

scrape_images()

Pagination

You need different strategies if you want to navigate through multiple pages, depending on the site architecture. For a standard next button setup, locate the Next Page element, click it, and loop the action until the button becomes disabled or vanishes from the DOM.

Numbered pagination gives you two options:

  • If the site is server-rendered, you can extract the total page count and iterate through a list of URLs.
  • If the site is a single-page application where the URL remains static, you will need to script clicks on each specific page number instead.

For infinite scroll layouts, your code needs to scroll the window down and wait for new network requests to finish or for new elements to render before scrolling again.

Concurrency

To scrape at scale, you need to run tasks in parallel, and Playwright handles it efficiently through browser contexts. Rather than launching a heavy browser process for every thread, you can open multiple lightweight contexts from a single browser.

Creating multiple pages inside one single context shares session data like cookies and cache. During large scrapes, you should spin up separate contexts for each task, and you can execute these concurrently using Python asyncio or Node Promise.all.

Always set strict concurrency limits like 5 to 10 active tasks, as too many concurrent pages will exhaust your system memory or trigger rate limits on the target server. When saving large media files, stream the binary data asynchronously to keep the main thread unblocked.

Retries and Error Handling

Network requests and page loads often fail during a scrape, so your code needs to catch specific Playwright exceptions like TimeoutError when navigation stalls or target elements fail to render.

When a request fails due to a rate limit or server error, use exponential backoff. Waiting progressively longer between retries gives the target server time to recover.

Keep in mind that delaying your requests will not stop bot protection systems from permanently blocking your IP address. To prevent permanent bans, you must combine your retry logic with proxy rotation so each attempt comes from a fresh IP.

Waiting and Dynamic Content

Extracting data from modern websites requires waiting for asynchronous elements to load.

Modern web applications constantly run background trackers, which means the network rarely stops entirely. Relying on an idle network state will frequently cause your script to time out. Older methods like wait_for_selector are also outdated. You should use Playwright Locators instead. Locators automatically wait for a specific element to become visible and actionable before clicking or extracting text.

To avoid page loading times completely, skip the JavaScript rendering. You can intercept the background network traffic and capture the raw JSON API responses directly. The ability to interact with dynamic elements and intercept network requests is exactly why you need browser automation instead of a basic HTTP client.

How to Set Up and Use Proxies in Playwright

Whenever you’re engaged in web scraping at scale, you’ll need proxies. No matter how well you optimize your Playwright scraper, whether you tune request rates, switch between a headful and headless browser, or adjust user agents, a single IP address sending many requests will eventually run into rate limits or IP bans.

A proxy server routes your requests through a different IP address, and a rotating pool spreads traffic across many addresses so no single IP carries the full load. Playwright has native proxy support, which is incredibly easy to make use of.

All you need to do is change a few lines when launching Playwright:

const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch({
    proxy: {
      server: 'http://proxy-server:port',
      username: 'username',
      password: 'password'
    }
  });

  const page = await browser.newPage();
  await page.goto('https://example.com');
  console.log(await page.title());
  await browser.close();
})();

You simply feed a few launch options related to proxies, and that’s it. If your proxy provider does not support proxy rotation, however, you may need to create custom logic.

Here’s a Python example:

from playwright.sync_api import sync_playwright

def launch_browser_with_proxy():
    with sync_playwright() as p:
        browser = p.chromium.launch(proxy={
            'server': 'http://proxy-server:port',
            'username': 'username',
            'password': 'password'
        })
        page = browser.new_page()
        
        # Example: navigate to a website to test the proxy
        page.goto('https://example.com')
        print(page.title())

        browser.close()

launch_browser_with_proxy()

FAQ

Can Playwright be detected?

Yes. Websites can recognize automated traffic from any browser automation or HTTP request library, including Playwright, through signals such as request frequency, IP reputation, and browser properties like navigator.webdriver. What you can control is how your scraper behaves: keep request rates reasonable, limit concurrency, follow the site’s terms of service, and route traffic through rotating residential proxies so requests are distributed across many IP addresses.

Does Playwright have a UI?

Playwright runs as a headless browser by default, but you can open a visible browser window by setting headless to false. It also ships with graphical tools, including the Playwright Inspector for stepping through scripts, Codegen for recording actions and generating locators, and the Trace Viewer for reviewing a run step by step. Playwright Test adds a UI Mode for running and debugging tests.

Is Selenium better than Playwright for web scraping?

For most modern web scraping projects, Playwright is the more practical choice. It offers built-in auto-waiting, browser contexts for lightweight session isolation, and network interception out of the box. Selenium still has advantages: a longer track record, support for more programming languages, and a larger archive of community tutorials. It remains a solid option if your team already uses it.

Can Playwright scrape JavaScript-heavy sites?

Yes. Playwright operates full web browsers, which means it renders client-side JavaScript and fetches dynamic network requests natively.

Why does my Playwright scraper get blocked?

The most common causes are request volume and IP reputation. Sending many requests from a single IP address in a short time triggers rate limits, and datacenter IP ranges are often treated more strictly than residential ones.

Websites may also check automation markers such as navigator.webdriver, or inconsistent settings like a user agent that doesn’t match the rest of the browser environment. To reduce blocks, add concurrency limits, retry with exponential backoff, and use a proxy server pool with rotating residential IPs.

Can I run multiple Playwright browsers at once?

Yes, but launching full browser instances consumes massive amounts of memory. To maximize performance, launch one main browser instance and open multiple isolated browser contexts.

Should I use headless or headful mode?

Headless mode runs the browser in the background without a graphical interface, whereas headful mode opens a visible browser window. Headless mode is faster and uses fewer system resources, which makes it the usual choice for production scraping. Headful mode is best when you’re writing and debugging your extraction logic, since you can watch every click and scroll. Some websites respond differently to a headless browser, so if a page renders incorrectly or returns errors in headless mode, run the same script in headful mode to compare the behavior.

Create Account
Share on
Article by IPRoyal
Meet our writers
Data News in Your Inbox

No spam whatsoever, just pure data gathering news, trending topics and useful links. Unsubscribe anytime.

No spam. Unsubscribe anytime.

Related articles