Court dismisses Google's DMCA lawsuit against scraper SerpApi.
Google established itself as one of the largest companies worldwide by scraping the entire internet without seeking permission initially. This week, a court ruled that Google cannot prevent others from scraping its search results.
On July 20, a federal judge dismissed Google’s lawsuit against SerpApi. This company scrapes Google search results and sells them as structured data. Google had filed the lawsuit in December under the Digital Millennium Copyright Act (DMCA), claiming that SerpApi bypassed its anti-bot defenses to extract copyrighted content. The court determined that the search results are not considered copyrighted material.
Why the case failed
The DMCA provision invoked by Google, Section 1201, only safeguards technology that protects copyrighted works. Google’s protective system, known as SearchGuard, guards its search results. Judge Yvonne Gonzalez Rogers ruled that basic results such as URLs, snippets, and factual index data are public facts and not copyrighted works. Those claims were dismissed without an opportunity to refile.
“Google has not provided a plausible violation of the DMCA,” she stated.
The judge was critical of SerpApi’s techniques, acknowledging that spoofing browser fingerprints, rotating IP addresses, and solving CAPTCHAs to bypass SearchGuard constitutes circumvention. However, it is not unlawful unless the barrier restricts copyrighted works with the owner's consent.
Google’s current predicament
Google does have a limited avenue to pursue. The court granted it 21 days to refile on a narrower claim concerning licensed snippets that occasionally appear in its knowledge panels.
To succeed, Google would need to assert that these panels contain copyrighted content it is entitled to protect. Meredith Rose, a DMCA expert at the nonprofit Public Knowledge, indicated that making such a claim could be risky for Google.
If Google contends that compiling these panels reproduces copyrighted material, it could prompt rights holders whose content appears in its results, as well as its AI Overviews, to sue Google for the very scraping it is contesting.
“They have gotten themselves into a bit of a predicament,” Rose told Ars Technica. Hence, to win this minor battle, Google might have to concede on a much larger scale.
Google has stated it intends to proceed. Spokesperson José Castañeda mentioned that the company was “pleased to see that the Court rejected nearly all of SerpApi’s legal arguments” regarding standing, and Google plans to submit an amended complaint.
What Google aimed to protect
The value of search data has never been higher. A chatbot cannot summarize online information without access, and Google does not provide an official search API. This lack of direct access makes third-party services like SerpApi the primary means to reach Google’s index. Its clients include Nvidia, Uber, Adobe, and the AI search engine Perplexity.
SerpApi characterized the dismissal as significant beyond its immediate implications. It stated that Google and Reddit “do not own the Internet,” accusing both of attempting to “weaponize the DMCA to restrict access to the open Internet.” Its CEO Julien Khaleghy pointed out that the damages Google proposed, if taken literally, could exceed the entire US economy.
This irony is clear, and critics have highlighted it. Google has spent two decades scraping publishers’ work to build its search dominance while now attempting to use copyright law to prevent a company from doing similarly.
Reddit’s scrutiny and exposure
Google was not the first to take this approach. In October, Reddit initiated a nearly identical DMCA suit against SerpApi and Perplexity concerning Reddit content visible in Google results. Facing its own dismissal hearing shortly after Google’s defeat, Reddit finds itself in a difficult position. It is neither the copyright holder nor the exclusive licensee of the content appearing in search results, which was the specific issue that undermined Google’s case.
SerpApi has indicated that Reddit aims to function as a “toll collector,” charging for access to user-generated content. Reddit has been increasingly restricting access to its data for some time. (Advance Publications, the parent company of Ars Technica, is Reddit's largest shareholder.)
The larger battle for the open web
Rose provided a broader context for the case. Since around 2023, there has been a surge in aggressive AI scraping. Content publishers have raced to protect their material, leading to a “re-enclosure” of a web that was primarily open.
This defensive reaction halts not just AI training but also disrupts anonymous crawling essential for research, archiving, journalism, and public health reporting. The safeguards designed to prevent scraping end up affecting everyone.
The stakes mirror those involved in Google’s AI search strategy and its confrontations with regulators regarding search data. The question of who gets to access the public web on a large scale and who gets to monetize that access is being resolved on a case-by-case basis. Google’s control is already beginning to diminish in the AI era, and it has 21 days to evaluate whether this confrontation is worth the
Other articles
Court dismisses Google's DMCA lawsuit against scraper SerpApi.
A judge dismissed Google's DMCA lawsuit against SerpApi, stating that search results are public information. Refilling the case might reveal Google's own scraping practices.
