Court dismisses Google's DMCA lawsuit against scraper SerpApi.

Court dismisses Google's DMCA lawsuit against scraper SerpApi.

      Google established itself as one of the largest companies worldwide by scraping the entire internet without seeking permission initially. This week, a court ruled that Google cannot prevent others from scraping its search results.

      On July 20, a federal judge dismissed Google’s lawsuit against SerpApi. This company scrapes Google search results and sells them as structured data. Google had filed the lawsuit in December under the Digital Millennium Copyright Act (DMCA), claiming that SerpApi bypassed its anti-bot defenses to extract copyrighted content. The court determined that the search results are not considered copyrighted material.

      Why the case failed

      The DMCA provision invoked by Google, Section 1201, only safeguards technology that protects copyrighted works. Google’s protective system, known as SearchGuard, guards its search results. Judge Yvonne Gonzalez Rogers ruled that basic results such as URLs, snippets, and factual index data are public facts and not copyrighted works. Those claims were dismissed without an opportunity to refile.

      “Google has not provided a plausible violation of the DMCA,” she stated.

      The judge was critical of SerpApi’s techniques, acknowledging that spoofing browser fingerprints, rotating IP addresses, and solving CAPTCHAs to bypass SearchGuard constitutes circumvention. However, it is not unlawful unless the barrier restricts copyrighted works with the owner's consent.

      Google’s current predicament

      Google does have a limited avenue to pursue. The court granted it 21 days to refile on a narrower claim concerning licensed snippets that occasionally appear in its knowledge panels.

      To succeed, Google would need to assert that these panels contain copyrighted content it is entitled to protect. Meredith Rose, a DMCA expert at the nonprofit Public Knowledge, indicated that making such a claim could be risky for Google.

      If Google contends that compiling these panels reproduces copyrighted material, it could prompt rights holders whose content appears in its results, as well as its AI Overviews, to sue Google for the very scraping it is contesting.

      “They have gotten themselves into a bit of a predicament,” Rose told Ars Technica. Hence, to win this minor battle, Google might have to concede on a much larger scale.

      Google has stated it intends to proceed. Spokesperson José Castañeda mentioned that the company was “pleased to see that the Court rejected nearly all of SerpApi’s legal arguments” regarding standing, and Google plans to submit an amended complaint.

      What Google aimed to protect

      The value of search data has never been higher. A chatbot cannot summarize online information without access, and Google does not provide an official search API. This lack of direct access makes third-party services like SerpApi the primary means to reach Google’s index. Its clients include Nvidia, Uber, Adobe, and the AI search engine Perplexity.

      SerpApi characterized the dismissal as significant beyond its immediate implications. It stated that Google and Reddit “do not own the Internet,” accusing both of attempting to “weaponize the DMCA to restrict access to the open Internet.” Its CEO Julien Khaleghy pointed out that the damages Google proposed, if taken literally, could exceed the entire US economy.

      This irony is clear, and critics have highlighted it. Google has spent two decades scraping publishers’ work to build its search dominance while now attempting to use copyright law to prevent a company from doing similarly.

      Reddit’s scrutiny and exposure

      Google was not the first to take this approach. In October, Reddit initiated a nearly identical DMCA suit against SerpApi and Perplexity concerning Reddit content visible in Google results. Facing its own dismissal hearing shortly after Google’s defeat, Reddit finds itself in a difficult position. It is neither the copyright holder nor the exclusive licensee of the content appearing in search results, which was the specific issue that undermined Google’s case.

      SerpApi has indicated that Reddit aims to function as a “toll collector,” charging for access to user-generated content. Reddit has been increasingly restricting access to its data for some time. (Advance Publications, the parent company of Ars Technica, is Reddit's largest shareholder.)

      The larger battle for the open web

      Rose provided a broader context for the case. Since around 2023, there has been a surge in aggressive AI scraping. Content publishers have raced to protect their material, leading to a “re-enclosure” of a web that was primarily open.

      This defensive reaction halts not just AI training but also disrupts anonymous crawling essential for research, archiving, journalism, and public health reporting. The safeguards designed to prevent scraping end up affecting everyone.

      The stakes mirror those involved in Google’s AI search strategy and its confrontations with regulators regarding search data. The question of who gets to access the public web on a large scale and who gets to monetize that access is being resolved on a case-by-case basis. Google’s control is already beginning to diminish in the AI era, and it has 21 days to evaluate whether this confrontation is worth the

Other articles

Why Home Smart Strength Training Is Becoming More Accurate Why Home Smart Strength Training Is Becoming More Accurate The contemporary home gym is influenced equally by available space and fitness needs. In apartments, multipurpose areas, and houses where workout gear must coexist with other activities, establishing a large setup can be challenging. This reality has led product design to focus on creating systems that remain compact yet effectively accommodate serious strength training. Reasons for the failure of the Starbucks AI inventory tool at full scale. Reasons for the failure of the Starbucks AI inventory tool at full scale. The AI inventory tool developed by Starbucks was discontinued following a complete national launch. NomadGo, the 30-member startup that created it, received the notification on April 3rd. Researchers have developed a tool capable of detecting the AI utilized to create a counterfeit video. Researchers have developed a tool capable of detecting the AI utilized to create a counterfeit video. Researchers at UC Riverside developed SAGA, a tool that identifies the specific system responsible for creating AI-generated fake videos by utilizing subtle visual patterns as digital fingerprints. Anthropic claims that the leaked Claude conversations functioned as expected. Anthropic claims that the leaked Claude conversations functioned as expected. Claude chats that were shared, including medical documents and children's phone numbers, could be found through Google searches. Anthropic asserts that the feature functioned as designed. How AI is optimizing corporate travel management How AI is optimizing corporate travel management AI is transforming traditional travel procurement portals into predictive systems that enforce policies at the time of purchase, rebook altered travel plans, and automatically negotiate vendor rates. ChatGPT now declines to imitate an author's writing style. ChatGPT now declines to imitate an author's writing style. ChatGPT has discreetly ceased mimicking the writing styles of named authors, including those who have passed away, as OpenAI confronts a series of copyright lawsuits related to its training data.

Court dismisses Google's DMCA lawsuit against scraper SerpApi.

A judge dismissed Google's DMCA lawsuit against SerpApi, stating that search results are public information. Refilling the case might reveal Google's own scraping practices.