Claude is becoming ambitious with watermarking, and I can sense the issues from far away.
Anthropic aims to make it simpler to identify AI-generated text, which I find generally acceptable. The company is testing an invisible watermark that can be integrated directly into text produced by Claude.
This seems like a logical approach. AI-generated text is ubiquitous, and understanding its origin could certainly provide insight. Additionally, Anthropic’s method doesn't just conceal a marker within a document; it modifies how Claude selects words to create a detectable statistical pattern.
However, one aspect concerns me. Anthropic is examining the persistence of this watermark, even after text alterations.
This aspect raises potential issues for me.
Nadeem Sarwar / Digital Trends
Did Claude influence my writing?
Consider the scenario of translation. Imagine someone writes an entire essay in Spanish and requests Claude to translate it into English. The ideas belong to the individual, the research is theirs, and so are the arguments. Claude's sole role is to translate.
Yet, the translated text could still bear Claude’s watermark.
The same applies to proofreading. What if someone authors a piece and asks Claude to correct the grammar? What if they need to shorten a paragraph, alter its tone, revise dictated content, or simply enhance clarity of a clunky sentence?
These scenarios are no longer niche applications for AI. People increasingly rely on assistants like ChatGPT, Gemini, and Claude for routine tasks that do not involve creating original content. A watermark might indicate that Claude played a role in producing a text, but it doesn't clarify whether Claude was responsible for the actual writing. Anthropic recognizes this, stating that the watermark signifies involvement, not the original creator's identity.
Now, envision trying to articulate that distinction to a professor after their detection tool flags your essay.
Unsplash
We already know how unreliable AI detection can be
I would be less concerned if our history with AI detection were particularly reliable. Unfortunately, it is not.
MIT Sloan offers clear guidance regarding current AI detectors, noting their high error rates and the potential to lead educators to mistakenly accuse students of misconduct.
We've seen the implications of this in the real world. Students have had to defend their own work after automated systems labeled it as AI-generated. In one documented case by The Guardian, a student's essay was marked as entirely AI-generated even though the student claimed they only utilized approved spelling and grammar aids. The appeal was ultimately accepted.
To clarify, Claude’s watermark represents a different concept. Traditional AI detectors analyze writing to gauge the likelihood that an AI generated it. In contrast, Anthropic is intentionally embedding a detectable signal in Claude’s output. In theory, this should enhance the system’s reliability. However, reliability isn't the sole concern; interpretation is also critical.
Nadeem Sarwar / Digital Trends
Using AI to prove we didn’t use AI
We're already in a somewhat absurd situation.
Students anxious about AI detection are resorting to what’s called AI humanizers, which rephrase text to reduce the likelihood of detector triggers. Some students even apply these tools to their own work due to worries about false positives. Naturally, detector companies are developing methods to spot humanizers.
Read that again.
A human can create something, worry that an AI will mistakenly attribute it to AI, modify it through another AI to enhance its human-like qualities, and then have yet another system evaluate whether AI made it appear more human.
It's a technological ouroboros.
Creating a watermark for Claude that can withstand editing and translation is an impressive technical feat. Previous studies have indicated that some text-watermarking methods can be circumvented via translation, so addressing that vulnerability represents significant progress.
Nevertheless, I believe that making the signal more difficult to remove doesn’t resolve the core issue.
Rachit Agarwal / Digital Trends
Context is essential for a watermark
There are valid reasons to apply watermarks to AI-generated content. It could aid in identifying widespread misinformation, undisclosed synthetic text, or AI-written materials that find their way into training datasets.
The challenge is that AI tools now perform far more than simply generating content from scratch. Users employ them to translate text, proofread documents, summarize research, assist with coding, enhance accessibility, or tidy up an email before sending it. In this context, recognizing AI’s involvement does not automatically indicate who the true creator is.
Each of these interactions involves AI to varying degrees. Knowing that Claude generated an essay from start to finish is beneficial. However, if Claude translates an essay created by someone who devoted three weeks to researching and writing it, realizing Claude was involved provides much less insight.
The watermark might effectively answer “Did Claude influence this?” My concern lies in what happens when people start interpreting that answer as proof of “Did Claude author this?”
Anthropic can engineer the most advanced watermark, but if users fail to comprehend that distinction, I fear we will encounter significant issues.
Other articles
Claude is becoming ambitious with watermarking, and I can sense the issues from far away.
Anthropic is working on the challenge of detecting AI-generated text, but a constant watermark might lead to a new issue if AI support is confused with AI authorship.
