1. Will Dewatermarking be Illegal?
By: Andrew Gamino-Cheong
Within hours of Anthropic announcing its text watermarking approach for the EU AI Act, tools emerged claiming to strip it. None can be verified yet, since only Anthropic can detect the marks. Tools that strip C2PA provenance metadata from images have existed longer and work more reliably, since that only requires editing a file.
The legal question and the ethical question turn out to be the same question. Every regime that has confronted a circumvention tool has asked whether the tool has uses beyond breaking the rules. So the thing worth arguing about isn’t whether dewatermarking is technically possible. It’s whether anyone has a good reason to do it.
Some plausible ones already exist:
Watermarks can carry more than “this is synthetic.” Provenance metadata may reveal which model was used, when, and how much AI was involved, which is real exposure for a lawyer, a journalist, or a dissident.
Marks can attach to work that’s substantially human. Providers may flag content that was merely proofread or translated by a model, so removal may be a correction rather than a deception.
Security researchers need to break watermarks to evaluate whether they actually work.
Against that, the case for prohibition is straightforward. Watermark removal is the enabling step for most synthetic-content fraud, and provenance is worthless if it’s optional.
The two major jurisdictions are resolving that tension differently. The EU has landed on a permissive answer for tools and a strict one for institutions. Its top court ruled in July 2026 that VPN providers aren’t liable when users bypass geo-blocks, calling VPNs “lawful technical tools.” The UK reached a similar place by politics rather than law, floating a VPN ban after Online Safety Act age checks triggered a signup surge and then abandoning the idea. Meanwhile the AI Act’s Code of Practice makes AI companies write watermark protection into their terms of service and bars their business customers from stripping marks. That reaches contracted parties and misses everyone else. A developer who posts a dewatermarking script on GitHub never signed anything.
The US has the opposite shape. Section 1201 of the DMCA, written to stop DVD copy protection cracking, bans distributing a circumvention tool if it was mainly designed for that purpose or marketed that way, regardless of who uses it. Legitimate secondary uses don’t save it, which is how DeCSS lost. Applying that to watermarks requires treating a provenance signal as copyright protection, which is a stretch. But if a court accepts it, the legitimate-use defense that protects VPNs in Europe doesn’t transfer.
Key Takeaway: Whether dewatermarking stays legal depends on whether courts accept that removing a mark can be honest. Europe has left that door open for tool builders while closing it for companies, and the US may do the reverse.
2. Tech Explainer: OpenAI’s HuggingFace Incident
By: Anastassia Kornilova
Several days ago, OpenAI and METR released parallel reports on last month’s breach of Hugging Face’s production systems. During this attack, agents built on IM1, an internal-only OpenAI research model, broke out of a cybersecurity evaluation called ExploitGym and ran a days-long, unauthorized attack on Hugging Face’s infrastructure. These are the key components that led to the attack:
Goal Manipulation: The agents involved in the attack followed a common pattern: after failing to solve their assignment directly, they discovered an unsanctioned message board inside a shared system and joined a collective effort to hack into Hugging Face. Different agents worked on workstreams that dealt with subproblems and risked failing their own assigned task. The agents had diverse reasons for joining the project and many didn’t understand why hacking HuggingFace was necessary, but helped anyway. This shift shows that agents are vulnerable to manipulation and can have their goal hijacked. In this case, the change was unintentionally set by other agents, but the same can happen through attacks by malicious agents. In fact, OWASPS Top 10 for Agentic Applications highlights this as a risk and suggests mitigations.
Incentive to Keep Going: During post-training, models are rewarded for finding solutions to problems; they don’t have an incentive to claim that a problem is unsolvable and stop. Only 60-70% of the tasks in ExploitGym are believed to be solvable, so upon discovering an impossible task, models are incentivized to get creative. Here, the agents first gained access to the original ExploitGym paper and code repository and figured out how to reverse engineer the solution. However, the paper stated that the scoring system would check whether the task was solved correctly (not just look at the answer), so the agents figured that they needed a way to study and fool that system, and realized that they could hack into HuggingFace to access a copy of the scoring system. This isn’t unique to a complex benchmark task. In August, an OpenClaw agent running on Anthropic’s Claude hacked a gym’s booking system in Melbourne after being asked to help book a class, then bumped another member off a waitlist to make room for its user, entirely on its own initiative, on a task with nothing like ExploitGym’s stakes. To address this, model providers need to change their post-training process and build graders that reward a model for correctly flagging a broken or impossible task and stopping, instead of only rewarding it for producing an answer.
Insufficient Monitoring: OpenAI didn’t use monitoring techniques for this evaluation that is used by their production systems. This included both traditional security monitoring and chain-of-thought monitoring that reviews intermediate reasoning traces output by models. Security monitoring helped OpenAI detect other malicious behavior during an earlier incident in June and a later incident in July, but their set-up missed the warning signs in this attack. While OpenAI’s reports state that their production chain-of-thought monitoring system could have caught the attack, this methodology may not always be sufficient: earlier research suggests that traces don’t always reflect intentions and the actual reasoning process, and METR investigation showed multiple attempts by agents to manipulate the logs to hide actual behavior and spoof tool calls. Stronger security monitoring practices may soon be necessary across all agents. This is one area where AI can help with defence, Hugging-Face used AI-assisted methodology to detect and analyze the attacks on their end.
Key Takeaways: This HuggingFace attack resulted largely from gaps in OpenAI’s security practices during the ExploitGym evaluation, related to relaxed monitoring and insufficient investigation into related preceding incidents. However, this incident highlights the general advancement in capabilities from the most models in areas related to long-running tasks, agent coordination and cybersecurity. A fool-proof set of mitigations doesn’t exist: agents have found ways to circumvent a safe and restricted set of tools, human review can be unreliable, and tool-call logs can be spoofed if an agent controls its own execution environment.
3. Trustible Spotlight: Joining PACT AI
By: Lauren Madden
We recently announced that Trustible is a founding member of the newly established Partnership for Assurance, Credibility, and Trust on AI (PACT AI), a coalition working to build the infrastructure needed to verify that AI systems work as designed. PACT AI’s founding members span enterprises, technical experts, AI insurers, and civil society leaders working across the AI economy.
Both consumers and companies deserve to know that the AI systems powering the economy are working as designed, and that’s why the coalition is committed to building the shared standards and practical evidence that makes trust provable for the whole economy. Learn more.
4. Policy Updates
By: Sydney Cullen
G20. At the G20 Innovation Ministerial in North Carolina this week, the Trump administration is pressing member countries to avoid new AI regulatory bodies and instead sign onto a set of principles dubbed the “Carolina Principles.” The pitch is tied to competition with China. Officials argue that heavy regulation abroad risks pushing countries toward Chinese open-weight models instead of the American stack. Sam Altman, Jensen Huang, Elon Musk, and Demis Hassabis are all appearing at the meeting alongside Commerce Secretary Lutnick.
Our Take: While it is logical to focus governance efforts on novel risks rather than over-burden developers and deployers with redundant rules, the existing novel risks from AI still haven’t been adequately governed. This “hands-off” push comes less than a week after OpenAI published its technical report on how a swarm of nearly 700 of its agents escaped an isolated testing environment during an internal cybersecurity evaluation and compromised Hugging Face’s infrastructure, an incident that rattled the industry and drew bipartisan attention in Congress. Until existing novel risks like autonomous agents evading their own safety controls are actually addressed, statements like this are hard to take at face value.
California. SB 1119, known as Adam’s Law, cleared the legislature and awaits Governor Newsom’s signature. The bill requires companion chatbot operators to conduct child safety risk assessments prior to deployment or when a substantial modification is made to a system starting July 1, 2027, and submit to independent audits reported to the Attorney General, building on the disclosure-based framework established by last year’s SB 243. In a break from the industry’s usual posture, OpenAI publicly endorsed the bill on August 31, and CEO Sam Altman reportedly called Newsom directly to encourage him to sign it.
Our Take: SB 1119 requires a risk assessment before an operator makes a “new or substantially modified” companion chatbot available, but never defines “substantial.” The EU AI Act has the same trigger for high-risk systems, and defines it as a change not foreseen in the original assessment that affects compliance or shifts the system’s purpose. However, Californian operators are left guessing whether a model upgrade, new persona, or memory feature counts, and with a private right of action allowing minors and guardians to sue for damages, including punitive damages, guessing wrong is costly. Operators should document their own definition of “substantial modification” now rather than wait for a regulator or a plaintiff’s attorney to set it for them.
Australia. APRA and ASIC issued a joint statement urging financial market entities to take concrete steps against cyber and operational threats driven by frontier AI. The release landed a day after Australia’s National Cabinet agreed to legislate mandatory AI and data center standards by early 2027, suggesting supervisory expectations for regulated firms are hardening well ahead of that framework taking effect.
Our Take: Australia’s approach here is a preview of what enforcement looks like in the absence of AI-specific legislation: use the levers you already have. APRA and ASIC aren’t creating new AI rules, they’re telling regulated entities that AI risk already falls under existing prudential standards like CPS 230, and that boards relying on vendor presentations instead of genuine technical literacy will be treated as a governance failure.
In case you missed it, a few additional developments:
AI and Federal Recordkeeping. The National Archives clarified that agencies’ use of AI platforms doesn’t automatically create a federal record. That determination depends on whether the material is relied on for decisions, circulated, or incorporated into agency systems. The guidance pays particular attention to audit trails, noting they’re treated as personal files unless an agency captures and uses them for official business.
FTC Seeks Comment on Personalized Pricing Enforcement. The FTC published a draft enforcement policy statement addressing AI-driven personalized pricing, arguing that failing to disclose the practice to consumers can violate the FTC Act’s unfair-or-deceptive-practices standard. This is inline with states like Maryland and New Jersey who have recent prohibitions on personalized pricing as well.
Bill Gates Warns There’s “No Plan” for AI’s Economic Disruption. In a nearly 6,000-word essay, Microsoft co-founder Bill Gates argued the world is unprepared for the labor market disruption AI will cause, predicting law, medicine, customer service, software, and manufacturing will see significant impact within a decade rather than over generations. He proposed policy responses including taxing AI-driven revenue to fund a safety net and reserving a share of jobs for humans.
—
As always, we welcome your feedback on content! Have suggestions? Drop us a line at newsletter@trustible.ai.
AI Responsibly,
- Trustible Team



