1. How should AI Governance Professionals React to the AI Safety Debate
By: Andrew Gamino-Cheong
The AI Safety debate reached fever pitch levels over the past few weeks following numerous reports of additional errant agent swarms that successfully caused cyber incidents around the world, a report on the ways Claude models have been misused by both State and non-State actors, and the recent departure of an AI safety researcher at Anthropic. AI Lab leaders posted essays on the need to slow down and there were calls for a special session of Congress just to address AI. Most of these discussions focus on frontier AI systems with advanced capabilities that pose catastrophic risks. Governing and regulating these systems will become important, and while the debate and drama makes for great media, it isn’t always clear how these discussions and novel frontier models may impact the day to day of AI Governance practitioners. Most AI Governance professionals are deep in the trenches trying to collect and review use case documentation, figure out whether to approve a new MCP server integration, or try to determine how to monitor their chatbot systems after deployment, rather than solve AI alignment problems. This begs the question, what are some actionable takeaways from the last few weeks of AI safety headlines for the day to day AI Governance professional inside an enterprise? Here are a few of our thoughts:
Limit Coding Ability
While AI is truly impressive in its ability now to write code, this capability introduces a potentially boundless amount of risk. Agentic coding with a human in the loop integrated into a software development lifecycle is fine, but an agent with the ability to write and execute its own code autonomously in production will not be within the risk tolerance of most organizations.
Set Cost / Time Limits
If a specific task cannot be done by an agent within a certain amount of time, or within a certain budget, then it may not be worth doing with an agent. Most of the time when an agent spends a lot of time on a problem, it’s because they encountered some obstacle like blocked network access, insufficient permissions, etc. Trying to fix these issues is where many agents have gone off the rails, and it’s often unintended as well. Setting fixed time/cost limits is just good business practice, and can also help catch a blocked agent before they try and go around legitimate protections.
Invest in Secure Sandboxes
At least some of the evaluations that broke out were able to do so because they were not running in a fully isolated environment. In some cases, the agents actually found a way to use a shared package repository they had access to in order to coordinate work. This shows that advanced AI evaluations likely need a much more secure environment to work inside of, with a higher degree of scrutiny of their infrastructure to prevent any attempted external access.
Evaluation Log Sampling
OpenAI let an evaluation run for days without anyone checking and analyzing the logs, and during that time, their agent swarm managed to hack HuggingFace. Complex evaluations that run over a long time should have some mandatory check-ins to review logs, and ensure things are running as expected.
Minimally Capable Model Necessary
In cybersecurity, there’s a concept known as ‘zero trust’ which implies that a user should be given minimum platform permissions necessary for them to do their job. The appropriate AI equivalent may be using the ‘least capable’ model that can still achieve its main goal. A lot of systems in production in particular will be automating basic tasks, and use AI only to smooth some of the uncertainty that sometimes happens. Using highly capable models like Mythos on these systems will be both cost inefficient, and actually increase the risk of the system considerably.
Plan for Disclosure Requirements
While US Federal action on AI is still unlikely, many States are responding to immense voter demand for action on AI. Data centers will be the first target, but AI disclosure will be a close second. Policymakers aren’t innovators and don’t necessarily know how to regulate AI, so they instead will reach for highly visible things that people want that will be easy to implement, and enforce. Disclosure of AI use will be the minimum baseline, and as more States adopt policies around it, others will piggyback.
Key Takeaway: The AI Safety debate is going to continue to consume the media’s attention, but there are actionable takeaways from some of the recent incidents for most enterprise AI Governance professionals, including how to prevent similar run-away agent swarms.
2. Trustible Spotlight
By: Lauren Madden
Introducing Trustible’s AI Monitoring Hub
An approved AI use case rarely stays the same. Vendors update models, usage expands, and output quality can drift long before a customer raises a concern or a regulator asks for evidence of oversight. Trustible’s new AI Monitoring Hub gives governance teams a structured way to catch those changes early, using the same use case record created during intake.
Each use case now includes a Monitoring tab that groups key metrics by category, including cost, quality, drift, and security. Every metric has a clear owner and a visible trend over time. Teams can track both internal and external signals, and when a metric exceeds its threshold, an alert is triggered to launch a formal investigation.
If you want to see how it works, watch our 3-min walkthrough inside a live use case record.
Your First 6 Months as an AI Governance Leader: A Checklist
Building an AI governance function from scratch? There is a lot to do, but achievable within 6 months if you have the right plan.
Start by grounding your policy in real regulatory obligations. Then build a governance structure with actual decision rights, centralize your AI inventory, and create an intake process that is easy to use for your team. From there, the job is to prove the program is efficient and build the case for more resources.
Our new blog post walks through all priorities in your first 6 months as an AI governance leader. Read the checklist.
3. Policy Updates
By: Sydney Cullen
Apple - Apple released iOS 27 on September 14 with a redesigned Siri AI, an opt-in beta that pulls from a user’s messages, emails, and calendar to answer requests locally, then reaches out to the cloud for open-ended or current-events questions. Apple’s announcement describes it as drawing on “personal context understanding” across a user’s own content, plus a new standalone Siri app that saves conversation history and syncs it across devices via iCloud. CNBC reports the cloud-based features carry daily usage limits, with expanded limits planned as a paid tier.
Our Take: Most shadow AI policies weren’t written for an OS feature that reads Mail and Messages by default once a user opts in. IT and security teams should confirm what MDM controls actually exist for this feature before assuming existing policy covers it, since the enterprise-management story here is still unsettled.
California - Governor Gavin Newsom signed SB 813 and AB 1405 on September 9, creating California’s first state-recognized framework for evaluating AI systems from the outside rather than relying on developers to self-report. SB 813 sets up a framework for independent verification organizations to assess AI systems and models for compliance with state law. AB 1405 creates a state registry for AI auditors with standards for their independence and integrity. Newsom also signed SB 1119 chatbot bill earlier this month and has a September 30 deadline for the rest of this session’s AI bill pile-up, including the SB 1000 transparency amendments.
Our Take: Organizations are currently utilizing ISO 42001 certifications as their AI governance certifications, and coming to compliance with a state specific certification could become burdensome. Likely we will see mandatory evaluations, like in Illinois, take priority for organizations over voluntary certifications like in California.
Texas - Texas AG Ken Paxton’s Senate campaign released an ad featuring an AI-generated version of Democratic opponent James Talarico “reading” a fictional Bible and reciting a string of remarks Talarico actually made in interviews, floor speeches, and sermons between 2021 and 2023. FactCheck.org’s review found the underlying quotes are real, but the ad strips out the surrounding context (a point about Hebrew grammar in Genesis became “God is nonbinary”; a comment on chromosomal variation became “God said, let there be six genders”) and stitches remarks from different years and settings into one continuous monologue delivered by a synthetic Talarico. A disclaimer noting AI use appears in small text roughly 20 seconds into the 30-second spot. Talarico has since walked back some of the original comments himself, telling Fox News he “didn’t say it correctly” on the sex-chromosome point.
Our Take: Texas’s own deepfake statute turns entirely on “intent to deceive,” the same standard courts have used elsewhere to carve satire and parody out of similar laws in Michigan, Oregon, and California. This ad tests that assumption: nothing about its presentation signals parody, it’s built to look like a real person’s words in his own voice. It also uses his own words, but lacks context and chops them together to seem like they occurred at the same time. Oregon is litigating a similar question right now over unlabeled synthetic attack ads from a former congressional candidate, and the outcome there will say something about how far “intent” alone can carry these laws. Separately, a disclaimer that surfaces for a few seconds in small type toward the end of a cable spot is easy to miss for any viewer, but would prove to be especially difficult for older viewers less practiced at spotting synthetic media, which matters here given the ad’s run on cable. Both issues will test how transparency laws and deep fake laws in the US are interpreted, especially as the midterms approach.
In Case You Missed It:
China - China’s Supreme People’s Court issued a 24-article opinion on September 7 instructing lower courts on how to apply existing civil law (the Civil Code, Personal Information Protection Law, and others) to AI disputes, rather than creating new statute. The guidance states that producing an identifiable digital likeness or cloned voice without consent can infringe personality rights, and extends to platform liability for failing to remove flagged content, algorithmic price discrimination, and AI-generated evidence. Notably, the court sidestepped whether AI-generated content itself qualifies for copyright protection.
Congress - Jacob Coxon resigned from Anthropic on September 8, telling more than 150 million X viewers that AI labs are “gambling with our lives,” and more than 20 members of Congress publicly called for tougher oversight within days. But with the House beginning their 6 week (now 7 week) recess early, Axios reports House leadership has shown no urgency, and many Democrats are framing binding AI legislation as 2027 business. Separately, House Democrats are floating a select committee on AI with subpoena power, contingent on retaking the House majority in November.
Payment Firms - The three payment networks announced on September 10 a joint “Know Your Agent” framework, an interoperability layer that would let a payment agent verified with one network be recognized across others, building on each company’s existing agent protocols (Visa’s Trusted Agent Protocol, Mastercard’s Verifiable Intent, Ant’s Agentic Mobile Protocol). The initiative is meant to let merchants and issuers confirm an agent’s identity, permissions, and transaction limits before approving a purchase, ahead of McKinsey’s projected $3 trillion to $5 trillion in AI-agent-handled commerce by 2030.
—
As always, we welcome your feedback on content! Have suggestions? Drop us a line at newsletter@trustible.ai.
AI Responsibly,
- Trustible Team


